{"slug": "study-shows-ai-is-as-good-as-human-tutoring-for-gre-learning-gains", "title": "Study shows AI is as good as human tutoring for GRE learning gains", "summary": "A study submitted to arXiv on 23 Sep 2026 by Curtis Northcutt introduces StudentBench, a suite of AI teaching evaluations built on over 175,000 student-AI messages, and reports that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015) across 2,383 human participants. In five of seven GRE domains the best-performing AI tutor surpassed the human tutor on average, and one AI tutor matched human tutoring gains (p = .044) at 918 times lower cost — USD 0.0052 for AI versus USD 4.81 for human per percentage point gained. A second study of 2,028 pairwise rubric evaluations by expert human tutors separated AI tutors on lesson planning, practice-problem creation, conversational pedagogy, cost and engagement, and the StudentBench platform is freely available at studentbench.org.", "body_md": "# Computer Science > Artificial Intelligence\n\n  [Submitted on 23 Sep 2026]\n\n# Title:StudentBench: AI and human tutoring yield equivalent GRE learning gains\n\n[View PDF](https://arxiv.org/pdf/2609.28470)\n\n[HTML (experimental)](https://arxiv.org/html/2609.28470v1)\n\nAbstract:Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student-AI messages to study whether large language models (LLMs) produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring. We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average. In a second study, expert human tutors compared LLM-generated lesson plans and practice problems through 2,028 pairwise rubric evaluations. Together, the two studies clearly separate AI tutors across: (1) lesson planning, (2) practice-problem creation, (3) conversational pedagogy, (4) cost, and (5) engagement. Surprisingly, one AI tutor achieved learning gains equivalent to human tutoring (p = .044) at 918 times lower cost (USD 0.0052 for AI versus USD 4.81 for human, per percentage point gained). For Quantitative GRE sessions, faster AI replies correlated with more student messages, more messages with more correct practice, and more correct practice with larger learning gains (all p < .002). The StudentBench platform is freely available at [this https URL](https://studentbench.org).\n    \n\n## Submission history\n\nFrom: Curtis Northcutt [\n[view email](https://arxiv.org/show-email/659c2238/2609.28470)]\n\n**[v1]** Wed, 23 Sep 2026 17:57:45 UTC (5,657 KB)\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/study-shows-ai-is-as-good-as-human-tutoring-for-gre-learning-gains", "canonical_source": "https://arxiv.org/abs/2609.28470", "published_at": "2026-09-25 02:01:16+00:00", "updated_at": "2026-09-25 02:29:59.640329+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-products"], "entities": ["StudentBench", "Curtis Northcutt", "arXiv", "GRE"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/study-shows-ai-is-as-good-as-human-tutoring-for-gre-learning-gains", "markdown": "https://wpnews.pro/news/study-shows-ai-is-as-good-as-human-tutoring-for-gre-learning-gains.md", "text": "https://wpnews.pro/news/study-shows-ai-is-as-good-as-human-tutoring-for-gre-learning-gains.txt", "jsonld": "https://wpnews.pro/news/study-shows-ai-is-as-good-as-human-tutoring-for-gre-learning-gains.jsonld"}}