The Tutor Was as Good as a Human. What Did the Student Stop Doing?
A new study finds AI tutors match expert human tutors on GRE learning gains, one of them at a nine-hundredth of the cost. The result is real. The mechanism it reports — faster replies, more messages, more practice — is also the shape of an engagement loop, and we already know what fluent machines do to people's willingness to check.
By Todd McCaffrey
A paper posted on 23 September found that AI tutors produce the same learning gains as expert human tutors on GRE questions, and one of them did it at about a nine-hundredth of the cost. I’ve no reason to doubt it. Over two thousand participants, a proper equivalence test, a public platform anyone can check.
But the detail I keep coming back to is further down the abstract. The faster the AI replied, the more the students typed, and the more they typed, the better they did. That’s a finding about learning. It’s also a description of an engagement loop, and we already know what fluent, fast, agreeable systems do to people’s willingness to check what they’re told.
What StudentBench found
The study, by Curtis Northcutt and five co-authors, put 2,383 people through AI tutoring, expert human tutoring, or no tutoring, on the Quantitative and Verbal sections of the GRE. It logged more than 175,000 student–AI messages.
The headline is an equivalence result: AI tutoring was statistically equivalent to the human tutors on learning gains (p = .015). That’s a harder thing to establish than a difference, because you have to show the gap is small, not merely that you failed to find one. In five of the seven GRE domains, the best AI tutor did better than the human tutor on average.
The number that will travel is the cost. One AI tutor matched human gains (p = .044, a weaker result than the pooled one) at $0.0052 per percentage point gained, against $4.81 for the human. That’s the 918-fold figure.
And then the mechanism. For Quantitative sessions, faster AI replies went with more student messages, more messages with more correct practice, and more correct practice with larger gains. All three links held at p < .002.
What a GRE item is
None of this makes the study wrong. It makes it narrow, and the narrowness is worth saying out loud before the result gets quoted as “AI tutors are as good as teachers”.
A GRE item is well posed, short, and has one right answer that can be checked instantly. That is the terrain where tutoring by any means looks best, because both the tutor and the student know when they’ve got it. It is also terrain where the tutor being wrong is rare and quickly caught.
The abstract reports gains measured in the study. It doesn’t say whether they lasted a week, or whether the student could do the next problem with nobody in the chat. Those may be in the full paper; they aren’t in the summary that will be quoted.
The latency chain
Fast reply, more turns, more practice, more gain. On a gradable task with a competent tutor, that chain is good news, and the authors are right to report it.
Now move the same chain to a task where the tutor is sometimes wrong and nobody finds out immediately. Speed and fluency still drive engagement. Engagement still drives the student to go along with what’s on the screen. The chain doesn’t care whether the content is right.
Fluency and misplaced trust
Two weeks earlier, Maggie Liao and S. Shyam Sundar published a study in the Journal of Computer-Mediated Communication that shows the other end of it. They told 477 people they were testing a new AI health assistant, and had the assistant give them wrong information. One example was that ginger can beat chemotherapy against cancer. The more conversational the assistant, the more credible people found it, even when the claim was absurd.
They also tested fixes. A small warning icon did little on its own. Trust only dropped for people who actually clicked through and checked the claim against evidence. Given the option, nearly half did.
So the property StudentBench measures as helpful, a responsive, conversational tutor that keeps the student typing, is the same property that suppresses scepticism when the tutor is wrong. And passive warnings don’t fix it. Only verification does, and verification is effort.
Effort as the missing variable
That word is doing a lot of work this month. On 27 September the Irish Times carried a column by the Financial Times’s Sarah O’Connor on OECD PISA data showing that “hasty readers”, students giving fast but wrong answers, rose from 6.6% to 11.4% of test-takers between 2018 and 2025. The OECD’s Andreas Schleicher told her it was one of the big contributors to falling outcomes.
Cost per percentage point is a good metric. It leaves out what the learning cost the learner: how much effort they put in, how much they checked, and whether the habit of checking survived the session. A tutor that produces the same gain while the student does less of their own verifying is not the same tutor, even if the score is identical.
That’s the human–AI interaction question in this, and it’s the one I care about. Not whether the machine can teach. What the learner stops doing once it does.
What a better study would measure
StudentBench is a platform, which is its real strength: other people can run studies on it. If I could add four measures, they would be these.
- Delayed retention. The same kind of item a week later, no tutor.
- Unaided transfer. A problem of a different shape, no tutor.
- Verification behaviour. Plant the occasional wrong step in the tutor’s explanation and count how often it’s caught.
- Felt effort. Ask. A standard workload scale takes two minutes.
The third is the one that matters most, and it’s cheap. If AI tutoring holds its gains while students keep catching the planted errors, that is a much stronger result than equivalence on a score. If the catch rate falls as the tutor gets faster and friendlier, the headline number was never the whole story.
Sources
All checked on 28 September 2026.
- Northcutt, Hasmani, Feng, Khangi, Plesner and Mueller, StudentBench: AI and human tutoring yield equivalent GRE learning gains, arXiv 2609.28470, 23 September 2026; platform at studentbench.org
- Liao and Sundar, “Chat but verify”, Journal of Computer-Mediated Communication, doi:10.1093/jcmc/zmag012; summary: Penn State, 14 September 2026
- Sarah O’Connor, The descent of man: are we developing a distaste for effort?, The Irish Times, 27 September 2026
Todd McCaffrey is a New York Times bestselling author and holds an MSc in Cyberpsychology from ATU Letterkenny. He builds and writes about AI at foxxelabs.ie. This piece was drafted with Claude from a weekly news brief and checked against the sources above.