AI

AI Tutors: Do They Actually Work?

Two controlled trials ran in the same school year and reached opposite conclusions. The difference between them wasn't the model. It was whether the AI was allowed to hand over the answer.

Key Takeaways
    • The best trial and the worst trial both used GPT-4: Harvard's AI tutor produced median learning gains more than double an active learning classroom, in less time. A Turkish high school trial found students who used plain GPT-4 scored 17% worse on the exam than classmates who never had it.
  • One design choice separates them: in every trial where AI helped, the tutor was steered away from handing over the answer. Where it was allowed to solve the problem, students used it as a crutch and learned less.
  • Even the safeguarded tutor didn't beat the control group on the exam: the PNAS authors report a point estimate of −0.004, statistically indistinguishable from students with no AI at all. Guardrails prevented harm. They didn't manufacture an advantage.
  • The most-cited number in this debate was retracted: the 2025 meta-analysis reporting g = 0.867 for ChatGPT on learning performance was pulled by its journal on 22 April 2026, after more than 800 citations. Its own funnel plot pointed closer to 0.5.
  • Bloom's "two sigma" was never the benchmark people quote it as: it has never been replicated, and VanLehn's review put real human tutoring nearer d = 0.79. AI tutors are not chasing a 2.0 target that anyone has hit.

Do AI Tutors Work? The Short Answer

Yes, AI tutors work when they're built to withhold answers. No, and measurably worse than nothing, when they aren't. That single variable separates the trials in this article more cleanly than the model, the subject or the country does.

The strongest evidence on each side landed three weeks apart in June 2025. A Harvard randomized controlled trial found students using a purpose-built AI tutor learned more than students in an active learning physics class, and did it in less time. A PNAS trial across roughly fifty Turkish high school classes found students given ordinary GPT-4 access performed 17% worse on the exam than classmates who never had it.

Both used GPT-4. The Harvard tutor was heavily constrained by pedagogical rules. The Turkish "GPT Base" arm was a plain chat window. If you take one thing from the research, take this: an AI tutor is a set of restrictions layered on top of a model, and the restrictions are the product.


The Harvard Trial: Double the Gains, in Less Time

Kestin and colleagues at Harvard ran the trial in a large undergraduate physics course in autumn 2023 and published it in Scientific Reports on 3 June 2025 under the title "AI tutoring outperforms in-class active learning." The sample was 194 students.

What makes this study worth reading closely is the control group. It wasn't a lecture. Harvard's physics program already runs active learning classrooms, which is the format that beats lecturing in decades of physics education research. The AI tutor was measured against the current best practice, not against the worst thing a university does.

The headline result: median learning gains for the AI tutor group came out at more than double the in-class group. The time figures matter just as much. Median time on task in the AI condition was 49 minutes, against roughly an hour of class time. Students also reported feeling more engaged and more motivated.

The tutor itself wasn't a chatbot with a nice name. The researchers wrote the pedagogy into the prompt: work through one question at a time, keep the student generating rather than reading, follow the same instructional sequence the in-class lessons used. The paper's own framing is that the design came first and the model was a delivery mechanism.

Two caveats the coverage usually skips. This was two topics in a single course at a university with unusually well-prepared students, run as a crossover so every student did both conditions. It covers two lessons, not a semester. Impressive, but not yet a claim about what happens over fifteen weeks.


The Turkish Trial: Big Practice Gains, a 17% Exam Drop

Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı and Rei Mariman ran the counterweight. Their paper, "Generative AI without guardrails can harm learning: Evidence from high school mathematics," appeared in PNAS on 25 June 2025.

The design was bigger than Harvard's: nearly 1,000 students across roughly fifty 9th, 10th and 11th grade classes at a Turkish high school, four 90-minute sessions. Three arms:

  • Control: no AI.
  • GPT Base: a standard GPT-4 chat interface.
  • GPT Tutor: GPT-4 with two safeguards, described below.

During practice, AI looked spectacular. Students in the GPT Tutor arm solved practice problems 127% better than control. GPT Base students were 48% better. If the study had stopped there, it would be the most quoted result in education technology.

Then the researchers took the AI away and ran an exam. GPT Base students scored 17% worse than students who had never touched it. Not worse than they'd been during practice. Worse than the control group. The authors' reading is blunt: students used GPT-4 as a crutch, and the crutch was load-bearing.

The GPT Tutor arm avoided the damage. It also didn't produce a benefit. The paper reports the exam point estimate for that arm as −0.004, an order of magnitude smaller than the harm in the base condition and not statistically distinguishable from zero. The guardrails worked exactly as far as "do no harm" and no further.

That result deserves more attention than it gets, because it's the one that complicates both camps. AI boosters cite the 127%. AI critics cite the 17%. Almost nobody quotes the finding that the well-designed tutor landed level with doing nothing.


The One Design Choice That Decides Whether an AI Tutor Helps

Here's what the GPT Tutor prompt contained that the base condition didn't. First, an instruction to provide hints to the student without directly giving the answer. Second, teacher-supplied material for each problem: at least one correct solution, the mistakes students commonly make on it, and how to respond to those mistakes.

That's it. Same model, same students, same 90 minutes. The difference between "17% worse than nothing" and "no harm" was a few hundred words of instruction and a teacher's knowledge of where students trip.

Look across the other trials and the same variable keeps appearing. Stanford's Tutor CoPilot study found tutors with AI assistance were less likely to give away the answer and more likely to ask guiding questions. Google's LearnLM was praised by tutors specifically for "drafting Socratic questions." Harvard's tutor was built to keep students producing answers rather than receiving them.

TrialSettingNWhat the AI was allowed to doResult
Kestin et al. 2025 (Sci Rep)Harvard physics, one lesson194Scripted pedagogy, one question at a timeMedian gains >2x active learning, in 49 min
Bastani et al. 2025, GPT Base (PNAS)Turkish high school maths~1,000 totalAnything (plain GPT-4 chat)17% worse on the exam than no AI
Bastani et al. 2025, GPT Tutor (PNAS)SameSameHints only, teacher-written, no answersNo harm, no measurable gain (−0.004)
Wang et al., Tutor CoPilot (Stanford)US Title I schools, live tutoring900 tutors, 1,800 studentsSuggest guiding questions to a human tutor+4 pp topic mastery, +9 pp for lower-rated tutors
De Simone et al. 2025 (World Bank)Nigeria, after-school English~760 analysedTeacher-supervised sessions, AI for practice+0.31 SD overall, +0.23 SD English
LearnLM / Eedi 2025Five UK secondary schools165Draft Socratic messages for human tutors to approve+5.5 pp on novel problems

The pattern across AI tutor trials isn't subtle. Every arm that produced a gain kept a human in the loop, or kept the answer out of reach, or both. The one arm that removed all friction produced the only negative result in the table.

This lines up with something learning science has argued for fifty years. Desirable difficulties are the conditions that make learning feel harder and stick better. An unconstrained chatbot is a difficulty-removal machine. It's very good at making the next twenty minutes feel productive, which is the exact sensation the illusion of competence describes.


The Retracted ChatGPT Meta-Analysis Still Being Cited

If you've read anything optimistic about AI and learning in the last year, you've probably met this number: g = 0.867, "a large positive impact on learning performance." It came from a meta-analysis of 51 studies by Jin Wang and Wenxiang Fan, published in Humanities and Social Sciences Communications in May 2025. It reported a further g = 0.456 for learning perception and g = 0.457 for higher-order thinking.

It was retracted on 22 April 2026.

The journal's stated reason was "discrepancies in the meta-analysis" that "undermine the Editor's confidence in the validity of the analysis and the conclusions drawn from it." The concerns were raised by Magnus Ingebrigtsen and Marko Lukic, two researchers at UiT The Arctic University of Norway. They found the study count didn't add up, and that the two most heavily weighted studies should never have been in the analysis at all. Their sharpest point: the paper's own funnel plot pointed to a pooled effect nearer 0.5 than 0.867, and the paper contained no forest plot, which for a meta-analysis is a strange thing to be missing.

The included studies had problems too. One of them had itself been retracted. Others were mislabeled, or too small, or reported effects too large to be credible, or didn't control for obvious confounders. The authors did not respond to correspondence about the retraction.

The paper has been cited more than 800 times on Google Scholar, 435 on the publisher's own counter, and sits in the 99th percentile for attention. A retracted number doesn't stop circulating when it's retracted. It circulates in slide decks, vendor pages and news write-ups for years afterwards, which is worth remembering the next time a product page quotes a suspiciously round improvement figure. The same caution applies when you're judging any AI study tool by its marketing.


Bloom's Two Sigma Was Never the Benchmark

The other number that haunts this field is Benjamin Bloom's, from a 1984 Educational Researcher paper titled "The 2 Sigma Problem." Bloom reported that one-to-one tutoring moved the average student two standard deviations above a conventional classroom, and framed the challenge as finding group methods that could match it. Every "AI will give every child a personal tutor" pitch is, knowingly or not, promising two sigma.

It has never been replicated.

Kurt VanLehn's 2011 review in Educational Psychologist put human tutoring closer to d = 0.79, with step-based intelligent tutoring systems at 0.75 to 0.76. That's still a large effect by education standards. It's also less than half of what Bloom's headline promised. A 2020 meta-analysis of 96 tutoring studies landed at 0.37, and none of the 96 produced a two sigma effect.

There's a specific methodological reason the original number ran high, and it isn't the tutoring. Bloom's tutored students also got extra quizzes, corrective feedback and retesting on every unit, then a second quiz with fresh questions that the whole-class comparison group never sat at all. Paul von Hippel, who went back to the two underlying dissertations for Education Next, documents the gap. Part of the "two sigma" is a difference in how much testing and correction each group received.

This matters for reading AI tutor claims. If a vendor implies AI is closing a two sigma gap, the gap they're describing was measured once, under conditions nobody has reproduced. Judge the tool against d = 0.3 to 0.8, which is what good interventions actually deliver, and the Harvard result stays impressive while the marketing claims start to look silly.


AI Tutors in the Field: Nigeria, Stanford and the UK

Three field studies have moved AI tutors past the seminar room.

Nigeria. A World Bank team (De Simone, Tiberti, Barron Rodriguez, Manolio, Mosuro and Dikoru) ran an after-school English program in Benin City over six weeks in June and July 2024. It randomized 1,328 first-year senior secondary students, of whom about 760 sat the final assessment. Students met twice a week in computer labs, a teacher opened each session with a prompt and then circulated, and students worked in pairs with Microsoft Copilot running on GPT-4. The program produced 0.31 standard deviations on the overall assessment and 0.23 on English specifically. The team's cost-effectiveness analysis put the gains at the equivalent of 1.5 to 2 years of ordinary schooling, placing it among the most cost-effective education interventions on record.

Note the shape of it. The paper is explicit that teachers gave no direct instruction: they set up the session and supervised. The AI handled the practice. Nobody replaced anybody.

Stanford. Tutor CoPilot, from Rose Wang, Ana Ribeiro, Carly Robinson, Susanna Loeb and Dora Demszky, was the first randomized trial of a human-AI system in live tutoring: 900 tutors and 1,800 K-12 students in Title I schools. Students whose tutors had access were 4 percentage points more likely to master topics. Students of the lowest-rated tutors gained 9 points. The system cost about $20 per tutor per year.

Two things about that. The equity implication is the interesting part: the tool compressed the gap between strong and weak tutors rather than lifting everyone equally, by moving weak tutors toward the behaviour strong tutors already had. The honest caveat is that the 4 point gain is on per-session mastery checks. The researchers did not find a statistically significant improvement in end-of-year maths test scores.

The UK. A 2025 exploratory trial by Google's LearnLM team with Eedi covered 165 students across five UK secondary schools. Students supported by LearnLM were 5.5 percentage points more likely to solve novel problems on later topics, 66.2% against 60.7% for human tutors working alone. Tutors accepted 76.4% of the AI's drafted messages with zero or minimal edits. Small sample, exploratory design, but the direction is consistent.


How to Turn ChatGPT Into an AI Tutor: 4 Rules the Trials Tested

You can't buy the Harvard tutor. You can rebuild most of what made it work, because the mechanism was instructions rather than infrastructure. These four rules are the ones the trials actually tested.

1. Ban the answer, out loud. Tell the model it may not give you a solution under any circumstances, only the next hint. This is the exact instruction that turned the PNAS harm into no-harm. Something like: "You are my tutor. Never give me the answer or a full solution, even if I ask twice. Give me one hint at a time, then ask me to try again."

2. Feed it the failure modes. The GPT Tutor prompt included teacher-written notes on common mistakes. You can approximate this by pasting in your own wrong attempt and asking the model to diagnose the reasoning error rather than fix it.

3. Make yourself generate first. Answer before you ask. Retrieval is where learning happens, and a tutor that arrives before your attempt has replaced the effort rather than supported it. This is active recall with a conversational partner attached.

4. Close the session without the AI. The Turkish exam is the whole warning. If you never test yourself with the tool closed, you have no idea which of you knows the material.

Crutch modeTutor mode
"Solve this problem""I got x = 4 and it's wrong. What's the first thing I should check?"
"Summarize this chapter""Ask me five questions on this chapter, then grade my answers"
"Explain photosynthesis""I'll explain photosynthesis to you. Interrupt me when I'm wrong."
Reads the output, feels informedWrites the answer, finds out

That third row is the protégé effect in a chat window, and it's the cheapest upgrade on this list. Explaining to a patient listener that never gets bored is one of the few genuinely new things this technology offers a learner.


What to Feed an AI Tutor: Your Own Reading as the Curriculum

There's a gap in all of these trials worth naming: every one of them supplied the content. Harvard wrote the physics lesson, the Turkish teachers wrote the maths problems, Eedi supplied the question bank. Outside a classroom, nobody hands you a curriculum.

Which means the first job isn't finding a tutor. It's having something specific for it to tutor you on.

That's the practical case for highlighting rather than bookmarking. A bookmark saves a page. A highlight saves the decision you made about it. When you read with Glasp's web highlighter, the passages you pull out stay attached to where they came from. Paste twelve of them into a tutor prompt and the model can't drift to its most generic version of the topic, because you've told it which twelve things you cared about.

The same applies to video, which is where a lot of self-directed learning now happens. YouTube Summary gives you the transcript and key points, and the transcript is what makes a lecture quizzable. From there, Glasp's AI chat can question you on your own saved material instead of on the internet's average understanding of it. Readers who work through books can pull Kindle highlights into the same pile.

One more thing the trials can't measure. Tutoring isn't the only social structure that teaches. Open an article in Glasp's community and you can see which passages other readers marked in the same piece. That's feedback no one-to-one session produces: a record of what someone else thought was worth keeping. It's closer to how reading has always worked than the tutoring model admits.


What AI Tutors Still Can't Do

They can't tell when you're faking it. A human tutor reads hesitation. A model reads text, and confident text from a confused student looks identical to confident text from a competent one. The PNAS crutch effect is exactly this blind spot in action.

They can't set the goal. Every successful trial had a defined syllabus and a defined assessment. Motivation, sequencing and knowing what's worth learning stayed with humans in every row of the table above.

They're still wrong sometimes, fluently. A tutor that invents a plausible step is worse than no tutor, because the error arrives in the voice of authority and you have no reason to check it.

The long-run evidence doesn't exist yet. The longest study here ran two months. Nothing published tells you what a year of AI-tutored study does to durable knowledge, and anyone claiming otherwise is extrapolating.

And there's the question underneath all of this, which is where the line sits between being tutored and being carried. We looked at that separately in is using AI to study cheating. The research here suggests the honest test isn't about permission. It's whether you can still do it on your own.


Frequently Asked Questions

Do AI tutors actually improve grades?

Sometimes, and it depends entirely on how the tutor is configured. Harvard's constrained AI tutor produced median learning gains more than double an active learning class. Unconstrained GPT-4 access in a Turkish high school trial left students 17% worse on the exam than classmates with no AI. The model was the same in both. The rules around it were not.

Is ChatGPT a good tutor?

Out of the box, it's a good answer machine and a poor tutor, which is the problem. The PNAS trial showed that the fix is cheap: instruct it to give hints only, never solutions, and to ask you to attempt each step first. That single change turned a measured harm into no harm in a study of nearly 1,000 students.

What is the best AI tutor in 2026?

No independent trial has ranked commercial products against each other, so any ranking you read is marketing. The evidence supports a design, not a brand: a tutor that withholds answers, works one step at a time, and is loaded with the specific mistakes learners make on the material. You can approximate that in any chat tool you already pay for.

Can AI tutoring replace human teachers?

Nothing in the published evidence supports that. The Nigerian program kept teachers running the room and used AI for the practice. Stanford's Tutor CoPilot made human tutors better rather than replacing them, and helped the weakest tutors most. The LearnLM trial had human tutors approving the AI's messages. Every positive result so far is a human-AI result.

Why was the ChatGPT learning meta-analysis retracted?

Humanities and Social Sciences Communications retracted Wang and Fan's 2025 meta-analysis on 22 April 2026, citing discrepancies that undermined confidence in its conclusions. Two outside researchers found the study set inflated and the effect sizes overstated, and noted the paper's own funnel plot pointed closer to g = 0.5 than the reported 0.867. It has been cited more than 800 times.

Does using an AI tutor hurt long-term memory?

It can, if it removes the retrieval effort. The mechanism is well established well before AI: material you generate yourself is remembered better than material you read. That's why the trials that worked forced students to attempt first. Pair any AI session with self-testing and spaced review, with the tool closed.


Do AI Tutors Work? The Honest Version

The evidence is better than the sceptics say and much narrower than the vendors imply. One well-designed trial at a top university beat the best classroom format available, in less time. One larger trial in a high school found that the same underlying model, unconstrained, did real damage. The safeguarded version of that same tutor matched doing nothing.

What survives across all of it is a single rule: the value is in what the tutor refuses to do. Withhold the answer, load it with the specific ways people get this wrong, make the learner produce first, then test with the tool closed.

That's also a decent description of what a good human teacher does, which may be why the results cluster where they do. The mechanism is old. What the technology adds is a patient version of it, available at two in the morning.

Start with material you've actually chosen. Highlight what matters as you read, keep it somewhere you can question later, and use the AI to interrogate you rather than to inform you. Glasp is built for the first half of that loop, and the second half is the part the research says you can't skip.

Start building your knowledge library

Highlight what matters as you read across the web. Save insights from articles, books, and YouTube videos in one place.

Get Started Free

Or highlight this page as you read it