Bright Learning.

Guiding education through the AI revolution, for the people who run, teach, and care about schools.

Explainers

Did an AI Tutor Actually Help Kids Learn? What the First Independent Khanmigo Trial Found

The first independent randomized trial of Khanmigo, an answer-withholding AI tutor, found Khan Academy helped modestly in math while the AI chatbot added nothing detectable over existing tools. Here is a calibrated reading of what the evidence does and does not show.

Photograph by Rosa y Dani, via Openverse.
Photograph by Rosa y Dani, via Openverse.

Why this one study matters

For years the promise has been constant: a personal AI tutor for every child, patient and always available. Khanmigo, built by Khan Academy, is one of the most prominent attempts to deliver it. It is deliberately designed not to hand out answers. Instead it asks guiding questions, nudging a student toward the solution the way a good human tutor might.

The question that matters is simple to state and hard to answer: does it help kids learn more? A new independent randomized controlled trial by economists Philip Oreopoulos and Hugh Low, circulated as a National Bureau of Economic Research working paper, is the first genuinely independent, rigorous test of this kind of purpose-built tutor Source. This explainer synthesizes that paper's public metadata, the general-audience reporting that broke it by Jill Barshay at The Hechinger Report (including her original interview with Sal Khan), and Khan Academy's own blog response SourceSource.

What an RCT is, in plain terms

A randomized controlled trial, or RCT, is the closest education research gets to a fair coin toss. You take a group of students and randomly assign some to use the new tool (the treatment group) and others to carry on without it (the control group). Because the assignment is random, the two groups should look alike on average in every way except the one thing you are testing. If they then end up different on a math test, the tool is the most plausible cause. RCTs are considered the strongest single design for establishing whether something actually works, rather than whether it merely looks promising.

Why the control group changes everything

Here is the detail that reshapes how you should read the headline. In this trial, the comparison was not against nothing. It was against what researchers call "business as usual." The control students were not sitting idle; they kept using the educational tools and instruction they would normally have access to Source.

That matters enormously. If you test a new tool against students who received no support, almost anything can look effective. But if you test it against students who already have decent tools, you are asking a much harder and more honest question: does this add anything on top of what already exists?

The three numbers, and why they disagree

Effect sizes in education are usually reported in standard deviations (SD), a way of measuring how big a difference is relative to the natural spread of scores. Roughly speaking, 0.1 SD is a small effect and 0.4 SD or above is large.

There are three readings circulating from this trial, and they are not contradictory once you understand what each measures:

  • Intent-to-treat, roughly 0.06 to 0.08 SD. This is the effect across everyone assigned to the treatment group, whether or not they actually used the tool much. It is the most conservative and arguably the most honest figure, because it reflects what happens when you roll a tool out to a real population where many people do not engage Source.
  • Sustained participants, roughly 0.14 SD. This is the effect among students who actually kept using the platform. It is larger, which is unsurprising, but it is also harder to trust, because students who stick with a tool may differ from those who drop off in ways that themselves boost scores Source.
  • The vendor's compelling read. Sal Khan's blog frames the findings as encouraging evidence for Khan Academy in math intervention Source. This is explicitly an advocacy reading and should be weighed as such.

The defensible claim

Put the pieces together and one statement survives scrutiny: Khan Academy helped modestly, but Khanmigo, the AI chatbot layered on top, added nothing detectable over the existing tools. The measurable gains appear to come from the broader Khan Academy platform used as math intervention, not from the AI tutor feature specifically Source.

This is a narrower and more useful finding than either "AI tutors work" or "AI tutors fail." It says the existing, non-AI scaffolding did some good, and the much-hyped conversational AI component did not demonstrably improve on it in this trial.

The human finding that is easy to miss

The most striking result is not about the technology at all. It is about students. Khanmigo was built to withhold answers and ask questions instead. The trial's telling detail is captured in Barshay's headline: students did not get answers from Khanmigo, and they did not want its questions either Source.

Help-seeking is a choice, and most students declined it. A tutor that is always available only helps the students who choose to engage with it, and many simply did not. This points at a motivation problem that no amount of clever answer-withholding design can solve on its own. The bottleneck may be human willingness, not AI capability.

What the vendor says

Khan Academy did not dispute that the study happened or hide from it. Sal Khan engaged publicly, both in his blog and in an interview with Barshay, and framed the trial as containing genuinely compelling signals for math intervention SourceSource. Khan also indicated that the product tested is not the product that exists today, because Khanmigo has since been redesigned Source. That is a fair point and an important limitation, but note that the redesign claims come from the CEO, not from any independent test of the new version.

Honest uncertainty

This is one study, of one tool, with one population, testing a version of a product that has already changed. A single RCT, however rigorous, is a data point and not a verdict. The sustained-participant figure invites optimism but rests on a self-selected group. And the people most invested in the tool are the ones telling us it has improved since the trial.

What we can say with reasonable confidence is modest and worth saying plainly: when tested fairly against tools students already had, the AI tutor did not demonstrably move the needle, and most students chose not to use it as designed.

Take it further