Bright Learning.

Guiding education through the AI revolution, for the people who run, teach, and care about schools.

Explainers

What a Rigorous Trial Actually Found About Khanmigo, the AI Tutor

An evidence-led look at an independent randomized trial of Khan Academy's AI tutor Khanmigo, separating what the study reports, what the company claims, and where those accounts diverge so parents can weigh AI-tutor hype against the reality that most students decline Socratic help.

Photograph by Caleb Oquendo, via Pexels.
Photograph by Caleb Oquendo, via Pexels.

The claim versus the question

Parents and educators hear a steady drumbeat that AI tutors will transform outcomes for struggling students. A recent independent study offers an unusually sober test of that promise. It examined Khanmigo, the generative-AI tutoring layer built on top of Khan Academy's practice software, inside real schools over two years. The headline finding is uncomfortable for the hype: the AI layer added little, in large part because students simply did not use the help it offered.

This explainer synthesizes three sources of differing provenance. The first is an independent randomized controlled trial circulated as a National Bureau of Economic Research working paper, titled "One Click Away: AI Tutoring with Khanmigo in a Two-Year School Experiment" Source. The second is an interpretive blog post by Khan Academy's founder, Sal Khan, explaining what he found compelling in the trial Source. The third is independent journalism from The Hechinger Report, with original reporting by education reporter Jill Barshay, under the headline "Students didn't get answers from Khanmigo. They didn't want its questions, either" SourceSource.

Reading these together lets us separate what the independent paper reports, what the company claims about effect sizes and its product changes, and where the two accounts pull apart.

What the independent trial reports

The NBER working paper describes a two-year school experiment testing Khanmigo-enabled tutoring, framed by its own title around the idea that help was always "one click away" yet frequently not taken Source. The core behavioral finding, quoted directly in the Hechinger Report's coverage, is that students largely declined to engage with the Socratic, question-based help the AI offered. As the journalism frames it, students did not get answers from Khanmigo, and they did not want its guiding questions either Source.

This matters because the entire theory behind an AI tutor rests on students choosing to lean on it. Help-seeking is a behavior, not a feature. If the tool is available but unused, the AI layer cannot produce the gains attributed to it, no matter how well it performs when engaged.

The Hechinger Report characterizes the measured learning gains as modest and no greater than what Khan Academy's underlying practice software had produced in prior work Source. In other words, the independent journalistic read is that the generative-AI component did not clearly outperform the practice platform it was layered onto.

What the company claims

Sal Khan's blog post offers a more encouraging interpretation of the same trial. In his telling, the results are compelling, particularly for below-grade-level students using the math intervention Source. His post is also the only opened source that reports specific effect sizes.

According to that company blog, the trial produced gains in the range of roughly 0.06 to 0.14 standard deviations, with a figure around 0.08 standard deviations cited among them Source. These are small effects by the conventions of education research. For comparison, Khan's post points to a separate Khan Academy field experiment in India, also circulated through NBER, which reported a much larger gain of about 0.45 standard deviations under in-school supervised ed-tech support SourceSource.

The company blog also describes involvement by Charles River Associates in the analysis, and frames the findings as motivation for a product redesign intended to increase student engagement with the AI tutor Source.

Where the accounts diverge

The two readings do not contradict each other on facts so much as on emphasis.

The independent paper and the independent journalism foreground the behavioral failure: help was available, students declined it, and the AI layer therefore added little beyond the practice software SourceSource. The company blog foregrounds the positive signal for below-grade-level students and the specific, if small, effect sizes, while positioning the disappointing engagement as a design problem to be fixed Source.

A careful reader should notice what each source can and cannot establish. The specific effect-size numbers (roughly 0.06, 0.08, and 0.14 standard deviations, and the 0.45 standard deviation India comparison) were verifiable only in Sal Khan's company blog, not in the opened text of the primary paper Source. The Hechinger Report corroborates the findings qualitatively, calling them modest and no greater than prior Khan Academy results, and quotes the paper's core line directly, but it does not state those numeric effect sizes Source. Charles River Associates' involvement appears only in the vendor blog Source.

This is the crux for parents weighing the marketing. The most favorable framing of the numbers comes from the party with the strongest interest in a favorable framing. The independent sources confirm the direction and the smallness of the effect, and they add the crucial behavioral caveat that the AI's distinctive contribution went largely unused.

What this means for parents and educators

Two honest takeaways emerge.

First, the underlying practice software appears to produce only modest gains for below-grade-level students, and the generative-AI tutor layered on top did not clearly improve on that in this trial SourceSource. An AI tutor is not, on this evidence, a transformative intervention delivered by default.

Second, and more usefully, context and supervision may matter more than the AI itself. The much larger gain in the separate India experiment came under in-school supervised support, where adults structured the use of the tool Source. The pattern across both studies points to the same lesson: technology that is merely available tends to be underused, while technology embedded in a supervised routine is more likely to help. Help-seeking is a choice most students decline when left to make it alone.

Take it further