When AI tutors arrived in classrooms, the claims arrived with them. Some promised personalized instruction that would close learning gaps at scale. Others warned that chatbots would hand students answers and hollow out learning. A two-year randomized controlled trial of Khan Academy's Khanmigo offers a chance to replace both the hype and the doom with evidence, and the picture it paints is more modest and more interesting than either story.
This explainer synthesizes what the trial reported. Bright Learning did not conduct this research and could not open the primary working paper directly (see Limitations). The findings below are drawn from the paper's listing and from secondary reporting and a party-interested blog, and they describe a 2023 version of Khanmigo that has since been redesigned.
What was tested
The study, titled "One Click Away: AI Tutoring with Khanmigo in a Two-Year School Experiment," is a National Bureau of Economic Research working paper Source. According to secondary reporting, it ran across 18 schools in Hamilton County, Tennessee, following students over two years Source.
The design matters because it does something many ed-tech studies do not: it separates two different things that often get bundled together under the phrase "AI tutor." One is the Khan Academy practice platform itself, the sequenced exercises and videos students have used for years. The other is Khanmigo, the conversational AI layer that sits on top of it and is designed to coach students through problems rather than hand over answers Source.
Separating these two lets the trial ask a sharper question. Not "does Khan Academy help?" but "does adding an AI tutor to Khan Academy help more than Khan Academy alone?"
What the trial found
The short version: the practice software moved the needle modestly, and the AI layer did not add measurably to it.
Secondary sources report a combined effect of roughly 0.06 standard deviations, rising to about 0.08 standard deviations in the second year, with a larger effect of about 0.14 standard deviations for students who were below grade level and stayed in the program through a sustained second year Source. These numbers were confirmed through secondary reporting rather than the primary document.
To calibrate: an effect around 0.06 to 0.08 standard deviations is small but not nothing, and it sits broadly in line with what prior evidence has found for well-implemented education technology. It is the kind of gain that is real and worth having, but it is not the transformational leap some AI marketing has implied.
The more pointed finding is about the AI layer. According to the reporting, adding Khanmigo on top of the practice platform produced no measurable additional benefit Source. The gains that appeared came from the practice software, not from the conversational AI coaching.
Why the AI didn't add anything: the help-seeking bottleneck
This is where the trial becomes genuinely informative rather than just disappointing. The authors identify a behavioral mechanism, and it is not the one most debates about AI tutors focus on.
The worry people usually voice is that an AI tutor will simply give students answers, short-circuiting the effort that produces learning. Khanmigo was designed specifically to avoid this. It withholds answers and instead asks guiding questions, the way a good human tutor does.
What the trial found is that once the AI withheld answers and offered questions instead, students largely declined the help Source. The headline of the secondary reporting captures it directly: students did not get answers from Khanmigo, and they did not want its questions either Source.
That is the bottleneck. The constraint was not the quality of the AI's coaching. It was student help-seeking behavior. Help that students do not accept cannot help them, no matter how well designed it is. A tutor who is one click away is still only useful if the student clicks.
This reframes the whole question. The problem of AI tutoring in this trial was less about the machine's pedagogy and more about the human decision of whether to engage with support that requires effort rather than supplying shortcuts.
How the people involved read it
It is worth noting where interpretations diverge. Khan Academy, in a company blog post, described finding the trial compelling and emphasized the gains for below-grade-level students who stayed in the program Source. That reading is not wrong, but it is a party-interested source, and it understandably foregrounds the positive subgroup result over the null finding for the AI layer. A reader synthesizing the evidence should hold both facts at once: a modest real gain from the practice platform, and no measurable incremental gain from the AI tutor as tested.
There is also related work pointing in a different direction about what makes ed-tech work. A separate NBER working paper examines a Khan Academy field experiment in India and, per its title, describes how in-school supervised support produced large learning gains Source. We were only able to partially verify that listing and could not confirm its abstract, authors, or date, so treat it cautiously. But the contrast is suggestive: the ingredient that may matter most is not the sophistication of the tool but the structure around it, including supervision that pushes students to engage. If the bottleneck is help-seeking, then the fix may be human and organizational rather than algorithmic.
The honest takeaway: hope and caution together
For readers trying to calibrate between sweeping promises and sweeping fears, the evidence supports neither extreme.
The hope: structured practice software produced modest, real gains, larger for the struggling students who stuck with it over a sustained period Source. That is a defensible reason to keep investing in these tools.
The caution: the AI tutor layer, as tested in 2023, did not add measurable value, and the reason was that students declined help that asked them to think rather than handing them answers Source. Buying an AI tutor does not automatically produce learning. Getting students to use it as designed is the harder, more human problem.
And the biggest caution of all is about what this evidence is. It is a working paper that has not been peer-reviewed. It tested a specific 2023 version of Khanmigo that has since been redesigned Source. The product students would encounter today is not the product the trial measured. That does not erase the central lesson about help-seeking, which is likely durable, but it does mean the specific null result should not be read as a permanent verdict on the current tool.
