Why this matters
Families and educators are being told that AI tutors will transform K-12 learning. A large, two-year randomized controlled trial of Khan Academy's Socratic AI tutor, Khanmigo, gives us a rare chance to separate a genuine research result from the marketing narrative. The short version: the underlying practice platform produced modest gains, but the AI chat layer added little measurable value, and the most likely reason is that students simply did not use it much.
This explainer states the numbers exactly as the paper's own published abstract reports them, then explains the engagement data that helps make sense of the result, and finally frames the disagreement fairly.
What the study measured
The study is titled "One Click Away: AI Tutoring with Khanmigo in a Two-Year School Experiment" and is published as an NBER working paper Source. It is also indexed in the RePEc/IDEAS research database Source.
A randomized controlled trial (RCT) is the strongest common design for asking whether an intervention causes a change, because students are assigned to conditions by chance rather than by choice. This reduces the risk that pre-existing differences between groups explain the outcome.
The effect sizes, stated precisely
Effect sizes here are reported in standard deviations (SD). A useful rule of thumb: in education research a full-year effect around 0.1 SD is small but real, and larger figures are uncommon.
According to the paper's own abstract, the intervention produced gains of roughly 0.06 to 0.08 SD over a school year Source. This is the intent-to-treat estimate, meaning it averages across everyone assigned to the treatment group regardless of how much they actually used the product. Intent-to-treat is the conservative, policy-relevant number because it reflects what happens when you roll something out to real classrooms where uptake varies.
The abstract also reports a larger figure, described as the implied effect of a full year of active participation, about 0.14 SD Source. This is a participation-adjusted estimate: it models what the effect would look like for a student who engaged fully across the year. It is not the same thing as the average effect of being offered the program. Both numbers are legitimate, but they answer different questions, and conflating them overstates what the average student experienced.
Critically, the paper's abstract frames these gains as resembling those from Khan Academy practice without AI assistance Source. In other words, the measured benefit looks like the benefit of the practice platform itself, not of the Socratic AI layer added on top.
Why the AI layer added little: the engagement data
The paper's abstract reports engagement statistics that help explain the null AI result. Nearly all students tried the tool at least once: 96 percent used Khanmigo Source. But usage was thin after that first try. The median student messaged Khanmigo on only about a third of their practice days, and engaged it in only about 17 percent of exercise sessions in which they made an error, often sending bare answers rather than working through the tutor's questions Source.
That last point matters for a Socratic tutor. Khanmigo is designed to withhold direct answers and instead ask guiding questions. If students respond with bare answers and stop, the intended learning mechanism never really activates. Independent reporting by The Hechinger Report captured this dynamic in its headline framing: students did not get answers from Khanmigo, and they did not want its questions either Source.
Bright Learning did not read the paper's full methods, tables, or robustness checks; the numbers above are drawn from the published abstract, and the interpretation of low engagement is corroborated by Hechinger's independent reporting Source.
How the developer responded
Khan Academy founder Sal Khan wrote a blog post describing what he found compelling in the trial Source. He emphasized the participation-adjusted framing, pointing to the larger effect for students who engaged actively Source. That is an understandable emphasis from an interested party, but readers should note that the abstract's neutral wording labels this the "implied effect of a full year of active participation," not a claim about the typical student's outcome Source.
Khan's post also describes contextual details, including that more than a quarter of participants were special-education students and that the business-as-usual comparison condition included other practice software such as IXL, Zearn, Waggle, and DeltaMath in many of the schools Source. Bright Learning could not verify these details against the paper's abstract; they come from Khan's blog and secondhand reporting, so treat them as context rather than confirmed findings.
An important open question
Even if a future version of Khanmigo drove much higher engagement, more messaging is not the same as more learning. A tool can be sticky without being effective. The trial's central lesson is not that AI tutoring cannot work, but that in this rigorous test, being one click away from an AI tutor was not enough: most students did not take the click, and the ones who did often did not engage in the way the tutor was designed for.
A note on a different result you may hear cited
A separate Khan Academy field experiment in India reported much larger gains, in the range of 0.45 SD, from in-school supervised ed-tech support Source. That specific figure reaches this explainer secondhand rather than from the paper's verified full text, so treat it cautiously. But the contrast is instructive: the India study centered on supervised, structured, in-school use, whereas the strongest lesson of the Khanmigo trial is that unsupervised, optional AI engagement tends to be low. Implementation and supervision may matter more than the AI feature itself.
The bottom line for families and educators
- The practice platform produced a small but real gain (0.06 to 0.08 SD over a school year) Source.
- The AI chat layer did not add measurable value on average, and the gains resembled those from Khan Academy practice without AI Source.
- The most likely explanation is low engagement with the Socratic tutor, not that the underlying math practice failed SourceSource.
- Be cautious with the 0.14 SD figure: it describes a hypothetical fully engaged student, not the average one Source.
