Bright Learning.

Guiding education through the AI revolution, for the people who run, teach, and care about schools.

Opinion

Your District Should Not Buy an AI Writing Detector

AI writing detectors fail asymmetrically and at scale, so districts should redesign assessment instead of purchasing a tool that manufactures wrongful accusations.

Every few months, a vendor arrives at a district office with a demo that promises to end the AI cheating problem. You paste in a student essay, and a confidence score tells you whether a machine wrote it. This is an opinion piece, and I will state my position plainly: your district should not buy that product. Bright Learning builds learning tools rather than detection tools, so we have a stake in this argument, and I would rather you weigh the evidence than take our word for it.

Start with the strongest case for buying one, because it is real. Teachers are facing a genuine surge in AI-assisted work that no gradebook was designed to catch, and asking them to police it by instinct is neither fair nor sustainable. A detector, the argument goes, does not have to be perfect to be useful. Treated as one signal among many rather than as a verdict, it gives an overwhelmed teacher a place to begin a conversation, and abandoning detection entirely can look like surrendering academic integrity at the exact moment it is under the most pressure.

That case deserves a serious answer, and here it is. A detector's output is not like a spell checker's suggestion. Once a score enters a disciplinary workflow, it stops being information and becomes an accusation, one the student must then disprove. And the accusation is not distributed evenly. When Stanford researchers tested seven popular detectors, more than half of the essays written by non-native English speakers on the TOEFL exam were flagged as AI-generated, and one detector flagged nearly all of them, while more than ninety percent of essays by U.S. eighth graders were correctly judged to be human (as ScienceDaily reported on the study published in Patterns). A district that adopts these tools is not buying a neutral integrity check. It is buying an instrument whose errors land hardest on the multilingual students who can least afford to be presumed guilty.

The scale of the problem is a matter of arithmetic, not opinion. When Vanderbilt University explained why it was disabling Turnitin's AI detector, it did the math out loud: the university had submitted about 75,000 papers to Turnitin in a single year, so even at the vendor's own claimed one percent false-positive rate, roughly 750 student papers could have been wrongly labeled as AI-written (per Vanderbilt's August 2023 statement). A district that buys a detector at that volume is not buying accuracy. It is buying a known error rate and applying it to its own children, and it will not be able to keep those errors quarantined as mere signals once they are sitting in a teacher's queue.

Then there is the question of who has already walked away. OpenAI, the company that built the model most of this anxiety is about, released its own AI text classifier and then shut it down in July 2023, citing a low rate of accuracy (as TechCrunch reported). Vanderbilt, one of Turnitin's own customers, turned the feature off and said plainly that it did not believe the software was effective, in part because the vendor would not explain how the tool reached its conclusions. When both the maker of the technology and its sophisticated heavy users have concluded the tool does not work, a school board approving a purchase order for it is overruling the people with the most information.

None of this means the cheating problem is imaginary or that teachers should be left alone with it. It means the money and the political will a district would spend on detection are better spent on assessment that is harder to fake: in-class writing, oral defenses of work, drafts that show their revision history, assignments built around a student's own context. Those approaches are slower and less tidy than a confidence score, and they ask more of adults than of an algorithm. But they do not manufacture accusations, and they do not aim their mistakes at the students with the least power to fight back. A district's job is not to catch the most cheaters at any cost. It is to run a system that is fair to the student in front of it, and a tool that fails this predictably and this unequally has no place in that system. Do not buy it.

Take it further