Bright Learning.

Guiding education through the AI revolution, for the people who run, teach, and care about schools.

Opinion

Districts Should Stop Using AI-Writing Detectors to Discipline Students

AI-writing detectors accuse the innocent along patterns of language and disability and produce scores no one can cross-examine, so districts should retire them as disciplinary tools even though the cheating problem is real.

Picture the sequence, because the sequence is the problem. A student turns in an essay. A detector scores it thirty-one percent likely to be written by AI, and a teacher who has never before read a line of this student's work now reads the whole paper hunting for proof of a crime the software has already alleged. That sequence, and not any single wrong score, is the case against AI-writing detectors, and it is why districts should retire them as tools of discipline. I want to be clear that this is an argument and not a neutral audit: Bright Learning's position is that these tools do more harm than good in schools, and we hold that view knowing the cheating problem they were bought to solve is real. We have no financial stake in any detection vendor.

Start with who the errors land on, because that is where the familiar concession that the tools are merely imperfect quietly fails. The imperfection is not random noise. A Stanford study of seven popular detectors found that they flagged 61.22 percent of TOEFL essays written by non-native English speakers as AI-generated, while rarely misjudging native writers, and the researchers concluded that schools should avoid relying on the tools, especially where many students are still learning English (Stanford HAI). Survey data compiled by the Center for Democracy and Technology and reported by K-12 Dive points the same way inside American classrooms, where English learners and students with disabilities face a higher risk of discipline (K-12 Dive). A tool whose mistakes track a student's first language or disability is not a neutral integrity check. It is a bias engine wearing the costume of objectivity, and it delivers its worst outcomes to the students least equipped to argue back.

This is where the popular compromise, use the score as one signal among many, comes apart. A detector returns a percentage, and a percentage carries no reasoning that a teacher or a student can question. When Vanderbilt University disabled Turnitin's detector, its central objections were that the vendor would not explain how the number was produced and that a claimed one percent false-positive rate, applied to the roughly 75,000 papers the university processes in a year, would wrongly implicate about 750 students; the university concluded plainly that detection software is not an effective tool to use (Vanderbilt). A number that no one can cross-examine, yet that still colors how a teacher reads the next paragraph, is not a signal. It is an accusation laundered through a decimal point, and the one-signal-among-many framing simply spreads that accusation around more quietly.

Consider, too, who has walked away from this technology and who has walked toward it. OpenAI, the company whose models these detectors claim to catch, shut down its own AI-text classifier in 2023 after acknowledging a low rate of accuracy (TechCrunch). Schools moved in the opposite direction: teacher use of detection tools jumped to 68 percent in the 2023-24 school year, and the share of schools reporting discipline tied to generative AI climbed from 48 to 64 percent (K-12 Dive). Districts are scaling up a method that the most sophisticated builder of the underlying technology judged unfit and discarded.

The strongest case for keeping detectors deserves a fair hearing, because it is not foolish. Cheating with generative AI is real and growing, teachers are exhausted and outnumbered, and pulling detection without a replacement can feel like announcing that dishonesty now carries no cost, which corrodes the worth of every honest student's diploma. A flawed deterrent, the argument runs, still deters, and it hands an overwhelmed teacher somewhere to begin.

But a deterrent that threatens the innocent alongside the guilty is not enforcement. It is a lottery with a student's record as the stake, and it purchases whatever deterrence it offers by spending the trust of the very students who did nothing wrong. The real problem is worth solving, and it can be solved without that price: through in-class and handwritten writing, assignments that ask for process rather than only a finished product, drafts and version history that show a paper being built, and honest conversation when something looks off. Those methods are slower and more human, and they do not ask a district to accept a machine that accuses English learners and students with disabilities at higher rates and cannot say why. Keep the standard of honesty. Retire the tool that enforces it by guessing.

Take it further