Somewhere in America right now, a student is running her own essay through three different AI detectors before she submits it. She wrote every word herself — badly, earnestly, in her second language — and she has learned that this is exactly the profile of an essay that gets flagged. So she pastes her conclusion into a free checker, watches it return “87% likely AI-generated,” and begins the strange new ritual of the honest student: vandalizing her own prose until a machine agrees she is human. She swaps in a longer word. She breaks a clean sentence in two. She is making her writing worse on purpose, for an audience of one, and the audience is software.
This is where the first year of the AI homework wars has landed, and it is worth pausing on how odd it is. The take-home essay — the load-bearing wall of humanities education for a century — stopped functioning as an assessment almost overnight, sometime between ChatGPT’s launch in late 2022 and the following spring semester. The institutional response was to buy detection software. The detection software, it turns out, does not reliably work, and fails hardest on the students least equipped to fight an accusation. The interesting question is no longer whether AI can write a passable essay. It is what the essay was ever measuring, and why that measurement broke so cheaply.
Start with the detectors, because their failure is documented with unusual clarity. OpenAI — the company with the most to gain from a working classifier — released one in January 2023 and quietly withdrew it that July, noting that it was “no longer available due to its low rate of accuracy.” Its own published evaluation had the tool catching 26% of AI-written text while falsely flagging 9% of human writing, which is roughly the performance profile of a coin flip with anxiety. Independent testing was not kinder. A study in the International Journal for Educational Integrity ran fourteen detection tools, including Turnitin and GPTZero, against human, machine-written, and machine-paraphrased text: none reached 80% accuracy, and performance collapsed once the AI text had been lightly paraphrased — which is to say, once a student spent ninety seconds covering their tracks.
The bias findings are worse. In a 2023 paper in Patterns, Stanford researchers fed 91 TOEFL essays — written by non-native English speakers — through seven widely used detectors. The detectors marked an average of 61.3% of them as AI-generated. At least one detector flagged 97% of the essays. The same tools produced near-zero false positives on essays by American eighth-graders. The mechanism is almost elegant in its cruelty: detectors lean heavily on “perplexity,” a measure of how predictable a text’s word choices are. Fluent, idiomatic writing is unpredictable. Careful, correct, second-language writing is predictable. The detectors were, in effect, grading English-language learners for sounding like they had learned English from a textbook — which they had.
Turnitin, whose AI indicator rolled out to institutions in April 2023, claims a document-level false-positive rate under 1%, a figure that sits uneasily beside the independent evaluations. Universities have been voting with their settings panels. Vanderbilt disabled the tool in August 2023, citing false-positive risk; UBC declined to enable it at all; since then, institutions from Yale to Northwestern have barred, discouraged, or switched off detector reliance. The market sold a smoke alarm that sounds off in some kitchens more than others, and the smarter buyers have started unplugging it.
The proxy problem
But fixing the detectors would not fix the thing that actually broke, and this is the part the panic columns keep missing. The take-home essay was never a direct measurement of thinking. It was a proxy — a receipt for thinking. The assumption baked into a century of pedagogy was that producing a coherent essay was expensive in a very specific way: it required the writer to have done the reading, formed the argument, and wrestled both into paragraphs. The artifact stood in for the process because faking the artifact cost roughly the same as doing the process. That is what a proxy is: a stand-in that works only while it is expensive to counterfeit.
What large language models did to the essay was not cheating, exactly — it was counterfeiting at industrial scale. The cost of producing a plausible essay fell to the price of a prompt, and a proxy that costs nothing to fake is not a weakened proxy. It is no proxy at all. The student who prompts her way to a B+ has not broken the essay; she has revealed that the essay was always an indirect instrument, and indirect instruments have a shelf life set by the price of faking them.
Seen this way, the detector debacle was a category error from the start. Schools responded to a collapsed proxy by shopping for a better counterfeit-detector — trying to defend the receipt instead of re-examining the transaction. Even a hypothetical perfect detector would only restore the arms race at a higher level, leaving the same measurement problem untouched. Meanwhile the imperfect detectors we actually have impose their costs asymmetrically, on the student who writes in careful textbook English and cannot afford a lawyerly appeal.
What thinking costs now
The honest position is uncomfortable in both directions. The boosterish line — that AI is just a calculator for words and schools should get over it — ignores that the essay’s entire purpose was to make a student do the thinking the calculator now does for her. The panicked line — that we must hunt down every synthetic sentence — has already wrecked real students on the strength of software that its own makers withdrew. Between those two is a duller, more expensive truth: if you want to assess thinking, you have to watch it happen.
That means in-class writing, oral exams, drafts with version histories, a conversation in office hours where a student has to defend her third paragraph out loud. These are the oldest assessment technologies we have, and they share one property: they consume teacher time, which is the one resource nobody has budgeted. A professor with 120 students cannot hold twenty-minute vivas and also grade the problem sets and also answer the email. The essay survived for a century partly because it was cheap to administer at scale. Its replacements are not. The proxy failed, and the non-proxy costs money.
So the real fight over AI and homework was never about catching cheaters. It is about whether institutions will pay — in staffing, in class sizes, in tuition — to measure the thing they always claimed to measure. The student quietly degrading her own essay to appease a detector is the tell. The essay is dead. What replaces it will be decided, as these things usually are, by whoever is willing to pay for the more expensive measurement — and by how long everyone else can pretend the receipt still means something.