Somewhere in the last few years, millions of people quietly changed how they search the internet. The query is the same as ever — best budget espresso grinder, why is my monstera dying, does this rash look bad — but the habit now ends with an extra word: reddit. It is a small, almost apologetic act of filtration. What it means is: show me something an actual person wrote. Not a content farm, not an SEO smoothie, not a page that reads like a brochure for itself. A human being, ideally a slightly obsessive one, with direct experience and no business model.
The machines, it turns out, want the same word appended. The entire AI industry now shares your problem — it is starving for text it can prove came from a person — and that convergence is more interesting than either of the stories usually told about it. Because the strange part was never that the web contained human thought. It’s that nobody put it there on purpose. The web’s corpus is the exhaust of ordinary life: arguments, confessions, product reviews, forum threads about broken dishwashers, written by people who had no idea they were contributing to anything. We are training the engines of the future on that exhaust, and the exhaust is starting to contain itself. The interesting question is what suffocates first.
The worry has a name: model collapse. In July 2024, Ilia Shumailov and his colleagues published a study in Nature with the blunt title “AI models collapse when trained on recursively generated data.” Their finding, demonstrated across several kinds of model, including a small language model fine-tuned over and over on its own output: when a model’s training data consists increasingly of what earlier models produced, degradation follows a specific and eerie order. In early collapse, the model loses the tails first — the rare cases, the minority patterns, the weird stuff at the edge of the distribution. In late collapse, it converges on a narrow, samey mush that barely resembles the original data. In the simplest case they analyze, the mathematics is unforgiving: the drift away from the true distribution is unbounded, the variance shrinks toward zero. Everything becomes average, then becomes a caricature of average.
This is the point where the doom version of the story usually takes over: AI eats itself, the ouroboros chokes, the whole industry is a pyramid scheme selling its own exhaust back to itself. It is a satisfying story, and it overstates the paper. The inevitability result is proven for a specific regime — indiscriminate recursive training, in which generated data replaces the human original and comes to dominate the training set. The authors are explicit that sustained access to genuine human data prevents collapse, and in one experiment reported by MIT Technology Review, letting each new generation keep just 10 percent original data alongside the synthetic stuff was enough to blunt the damage. The labs, which employ people who can read, are filtering their crawls, curating their mixes and paying for text with a verifiable human pedigree. The doom story is comforting in a backhanded way — it imagines the problem correcting itself, the machine poisoned by its own tail. The paper describes something more conditional: a machine that gets sick only if you stop feeding it us.
But the complacent version — relax, the labs have it handled — misses the part of the paper that should actually worry you. Tucked into its conclusion is an acknowledgment that telling machine text from human text at scale remains unsolved, and that tracking the provenance of crawled data is now essential. In other words: even if every model on earth stays healthy, the open web is losing its value as a record of what humans independently thought. The contamination is not of the machines’ food supply but of the commons — the thing historians, linguists, social scientists and your espresso-grinder search all drink from. The authors put it in unusually plain terms for a scientific paper: data collected about “genuine human interactions with systems” will become “increasingly valuable,” and whoever trained early, on mostly human text, may hold a durable “first mover advantage.” That is a strange thing to read in Nature. It is an obituary for a resource, written by the people measuring its depletion.
Bruce Schneier sketched the social version of this future in an April 2024 essay — an argument, not a finding, but a hard one to wave off. If everything published in the open is instantly absorbed into the machines, he wrote, creators will retreat to walled-off audiences and “the great public commons of the web will be gone.” The timeline is debatable; the direction is not. Platforms increasingly treat crawling as trespass while the open crawl gets strip-mined, and there is still no foolproof way to detect machine-written text — which means doubt alone does the damage. A corpus you cannot audit is a corpus you cannot trust, whether or not it is mostly clean. Provenance, as one recent ACM commentary put it, has become a systems problem rather than a model-quality problem, which is to say: everyone’s problem.
So the scarce resource turns out not to be quality — that can be engineered, filtered, bought. It is provenance: proof of human origin, a chain of custody for thought. “You need to know where those data are coming from,” Shumailov told MIT Technology Review, which is a sentence you could embroider. Your reddit suffix was a homemade provenance filter before the industry had a name for the need. That is the real irony of the feedback loop. We spent three decades building the largest record of human life ever assembled without ever treating it as one, and we are noticing what it was only now that a machine wants to read it. The question ahead is not whether the models choke on the exhaust; fed carefully, they won’t. It is whether “written by a person” becomes a certification — the wild-caught label of text — while the default supply turns farmed. The early movers already have their freezer: the web as it was, back when it was mostly us. Everything published since is out in the weather.