The Web Is 31% AI Slop — And It’s Poisoning Future Models

The Web Is 31% AI Slop — And It’s Poisoning Future Models

Nearly a third of the tokens crawled from the open web this August are AI-generated. That’s not a warning about the future — it’s the data your next model is training on right now. The researchers ran 800 pretraining experiments to find out what happens next, and the answer is not “fine.”

What happened

Figure 5: Filtering AI text pays off at larger human budgets, and repeating human text beats adding AI text. Left: change in loss from removing the AI documents of a 22.3% AI mix without replacing the
Figure 5: Filtering AI text pays off at larger human budgets, and repeating human text beats adding AI text. Left: change in loss from removing the AI documents of a 22.3% AI mix without replacing the

A team from Pangram Labs and UMass Amherst crawled web data through mid-2026 and found that 27.5% of tokens from June web data were flagged as AI-generated content by their own Pangram detector — rising to 31.1% by August. This isn’t controlled synthetic data; it’s “wild” AI text written for human readers, produced by many different models, arriving unlabeled in pretraining corpora. To measure the damage (and the upside), they pretrained 800 language models at varying ratios of AI-to-human tokens and fit scaling laws to held-out losses on both text types. The headline finding: for data-starved models, a little AI text helps — but the benefit saturates fast and then reverses. For models that have seen a lot of human text already, AI tokens raise held-out human-text loss almost immediately, while fresh human tokens keep lowering it. Existing scaling laws, including Chinchilla (Hoffmann et al. 2022), cannot predict this sign-flipping behavior; the authors propose a new law with separate benefit and harm terms that reduces 41% prediction error over all AI ratios on models up to 3.6× larger than those used for fitting. They release WildAI, an 83B-token labeled corpus, all 800 models, and code.

Cold read

The AI content detection step is load-bearing and unvalidated here: the Pangram detector labeling 31% of web tokens as AI-generated is the foundation of every downstream finding, and the abstract tells us nothing about its false-positive rate or how it handles the genuinely murky middle — heavily AI-edited human text, templated content, or non-English pages. The scaling law extrapolates from smaller models to ones 3.6× larger; “41% lower error than the best existing law” sounds good until you realize existing laws were already failing badly on this data distribution, so the baseline is weak. The experiments vary the ratio of AI tokens added to a fixed corpus — this is not the same as the real pretraining scenario where you can’t separate or label the AI text at all, because it arrives unlabeled. The recommendation to “filter AI text” assumes you can reliably identify it at scale; the paper doesn’t establish that you can. And the finding that AI text remains valuable “when the target is AI text” opens a convenient escape hatch that every AI-output-heavy product will exploit without asking whether their users actually want a model optimized for AI-flavored prose.

What it means for you

  • Signal maturity: 3/5 — Solid empirical base, but the detector dependency and scale gap limit immediate applicability
  • Who gets hurt: Founders building proprietary training pipelines on CommonCrawl or FineWeb derivatives without provenance filtering — your data quality is degrading faster than your evals will catch
  • What breaks if this is true: The implicit assumption behind every “more data = better model” roadmap collapses; past a threshold, crawling more web text actively harms your human-text benchmark scores, which means your next training run could be worse than your last even with more compute
  • Why it might not land: If the Pangram detector is significantly miscalibrated (plausible — benchmark contamination of detectors is a real problem), the 27–31% figures are fiction, and the scaling law is fit to noise
  • Watch for: A major lab publicly announcing a “human-text-only” or “AI-ratio-capped” data policy for their next frontier model — that’s the signal that this finding has been internally replicated at scale

Forecast as of 2026-10-01

By Q3 2027, at least two frontier model training efforts (OpenAI, Anthropic, Google, Meta, or xAI) will publicly disclose AI-content ratio caps or provenance filters as a named component of their data pipeline — either in a technical report or a regulatory filing — citing contamination of web corpora as the driver.


Source: How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text — Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi. https://arxiv.org/abs/2609.40295v1

Similar Posts