LLMs could theoretically write as diversely as humans, but they don't. Post-training and safety guardrails keep their text detectable, argues Bradley Emi, CTO of AI text detector Pangram, in a blog post. Systems like ChatGPT, Claude, or Gemini learn behavioral rules to avoid dangerous outputs or censor certain political statements. This sharply narrows their expressive range, an effect called "mode collapse."
So-called base models, the raw models before post-training, write with more variety, so Pangram's detection doesn't flag them, Emi says. The same goes for narrowly specialized fine-tunes trained only on Hemingway or certain subreddit texts, and for broken outputs like incoherent text. This only applies to non-watermarked AI text, though. Watermarks will likely always work, even with a base model's variety.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Subscribe now