LLMs could write like humans but post-training guardrails make their text detectable Bradley Emi, CTO of AI text detector Pangram, argues that large language models like ChatGPT, Claude, and Gemini could write as diversely as humans but post-training guardrails cause 'mode collapse,' making their text detectable. Pangram's detection does not flag base models, narrowly specialized fine-tunes, or broken outputs, but watermarks will likely always work even with base model variety. LLMs could write like humans but post-training guardrails make their text detectable LLMs could theoretically write as diversely as humans, but they don't. Post-training and safety guardrails keep their text detectable, argues Bradley Emi, CTO of AI text detector Pangram, in a blog post https://pangram.substack.com/p/no-llms-dont-just-mimic-human-text . Systems like ChatGPT, Claude, or Gemini learn behavioral rules to avoid dangerous outputs or censor certain political statements. This sharply narrows their expressive range, an effect called "mode collapse." So-called base models, the raw models before post-training, write with more variety, so Pangram's detection doesn't flag them, Emi says. The same goes https://x.com/max spero /status/2089938824474263644 for narrowly specialized fine-tunes trained only on Hemingway or certain subreddit texts, and for broken outputs like incoherent text. This only applies to non-watermarked AI text, though. Watermarks will likely always work https://the-decoder.com/anthropic-watermarks-claudes-output-but-critics-question-the-tradeoffs/ , even with a base model's variety. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Pangram https://pangram.substack.com/p/no-llms-dont-just-mimic-human-text