cd /news/ai-policy/ai-text-watermarking-and-quality-los… · home topics ai-policy article
[ARTICLE · art-109327] src=blog.keyvan.net ↗ pub= topic=ai-policy verified=true sentiment=· neutral

AI text watermarking and quality loss

Anthropic's watermarking method for Claude, introduced to comply with the EU AI Act, may reduce diversity across multiple responses to the same prompt, according to the company's own supplementary documentation. Technology writer John Gruber argues that any quality loss is unacceptable, while critics contend the method preserves quality. The trade-off between diversity and detectability could affect systems that generate multiple responses and select the best one.

read5 min views5 publishedAug 24, 2026
AI text watermarking and quality loss
Image: Blog (auto-discovered)

After the EU AI Act’s transparency guidelines went into effect, and following the announcement that most major AI labs will soon be watermarking their text output to comply, there has been some debate about how much of an impact this is going to have, if any, on the quality of the text produced by AI models.

Since the focus of AI labs shifted to code production, I think the written output of the models has gotten much worse.1 But it’s still interesting to try to understand if watermarking could make it even worse.

I found the discussion around technology writer John Gruber’s recent post, Anthropic’s ‘Watermark’ Text Adulteration in Claude Is a Perversion of Writing, very interesting.

On watermarking, Gruber wrote:

It’s unacceptable for a tool to sacrifice an iota of clarity, coherence, meaning, quality, etc. for the purpose of embedding hidden clues within the text to suggest its provenance.

He got a lot of pushback, with many debating how you would even measure quality, or define it.

But some argued that watermarking does not affect quality at all: it’s essentially just using different random numbers to generate the text. Daniel Jalkut, whom Gruber linked to in his follow-up, wrote:

…there are also ways of tweaking the algorithm that would not diminish its randomness. For example if a random number generator routinely reversed the order of the digits in a generated number, the randomness would remain the same.

Similarly, in another discussion, someone described it as essentially swapping two faces of a die.

In a follow-up piece, Gruber wrote: I hope that’s true. I believe it’s possible that it is true. I think it’s highly unlikely that it is true.

I think Gruber is right.

The paper linked by Anthropic assures everyone that the method “preserves response quality”, based on some human evaluations. But that’s not the same as saying that watermarking has no effect on the output, that its effect is essentially swapping one set of random numbers for another set of random numbers, the way described above.

The 56-page supplementary PDF they link to at the bottom provides some more info on the method used, and it suggests that there’s more to this than the interpretation above. Notably, it says that there is a “diversity/detectability trade-off”:

While single-sequence non-distortion guarantees the quality of each individual response, it does not necessarily preserve diversity across multiple responses Okay, so what does that mean? They go on to explain:

when sampling several responses to the same prompt, the similarity between the responses is greater for the watermarked responses than the unwatermarked responses.

Now, if they had truly only swapped one set of random numbers for another, we would not be seeing any such difference compared to the unwatermarked responses.

They go on to write:

This could be problematic in scenarios where inter-response diversity is important, or could lower the overall quality of a system which generates many responses then selects the best one

Okay, so in their own words, it can affect overall quality after all. But nothing to worry about, because this would only concern you if you cared about “inter-response diversity”, for which they only provide one example: a system which generates and selects responses.

But what if you, the user, want to generate different options for you to then choose from? Some argue the model has no notion of the best next word to use for what each individual is trying to convey in their specific work, and so we shouldn’t get too hung up on the next word choice. But surely they’d agree that you, the individual, would know if it’s suitable or not once you see it. Let’s say you prompt: “I don’t like the ending to this sentence. Please change it: [sentence].” But you don’t like the example given. No problem, hit the retry button. Well, now you’re affected by watermarking, because you have now hit the “inter-response diversity” problem: you won’t get as many creative options as you would with an unwatermarked model.

Here’s another example of what this affects: one way people try to determine the confidence level of a response, to lower the chances of hallucinations, is to repeat the same prompt a number of times, to explore the different pathways the model can take, and then see how often it lands on the same answer. They’d then pick the most common one. Well, this is also affected. Now, you’re getting a more limited set of responses weighted by a secret key.

What’s interesting is that while they point to this affecting quality, I couldn’t find anywhere in the paper where they tested this particular aspect of the watermarking method with humans. If they didn’t, the wonderful human results they point to did not involve testing the most affected part of the watermarking method.

So it seems to me that they are limiting the creative range of their models by implementing watermarking, and they don’t think you’ll notice or care much.

1 Jasmine Sun recently looked at why AI models can’t write well. Here’s what Katy Gero, a poet and computer scientist, had to say:

The same whimsicality that made GPT-2’s voice fresh also made it prone to other unpredictable behavior. “If you’re a big corporation like Google or OpenAI, you want a chatbot that’s going to make money. The chatbot that’s

notgoing to make you money is the one that’s a weirdo,” Gero said.

── more in #ai-policy 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-text-watermarking…] indexed:0 read:5min 2026-08-24 ·