cd /news/artificial-intelligence/anthropic-plans-to-watermark-claude-… · home topics artificial-intelligence article
[ARTICLE · art-99006] src=eshumarneedi.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Anthropic Plans to Watermark Claude-Generated Text

Anthropic announced that future Claude models will watermark generated text to comply with the EU AI Act, effective August 2, and the method will not affect output quality, add hidden characters, or increase cost. The watermark is created through subtle word choices and is detectable only with a key, and it will not apply to code. Other major AI providers have signed the same Code of Practice and will implement their own watermarks.

read5 min views1 publishedAug 16, 2026

Future Claude models will generate text that contains a watermark. This is a way of determining the likelihood that Claude was involved in writing the text, and we, along with several other major AI providers, are implementing this change to comply with the EU AI Act.

In this article, we share answers to some of the questions we’ve received about how our chosen watermarking method works, whether it affects Claude’s outputs, and why we’re making this change. To summarize:

We use a method of watermarking that does not have any practical impact on the quality or content of Claude’s outputs;

The difference between watermarked and un-watermarked text will not be distinguishable to readers;

Nothing is added to the text and there are no hidden characters;

Watermarking doesn’t require extra tokens, and will not be more expensive;

Watermarking carries no identifying information and can’t be traced to a specific person, organization, or chat;

Watermarking won’t be specific to Claude. As of August 2, the EU requires AI providers serving its market to mark AI-generated content. Other major model developers have signed the same Code of Practice and will be implementing their own watermarks.

Large language models like Claude work by generating one word at a time. Each time the model decides on the next word, it chooses among a list of potential candidates, ultimately selecting the most sensible or likely based on the preceding text. Take the sentence “The weather today was cold and…”. The next word is very unlikely to be “sugary.” But it is quite likely to be “overcast” or “grey.” Under most circumstances, it doesn’t matter much to the reader which of these latter two words the model ultimately chooses—the meaning of the sentence is largely the same either way. In cases like this, the choice is settled by a random number.

Watermarking uses low-stakes choices like these—which occur many times over a piece of generated text—to leave a pattern in Claude’s responses. That pattern is undetectable to the reader, but is detectable to anyone who has a key that encodes it. When watermarking is used, choices are still made at random, but the source of the randomness is different. Instead of using an arbitrary random number generator to pick the next word, watermarking uses the key and a few words that come before to settle what word the model should pick. That is, the words that Claude picks are still random, but now, one can check the sequence of words and see if it’s consistent with the choices Claude would make if it was using the key. If it is, one can assign a probability that the text was generated by Claude.

Because the watermark is created through subtle differences in word choice, it won’t apply to code, which requires exact syntax to function properly.

I think it’s important for there to be a foolproof way to identify artificial intelligence-generated outputs, whether they be images, text, or audio. Social media platforms should also prominently mark AI-generated images and text as synthetic, especially if the content is political or official. I stipulate that identifying AI-generated content should be foolproof because false positives can have detrimental consequences. Pangram, the leading AI text detector which doesn’t — at least for now — rely on watermarking, is widely regarded as the most accurate but still can occasionally flag human-created text as AI-generated. Universities across the globe punish their students for passing off AI-generated work as their own; the consequences of these systems are enormous.

SynthID, a framework developed by Google DeepMind, embeds an invisible watermark within AI-generated images and videos. It remains intact even after manipulation or screenshots, and is completely random, meaning it cannot be faked or removed. SynthID is used by Google’s Nano Banana image generator and ChatGPT’s image generation feature, two of the most popular tools. Social media websites like Instagram, Facebook, X, and YouTube scan uploaded content for SynthID watermarks and display a badge next to synthetic imagery. SynthID is foolproof and one of the best systems for detecting AI-generated imagery online, but it just can’t be replicated via text.

This new watermarking framework outlined Friday is Anthropic’s first try at developing a text watermarking technology. The problem is that it objectively isn’t foolproof — again, SynthID cannot be replicated with text. I’m less concerned that the outputs will be lower in quality than those without the watermark, but I am worried about false positives. Text that has been edited by Claude shouldn’t have a watermark at all, whereas Claude-generated text edited by a human absolutely should. Here’s what Anthropic says:

Light editing probably won’t remove the watermark completely; a complete rewrite where every word is replaced will. In the latter case, of course, it’s arguable whether the text can any longer be described as AI-generated…

The watermark only applies to words Claude chooses. When Claude proofreads text written by a person, what it gives back has generally only been lightly edited; because nearly all the words are the person’s, there’s very little (if anything) for the watermark to attach to. Depending on the length of the text and how heavily Claude has edited it, those changes might not be enough to make Claude’s involvement detectable. The more Claude writes, the more decisions it has to make, and the more space there is for a watermark.

These paragraphs don’t project much certainty. “Probably won’t,” “generally only,” “might not be enough.” I don’t blame Anthropic for this — watermarking text is nearly impossible! Claude relies solely on word choice to watermark its outputs, and word choices can be changed trivially. But I’d argue that releasing a first-party watermarking tool that isn’t 100 percent accurate is a dangerous precedent. People will use this tool to check for AI-generated text, but because it’s a first-party framework, arguing against its judgment will be an uphill battle. This is not Pangram; it’s not a guess.

Again, I do think all AI-generated text in the future ought to be watermarked, but with the important caveat that the watermark is foolproof. This watermark doesn’t guarantee that, and I think it ought to be scrapped.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/anthropic-plans-to-w…] indexed:0 read:5min 2026-08-16 ·