Even Claude says AI watermarking is no ‘silver bullet’ Anthropic's Claude chatbot has questioned whether watermarking AI-generated text can ever be a foolproof identification method, saying simple marks can be cropped or edited out and sophisticated text watermarks can be defeated by paraphrasing, just days after Anthropic began embedding invisible marks into Claude's output. The chatbot backed watermarking as 'worth doing' but not a 'silver bullet', noting that those most likely to misuse AI are exactly the people motivated to strip watermarks. Anthropic is the first major AI company to detail watermarking under new EU transparency rules, adapting Google DeepMind's technology, while OpenAI and Google are also working on similar measures. Even Claude says AI watermarking is no ‘silver bullet’ Claude has questioned whether watermarking AI-generated text can ever provide a foolproof way to identify machine-written material, just days after its maker, Anthropic, began embedding invisible marks into the chatbot’s own output. Asked by City AM whether AI-generated content should be watermarked, Claude warned that simple marks can be “cropped, screenshotted, or edited out”, while more sophisticated text watermarks can often be defeated by paraphrasing or reformatting. The chatbot ultimately backed watermarking, but described it as only one part of a wider system for identifying AI content, not a “silver bullet”. “I think watermarking is worth doing, but I’m sceptical of it as a headline solution,” Claude said when pressed for its view. It argued the technology was weakest against precisely the people most likely to deliberately misuse AI. “The people most likely to weaponise synthetic content for fraud, disinformation, or non-consensual imagery are exactly the people motivated to strip a watermark,” Claude said, adding that for text this could be as simple as rephrasing it or translating it into a different language and back again. The response comes after Anthropic became the first major AI company to detail how it will watermark Claude-generated text under new EU transparency rules. Claude leads watermark push Anthropic has begun embedding an invisible mark into text generated by newer Claude models, with the system applying globally rather than just to users in Europe. Rather than inserting hidden characters or a visible label, Claude slightly changes the way it chooses between possible words as it generates a response. The technique is adapted from Google DeepMind’s technology, and is designed to leave a pattern that can later be detected with the appropriate verification system. Anthropic also plans to allow third parties to check text for the signal. But the tech behemoth itself has stressed that finding the mark is not definitive evidence that Claude wrote a piece of content from scratch. Human-written work can be watermarked after being translated or substantially edited by Claude, while text originally generated by the chatbot can eventually lose its detectable signal if it is heavily rewritten. Short passages and types of output where Claude has little choice over its wording can also be more difficult to mark reliably. Code, for example, generally carries less watermarking because producing functioning software can require exact outputs. The chatbot recognised those shortcomings when City AM questioned it, saying reliable text watermarking that survives editing was not “close to solved”. It was more positive about applying the technology to images or video, where invisible marks can be more difficult to remove without degrading the underlying content and where deepfakes or fake evidence present potentially greater risks. For text, Claude said it would put greater emphasis on disclosure when AI-assisted material is published, provenance standards and media literacy. AI firms split on following suit Anthropic’s move follows the introduction of transparency requirements under the EU AI Act https://artificialintelligenceact.eu/ , which is pushing developers towards making machine-generated material identifiable. The company has signed the EU’s voluntary Code of Practice and said it will ultimately extend watermarking to older Claude models. OpenAI is also working towards marking text outputs as part of its commitments under the European rules, while Google is expected to introduce measures covering AI-generated content. OpenAI has said its goal is to expand existing provenance signals across different forms of content, including text, although it has not said it will adopt the same method as Anthropic. But City AM revealed last week that Elon Musk’s xAI has taken a different route. https://www.cityam.com/chatgpt-might-follow-claudes-watermark-pledge-but-grok-to-swerve-it/ The Grok developer was the only major large language model maker not to sign the EU code, meaning it has made no equivalent voluntary commitment to watermark its chatbot’s output. It will still have to comply with applicable AI Act requirements. Anthropic’s decision has meanwhile prompted criticism from some Claude users concerned that genuinely human work could be identified as AI-assisted simply because the chatbot had subsequently edited it. The tech giant stressed that its watermark contains no information identifying a particular person, company or conversation. It has also acknowledged that light editing may leave the signal intact while a complete rewrite can remove it.