EU-enforced update leads to fears that people will be accused of using artificial intelligence with their own writing
- Bookmark
- CommentsGo to comments
The new update is part of the EU's AI Act, which requires that artificial intelligence companies introduce ways of spotting AI-generated content. But it comes amid broad concern that the internet is filling up with low-quality media generated with artificial intelligence, which has come to be referred to as "AI slop".
In keeping with AI rules, those invisible watermarks will be placed onto both text and images. It will be implemented across the world, and even in specific uses of the Anthropic chatbot, such as Claude Code.
Watermarks in AI-generated images have become widespread and are relatively simple. They leave a set of clues in the image that are invisible to a human but can be detected by automated systems.
They do however come with limitations. The watermark may be removed when a user converts an image or takes a screenshot of it, for instance.
But perhaps the most surprising and controversial change is that Anthropic said it would introduce similar features for the generation of text. That too will include an invisible trail that automated systems will be able to spot but will be invisible to humans, the company said.
The feature has already proven controversial among some users who fear that asking Claude for edits on a text might mean that it is falsely flagged as having been produced by an automated system, for instance. Anthropic stressed that was a limitation of the system, and that the "Claude mark" might be visible even in examples where "the underlying ideas, text, or data originated from another source".
"A detected mark provides a signal that content was processed by Claude, but is not fully conclusive," Anthropic warned. "Detecting a Claude mark tells you that the content may have been processed by Claude," it said, noting that a mark "does not, on its own, confirm the full provenance of the content.
It also noted that the lack of such a watermark does not necessarily mean that the text was not generated or processed using AI. It might have come from a model before the mark was introduced, heavily edited or changed, or be too short for it to include a reliable signal, for instance.
Anthropic did not give any specific detail on how exactly it will integrate the watermarks into its text, presumably at least in part to avoid people finding ways to remove it. But similar technology tends to work either by using specific, invisible characters or choosing specific words more than others – and it is thought that Claude is using the latter.
That might include using a higher or lower proportion of a particular word, for instance, so that it would be picked up in a statistical analysis. Anthropic stressed that it would not be perceptible and that it would not "change the meaning, quality, or readability of Claude’s response".
In the past, other technologies for detecting AI-generated text have run into a variety of problems.
It might be possible to save all generated messages from a chatbot and check suspect text against it, for instance, but that brings both practical and privacy concerns about how to store all of those conversations. Other options already offered include checking text after it is generated for the characteristic signs of AI writing – such as particular words or sentence constructions – but that too can require significant computing resources and does not always work consistently.
Initially, Anthropic will add the marks to any content generated with models that were introduced after 2 August of the year, it said. But it is looking to add support to older models, it said.
Join our commenting forum #
Join thought-provoking conversations, follow other Independent readers and see their replies
Comments