# The EU Wanted A Deepfake Detector. It Got An AI Scarlet Letter.

> Source: <https://www.techdirt.com/2026/08/20/the-eu-wanted-a-deepfake-detector-it-got-an-ai-scarlet-letter/>
> Published: 2026-08-20 18:03:16+00:00

# The EU Wanted A Deepfake Detector. It Got An AI Scarlet Letter.

### from the *won't-end-well* dept

I spoke a little about this on [last week’s Ctrl-Alt-Speech](https://podcast.ctrlaltspeech.com/2315966/episodes/19644518-watermark-my-words), but now that Anthropic has come out with more details about how its [text watermarking works](https://www.anthropic.com/news/claude-text-watermark), the debate has shifted into overdrive. [Many people are upset about it](https://www.businessinsider.com/claude-users-cancel-subscriptions-citing-anthropic-new-ai-watermark-2026-8), even though it appears that Google has already been doing something similar with the output from Gemini. I think that people are right to be upset, but for the wrong reasons, aimed at the wrong target.

The real problem is that the EU’s AI Act is aimed at a threat that never really materialized, and the end result will fall hardest on the people who get the most genuine benefit from these tools. Also, it doesn’t help that Anthropic chose to comply with the law in a manner that looks broader than what the law requires — though it did so for reasons that are more understandable than some of its critics suggest. Either way, though, the costs will fall most heavily on people who are using the tech properly.

Let’s take a few steps back first. If Anthropic’s explanation of how they watermark text isn’t clear enough for you, I think this explanation of [how text watermarking works](https://declaude.org/watermarking/) is much better. The fundamental thing to understand about generative AI is that it’s always trying to generate the next token, and it does so probabilistically, not deterministically, meaning that each time you’ll get something slightly different. The results already have biases in them (that’s part of the weights part of LLMs), but the companies can deliberately bias them in a manner that the tool is ever so slightly more likely to choose certain words based on a prompt than without that bias.

Think of it in the way that most random generators are not, in fact, “random.” With a little bit of effort, people can often figure out the “bias” of a random number generator, giving them an advantage in determining what number will be generated. Here, it’s the same sort of thing, but the “bias” is impacting the randomness of which word the tool will choose next.

With a long enough text, and a key regarding the bias, you can look at the text and see that enough of the very slight changes match the “watermark” bias, as to suggest the text was likely generated by a model carrying that watermark. Such a system is hardly foolproof, but it can absolutely call out text likely generated with that specific bias. Editing text after the fact may or may not get rid of the watermark, depending on whether or not the edits remove/change enough of the “biased” words.

Many of the people who are upset are so because they think they won’t be able to cheat any more, and… I don’t care one bit about them. There is some concern that the biasing will [make you choose worse words](https://buzzmachine.com/2026/08/15/words-matter-damnit/), but I’m also not all that concerned about that. As far as I’m concerned, if you’re using the AI to write for you, in some cases you’re already having it choose words, and it should be on you to know when and how to choose better words. The stronger version of that objection is about how the watermark will apply when merely editing content, in which case the watermark bias may nudge your prose towards specific synonyms. But even there, that seems much more like an argument to use the tools differently, not as a fully damning issue.

My larger concern is in how this will almost certainly be used to simply attack and denigrate people who actually are using the technology in a reasonable manner, but will be falsely accused of “cheating.” Just this week I wrote about the ways in which [I use AI tools](https://www.techdirt.com/2026/08/18/an-update-on-how-i-use-ai-to-help-with-techdirt-its-still-not-writing-articles/) to help with my work, not as a tool for writing, but helping me review and edit stories. In that article, I mentioned the [REAL Rating](https://www.realgoodai.org/real-rating) site, which offers a one-to-five scale to make clear that not all AI use is the same.

But a watermark reduces all of that to a simple binary: “did this use AI at all.”

And I worry about the fallout from that.

Which brings us back to where all this came from. As LLM tools started becoming popular, the EU rushed out its [EU AI Act](https://artificialintelligenceact.eu/), making it clear that they didn’t want to be slow to regulate, in the way that they (falsely) feel they should have regulated social media much earlier. Part of the early concerns with AI was that it would be used for “deepfakes” to fool people. As [we predicted](https://www.techdirt.com/2019/02/28/deception-trust-deep-look-deep-fakes/), that fear has been largely overhyped. There are still concerns that it *could* become a problem, but to date it really has not been. Most AI-modified content has been rightly called out as such. Indeed, the biggest thing with deepfakes tends to be people who are legitimately caught on video behaving badly [blaming deepfakes](https://isps.yale.edu/resource/the-liars-dividend-can-politicians-claim-misinformation-to-evade-accountability) (usually unsuccessfully) for their actual actions.

But because that was one of the biggest initial fears, the EU included a provision demanding “[transparency of AI-generated content.](https://digital-strategy.ec.europa.eu/en/policies/code-practice-ai-generated-content)” Providers are now required to mark LLM-generated content so that someone can identify it as such. And if deepfakes are a problem you’re trying to defeat, that could make some amount of sense.

However, deepfakes aren’t really much of a problem.

What we do have is a large and understandably frustrated group of people who are hostile to *any* AI use at all. And now they’ve been handed a detector with which to accuse anyone of being an AI user. So instead of using AI detectors to prevent people from being fooled, the tech is now being turned around to accuse people of using the tools in any way, shape or form. And that seems like a problem.

To be fair, the Code of Practice does not demand transparency for text that a model merely cleaned up, translated, or spellchecked. Anthropic still went ahead and watermarked all of it anyway. You could argue that this is a form of malicious compliance, but the reality is that it’s a bit more complex than that. Adding in this feature just for fully generated text would require separating out different kinds of text generation, which may also confuse things. Also, for most (but not all) merely “processed” text, there’s a lower likelihood of the watermark showing anyway, since identifying the watermark requires a decently long string of generated words.

Along those lines, Anthropic admits that the watermark is less likely to show up in software code, since code usually needs to be more exact. There are just fewer interchangeable ways to write a working function, which means there are fewer opportunities to embed the kind of word choice randomness to make the watermark effective.

The company also says that it has no “durable” way to limit the watermark such that it only impacts EU users covered by the EU’s law, which is why it’s rolled the feature out globally. I have a little difficulty believing that. Geofencing websites is pretty common practice to deal with geographic regulatory restrictions, such that even if it’s not perfect, it’s not clear why this needs to be rolled out globally. Indeed, a company clever enough to figure out how to do a generated text watermark can also figure out how to do some basic geoblocking.

Back to the potential fallout of all this: First, let’s acknowledge that plenty of AI generated content does, in fact, suck. And I absolutely get the same instinctual negative reaction many people get when I see something that is clearly AI generated. It feels lazy and often annoying.

But not every use case is the same, and plenty of people reach for these tools for reasons that have nothing to do with laziness. For example, [non-native speakers](https://www.nature.com/articles/d41586-023-02320-2) find these tools genuinely helpful in communicating more clearly with native speakers.

But now, attempts to use these tools to be a better communicator, may be dismissed as “fake” or “AI generated.” This is already a concern. A (potentially outdated) study from a few years ago found that non-native English speakers were [more regularly accused](https://www.cell.com/patterns/fulltext/S2666-3899(23)00130-7) of using AI tools. The same goes for Black students, who were [more likely](https://www.edweek.org/technology/black-students-are-more-likely-to-be-falsely-accused-of-using-ai-to-cheat/2024/09) to be accused of using AI tools.

And it’s not just non-native speakers. AI has been [important within the accessibility community](https://makeitfable.com/article/insights-ai-and-accessibility/), even as some have (totally understandable) [reservations](https://afb.org/research-and-initiatives/ai-series/ai-quagmire) about the tech, but many studies have found that the tech is [widely used](https://www.sciencedirect.com/science/article/pii/S1096751625000235) among those with disabilities.

Yet, the setup of an AI watermarking tool is that it gives you a simple binary “yes” or “no.” And as much as Anthropic says that finding its watermark simply *suggests* that its tools were used, rather than making it conclusive, we all know that most people will use it as a sign of proof.

On top of that, the more sophisticated users — generally those with more knowledge and understanding — will likely find it relatively easy to use tools that will effectively remove the watermark. In fact, such tools already exist. That same site I linked to above with the clearest explanation of how the watermarking works, also hosts a tool built to strip the signal out. So the people who actually get dinged by this system will be the less sophisticated, less resourced users — exactly the ones with the most innocent reasons for using the tool in the first place.

And the impact there can be massive. There are already studies that suggest the mere labeling of something as having been mediated by AI [causes people to trust it less](https://www.sciencedirect.com/science/article/pii/S2949882126000551), even if [the underlying content](https://arxiv.org/pdf/2506.16202) is the same. There’s even one study that found that truthful information, if it has an “AI disclosure” on it, is [often deemed as false](https://jcom.sissa.it/article/pubid/JCOM_2501_2026_A09/), even if the actual information is true. That seems quite unhelpful!

So the EU’s effort to ward off the potential threat of deepfakes that hasn’t materialized, has created a very real problem: now designating people who are using the tech in useful, helpful, ways as “cheaters” whose content can’t be trusted. And I’m not the only one noticing this. As I was finishing up writing this, [John Gruber pointed](https://daringfireball.net/2026/08/anthropics_watermark_text_adulteration_in_claude_is_a_perversion_of_writing) me to a blog post written by James Padolsey (the creator of the DeClaude tool above, and the writer of the excellent explanation of watermarks), with the fantastic title, “[Anthropic’s weak watermarks appease a weak law](https://blog.j11y.io/2026-08-12_Anthropics-weak-watermarks-appease-a-weak-law/)” in which he writes:

That is to say: this rule risks penalising people who are already less able to produce conventional prose unaided for using technological means to make their lives easier. The Act exempts standard editing, but many legitimate assistive uses require more substantial rewriting while leaving the ideas, judgment and responsibility with the human.The same thought that led to this law could have applied to calculators at the time of their inception, had their outputs revealed themselves through artefacts. Thankfully, a sum borne of the brain is treated no differently from one produced by a calculator. Likewise with spellcheckers. To make assistance suspect only once the tool becomes capable enough to compose a whole sentence is not a principled boundary. It is a moral premium placed on difficulty itself.

Anthropic has nevertheless chosen a blanket, model-level implementation that appears broader than the law’s minimum requirement. That may be convenient compliance engineering, but it discards distinctions the law expressly attempted to preserve.The result is a signal broad enough to implicate harmless and assistive use, yet fragile enough to be removed by a motivated person through substantial recomposition. It risks concentrating suspicion on ordinary and assistive users while remaining weakest against deliberate deception.

Indeed. So we have a mistargeted law, badly drafted to go after a theoretical problem that hasn’t proven to be real, combined with penalties harsh enough that companies like Anthropic comply in the broadest, bluntest way possible — and all of it dropped into a world that has increasingly decided that AI use is a binary good-vs-evil question, and where no nuance is allowed.

The watermark won’t actually catch the people it’s ostensibly aimed at — those seeking to deceive people. Those people will likely be sophisticated enough to remove any such watermark and walk away clean. However, it will likely catch the people who had every right to use the tool in the first place: the non-native speaker who ran their draft through Claude to sound more fluent, the assistive-tech user who needed help composing a sentence, or the writer who wants an extra level of review on any text they’ve written. They’re the ones who will wear the scarlet letter. The EU wanted to look tough on AI and Anthropic wanted to look compliant. But neither of them has to answer to the the disabled user who used the tech to help them communicate, who now gets accused of lying and cheating.

Filed Under: [ai](https://www.techdirt.com/tag/ai/), [deepfakes](https://www.techdirt.com/tag/deepfakes/), [disclosure](https://www.techdirt.com/tag/disclosure/), [eu](https://www.techdirt.com/tag/eu/), [eu ai act](https://www.techdirt.com/tag/eu-ai-act/), [generative ai](https://www.techdirt.com/tag/generative-ai/), [transparency](https://www.techdirt.com/tag/transparency/), [watermarks](https://www.techdirt.com/tag/watermarks/)

Companies: [anthropic](https://www.techdirt.com/company/anthropic/)
