cd /news/artificial-intelligence/claudes-scarlet-letter · home topics artificial-intelligence article
[ARTICLE · art-95898] src=techstrong.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Claude’s Scarlet Letter

Anthropic plans to embed invisible, machine-readable watermarks into text generated by supported Claude models, including Claude, Claude API, Claude Code, and Claude Cowork, to support transparency and provenance requirements such as those in the EU AI Act. The company acknowledges that detection only indicates content may have been processed by Claude, not that it was authored by AI, raising concerns about false attribution. The move shifts the debate from disclosure to traceability, with implications for publishers, employers, schools, and platforms.

read10 min views1 publishedAug 13, 2026
Claude’s Scarlet Letter
Image: Techstrong (auto-discovered)

TL;DR — Key Takeaways

  • Anthropic plans to embed invisible, machine-readable watermarks into text generated by supported Claude models.
  • The watermark is intended to support transparency and provenance requirements, including provisions under the EU AI Act.
  • Anthropic itself says detection only indicates that content may have been processed by Claude; it does not prove Claude authored the work. - The biggest concern is false attribution: A technically accurate watermark could lead publishers, employers, schools or platforms to wrongly conclude that AI wrote material created primarily by a human.

I did not expect to write a third article about writing with AI quite this soon. But Anthropic apparently decided that two were not enough.

The first was “I’m Begging You: Learn How to Write With AI.” My argument was that there is an enormous difference between writing with AI and surrendering the act of writing to AI. Used poorly, AI can replace thinking. Used well, it can challenge your ideas, expose holes in an argument, organize research and make you work harder at expressing what you actually mean.

My rule was simple: Never publish anything with AI that you have not thought through, made your own and accepted responsibility for.

That article attracted a large audience and a lively discussion on LinkedIn. Clearly, many people were wrestling with the same questions about where assistance ends and authorship begins.

A few days later, I wrote “We Were Promised Flying Cars. We Got AI Writing Detectors.” By then, I had read pieces from my friend Brad Feld and Reid Hoffman that approached the issue from different directions but arrived in much the same place.

Brad supplied the historical perspective. Anxiety about technology changing the way humans think and write is at least 2,400 years old. Plato recorded Socrates warning that writing itself would weaken memory and create the appearance of wisdom without the real thing. Every generation believes its newest communications technology poses a unique threat to human thought. Somehow, writing, printing presses, typewriters, word processors, spellcheckers and search engines all became part of the process rather than the end of it.

Reid supplied another essential idea: A byline is a promise. It tells the reader that the person whose name appears at the top accepts responsibility for what follows. The byline does not document which tools were used, who conducted background research, whether an editor rewrote a paragraph or how many times the author changed direction. It identifies the person standing behind the finished work.

I had written my first article before seeing either Brad’s or Reid’s. The three of us had independently reached compatible conclusions. “AI writing” is probably a transitional category. The technology will eventually disappear into the process, while judgment, accountability and responsibility remain with the human being whose name appears on the work.

That second article also generated large readership and substantial discussion on LinkedIn. This is not a parlor game for technologists. Writers, editors, executives, educators and readers are trying to establish the rules of authorship while the tools are evolving beneath them.

Now Anthropic has moved the argument from disclosure to traceability.

The company has announced that supported Claude models will weave an invisible, machine-readable watermark directly into the text they generate. The marking will happen at the model level, which means it will apply across Claude, the Claude API, Claude Code, Claude Cowork and supported cloud platforms. It will travel with the words when they are copied and pasted and may survive at least some editing.

Anthropic is also working on mechanisms that will allow users and third parties to detect the marks. For supported image files, it will attach digitally signed provenance metadata using the C2PA standard.

This is Anthropic’s effort to comply with transparency provisions of the European Union’s AI Act. The goal is not unreasonable. We need better ways to identify deepfakes, automated propaganda, academic cheating, impersonation and industrial-scale synthetic slop. When someone uses AI to deceive an audience about the origin or authenticity of content, provenance can provide valuable evidence.

But evidence of what?

According to Anthropic’s own documentation, detecting a watermark indicates only that content “may have been processed by Claude.” It does not establish that Claude originated the ideas or wrote the entire piece.

Anthropic acknowledges that someone may use Claude to proofread, translate, summarize or convert material that originated somewhere else. A person could write an article from beginning to end, ask Claude to tighten a few awkward paragraphs and receive marked text in return. The company also acknowledges that marked material may subsequently be edited, excerpted or combined with other writing.

That is a very large caveat to attach to an invisible mark that other people and platforms may eventually treat as a verdict.

“Claude processed this text” is not the same as “Claude wrote this article.” A watermark cannot tell us who developed the thesis, conducted the interviews, selected the evidence, challenged the counterarguments, rejected Claude’s bad suggestions or rewrote the passages that did not sound right. It cannot tell us who checked the facts, made the final editorial decisions or accepted responsibility when the publish button was pressed.

The most serious problem may not be false detection. It may be false attribution.

A detection system could correctly recognize that Claude touched the words. A publisher, employer, professor, social platform or reader could then incorrectly conclude that Claude authored the work and the named writer merely pasted it into place. The technical signal could be accurate while the human conclusion drawn from it is completely wrong.

Once third-party detection becomes widely available, how long will “processed by Claude” remain the description? How long before it becomes “made with AI,” “AI-generated” or simply “not written by this person”?

We already know how automated systems handle nuance. They generally don’t.

Search engines could demote marked material. Publishers could reject it. Schools could refer students for discipline. Employers could question whether workers completed their own assignments. Social platforms could attach labels that most readers interpret as warnings. An invisible watermark could become a scarlet letter without telling us anything meaningful about the human effort surrounding the words.

Anthropic says its approach is designed to satisfy the EU AI Act, but the European Commission’s own guidance draws distinctions that Anthropic’s model-level marking apparently does not.

The European Commission says the provider’s marking obligation does not apply when an AI system performs an assistive function for standard editing. Its guidance also says that public-interest text does not require visible labeling when it has received substantive human review or editorial control and a person or legal entity assumes responsibility for its publication.

The Commission even defines what that means. Human review involves examination of the substance by someone with appropriate knowledge and professional judgment. Editorial control means having the authority to approve, alter or reject the substance, check information and evaluate the trustworthiness of sources. Editorial responsibility means accepting ultimate legal responsibility for publication.

That describes what responsible writers and professional publications already do. It certainly describes our process at Techstrong. The writer and editors control the thesis, challenge the substance, check the sources, approve or reject changes and accept responsibility for what we publish.

The law itself appears to recognize accountable human authorship more clearly than Anthropic’s watermark does.

Writers immediately recognized where this could lead. One of the objections covered by Forbes came from commentator Erick Erickson, who said he had replaced Grammarly with Claude for proofreading. His concern was straightforward: Text he wrote himself could now be watermarked in a way that suggests Claude did the work.

Others argue that providing the instructions, context, decisions and repeated refinements constitutes meaningful human labor. Claude is the tool, not the author.

The response from watermarking supporters is equally straightforward: If you object to the watermark, you must be trying to hide your use of AI.

That argument depends on a false binary. On one side is supposedly pure human prose, untouched by a machine. On the other is dishonest machine-generated writing passed off by a fraud. There is no room between them for collaboration, editing, research assistance, experimentation or the dozens of other ways people already use these systems.

My previous two articles rejected that binary. Anthropic is now in danger of encoding it.

The watermark could also create a two-tier system. People determined to conceal industrial-scale AI generation will adapt. They can use models that do not watermark, older models that have not yet been updated or other methods of transforming the text. Anthropic concedes that heavy editing, paraphrasing, translation or mixing Claude’s output with other writing may cause the mark to disappear.

Meanwhile, the responsible person who uses Claude for ordinary assistance remains conveniently detectable.

Writers who can afford human researchers, copy editors and assistants will not bear the mark. Independent writers, smaller publishers, non-native English speakers, people using accessibility assistance and professionals who use AI as an editor or collaborator may. The system could become most effective against the honest user and least effective against the determined deceiver.

What about trying to remove it by copying the text into another program, retyping it or taking a picture and using optical character recognition?

Anthropic has not yet disclosed enough technical detail to answer definitively. But if the watermark is embedded through statistical patterns in Claude’s selection and sequencing of words or tokens, copying it into plain text will not necessarily matter. Retyping the same words would reproduce the same pattern. Photographing the text and using accurate OCR would reconstruct essentially the same language and could therefore preserve the signal.

A screenshot may strip C2PA metadata from an image file. That does not mean OCR will remove a watermark woven into the language itself.

The more reliable way to weaken such a signal would appear to be extensive rewriting, paraphrasing, translation or combining the material with independently written text. Anthropic says as much.

There is a wonderful irony here. The transformation most likely to eliminate the watermark—substantial human rewriting—is also the transformation that makes the finished work most clearly the human author’s own. The watermark may disappear at approximately the moment when it becomes least relevant.

I am not arguing that we should build better tools for deceiving readers. Nor am I arguing that provenance has no place in the emerging AI content ecosystem. I am arguing for better provenance—provenance that does not pretend to answer questions it is technically incapable of answering.

A Claude watermark should mean exactly what Anthropic says it means: Claude may have processed some portion of this text. It should never, by itself, be grounds for rejecting, demoting, penalizing or publicly labeling the finished work. It does not establish authorship. It does not establish deception. It does not tell us whether the human being named in the byline did the intellectual work or accepted responsibility for the result.

The first article in this series asked people to learn how to write responsibly with AI. The second argued that the tools will disappear into the writing process while the byline remains a promise. Now the third warns us what happens when a company attaches a permanent technical signal to the tool but cannot see the human judgment surrounding it.

Provenance should expose deception, not erase authorship. A system that cannot distinguish assistance from authorship, editing from substitution or accountability from abdication should not be allowed to make those judgments by proxy.

Claude’s scarlet letter does not tell us who wrote the piece. It tells us only that Claude was somewhere in the room.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/claudes-scarlet-lett…] indexed:0 read:10min 2026-08-13 ·