ChatGPT on a phone, May 2023 (illustrative). Image: Jernej Furman / Wikimedia Commons, CC BY 2.0, cropped
Text written by ChatGPT and Codex in the European Union will soon carry a hidden watermark. OpenAI said on Monday that it will add the invisible signal to “eligible” output for users on every plan in the EU over the coming weeks, to meet the EU AI Act’s rule that AI-generated text must be identifiable by machines.
It is a regional move, not a global one. OpenAI said it is “not making text watermarking a global default at launch”, so ChatGPT users outside the EU won’t get it for now. Developers using OpenAI’s API anywhere in the world can switch it on for some models from Monday, but it stays off unless they choose it.
How the watermark works #
The system is called textGrain. Instead of adding anything you can see, it nudges the model’s choice of words in a statistical pattern that a detector can later look for. OpenAI has published a technical report and says it plans to release the technology as open source.
The company claims textGrain “matched or exceeded” the other methods it tested, including Google DeepMind’s SynthID for text, which DeepMind has also adapted for proteins. That comparison is OpenAI’s own and hasn’t been independently checked. It also says the watermark makes no meaningful difference to quality: on benchmarks for its GPT-6 Astra model, scores with and without it were within about a point of each other on most tests.
Easy to weaken #
OpenAI is unusually frank about the limits. With the detector set to wrongly flag about 1% of unwatermarked text, it caught the watermark in about 80% of 200-token passages and about 95% of 400-token passages on a subject like psychology, and “substantially lower” rates for maths, where there are fewer ways to word an answer.
Light editing does real damage. Swapping 10% of the words in a 400-token passage for synonyms cut detection from about 92% to 66%, and swapping a quarter of them cut it to 17%. Text that is short, edited, translated or made with another company’s tools may not be caught at all, and OpenAI warns that “the absence of a detected watermark does not prove human authorship.”
No public detector, for now #
Because of those error rates, OpenAI isn’t releasing the detector to the public. From Monday, approved researchers and “expert organizations” can apply for access, granted case by case under the EU’s code of practice on AI-generated content. The tool only says whether it finds an OpenAI watermark; it doesn’t identify the user or reveal prompts or conversations.
OpenAI also spells out what a positive result can’t tell you: how much a person edited or wrote, who owns the text, who is responsible for it, or whether it is accurate. Its existing checks for images and audio, the openai.com/verify tool and its Content Provenance API, stay open to the public.
Why it matters #
This is the first time the most widely used chatbot will mark its own writing, and Europe’s AI Act is the reason. For students, teachers and employers the practical effect is limited for now: there is no public checker, and OpenAI’s own figures show a quick rewrite can wash much of the signal out. But it sets a template other AI companies serving Europe will be pushed to follow.
Sources: OpenAI; OpenAI textGrain technical report; European Commission.