cd /news/artificial-intelligence/watermarking-text-generation-efficie… · home topics artificial-intelligence article
[ARTICLE · art-103300] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Watermarking Text Generation Efficiently

Enzo Lombardi, in an article published on Towards AI, explains that real AI watermarking works by biasing a model's token sampling rather than hiding characters, and that detection relies on a z-score test using a secret key. The article details the 'green list' method and notes that robustness against paraphrasing and other attacks has a theoretical ceiling.

read2 min views1 publishedAug 19, 2026

Author(s): Enzo Lombardi Originally published on Towards AI. Or Why What You Recently Read About AI Watermarking Is Probably Wrong Most explanations of AI watermarking describe something that does not exist. They talk about hidden Unicode characters smuggled between words, or invisible zero-width spaces, or secret vocabulary the model is forced to use, or a classifier that has learned what machine prose smells like. Some of those are real techniques for other problems. None of them is how watermarking actually works in the systems that ship it, and the confusion matters, because every one of those imagined mechanisms would be defeated by pasting the text into a plain editor. After the introduction, the article explains that real watermarking relies on biasing a model’s token sampling: instead of altering the written words directly, it nudges the “slack” in each next-token choice so the resulting text carries a statistical signature. It walks through the classic “green list” method (hashing the previous token plus a secret key to split the vocabulary each step, then adding a delta to logits for “green” tokens), shows how detection becomes a coin-flip-style z-score test that only needs the key and text, and discusses the fundamental tradeoffs controlled by delta and gamma (too much bias harms quality and markability). It then contrasts this with distortion-free approaches that preserve the distribution while encoding a correlation signal, argues that the scheme used by any given model can’t be reliably inferred from outputs alone due to cryptographic unpredictability, and outlines both defenses and attacks—especially paraphrasing, translation, mixing sources, and tokenization tricks—that can destroy or weaken the watermark. Finally, it frames watermarking as a limited but practical measurement tool for platforms rather than a universal “AI detector,” and emphasizes that robustness against determined attackers has a theoretical ceiling. Read the full blog for free on Medium. Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor. Published via Towards AI

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @enzo lombardi 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/watermarking-text-ge…] indexed:0 read:2min 2026-08-19 ·