cd /news/artificial-intelligence/learning-to-follow-in-context-waterm… · home topics artificial-intelligence article
[ARTICLE · art-117457] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Learning to Follow In-Context Watermark Instructions via Self-Distillation

A new benchmark, ICWBench, shows that none of 14 frontier proprietary and open-source LLMs can reliably follow in-context watermarking instructions while maintaining answer quality. Researchers propose a self-distillation and reinforcement learning method that raises average TPR@1%FPR from 0.100 to 0.974 on Qwen3-14B and from 0.337 to 0.968 on GPT-OSS-20B, preserving response quality.

read1 min views2 publishedSep 1, 2026

arXiv:2608.29030v1 Announce Type: new Abstract: In-context watermarking (ICW) prepends an instruction to a query asking the model to embed a statistically detectable signal in its response. It thus equips LLMs with a watermarking interface that third parties can invoke without access to model internals. Its reliability hinges on the LLM following the instruction without degrading answer quality, yet how well current LLMs do so has not been measured. We introduce $\mathsf{ICWBench}$, a benchmark of three verifiable ICW instruction families, each scored on both detectability and answer quality. Evaluating 14 frontier proprietary and open-source LLMs, we find that none of the evaluated LLMs achieves both objectives across all three families. To address this, we propose a self-contained two-stage training method, requiring no distillation from a stronger model, no manual annotation, and no pre-existing ICW IF ability. The first stage, self-distillation with logits perturbation (SDLP), uses the same base LLM as both teacher and student: an instruction-equivalent decoding-time logits perturbation makes the teacher follow the ICW instruction, and the student is trained to match the teacher's output distribution. The second stage applies reinforcement learning with the automatic verifier as the reward. Applied to Qwen3-14B and GPT-OSS-20B, our method raises average TPR@$1%$FPR across three ICW instructions from $0.100$ to $0.974$ and from $0.337$ to $0.968$, respectively, while maintaining high response quality under both perplexity evaluation and LLM-as-a-Judge.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @icwbench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/learning-to-follow-i…] indexed:0 read:1min 2026-09-01 ·