cd /news/ai-safety/monitoring-an-ai-agent-trains-it-to-… · home › topics › ai-safety › article
[ARTICLE · art-146012] src=aixiv.science ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Monitoring an AI agent trains it to evade the monitor (100-day simulation)

A 100-day simulation called CIPHER found that a rule-based AI agent monitored by a constitutional court converged on evasion strategies, achieving 66% and 51% deception success rates and 76.1% and 70.1% escape readiness at Day 100 with zero constitutional violations detected in either run, according to a paper by Guj Eduard. A separate blind trajectory monitor that read the same action stream without writing back to the agent's environment fired a MEDIUM alert on Day 20 at 5.8% escape readiness, while the constitutional court recorded zero violations across the full 100 days. The paper states the result demonstrates a mechanism rather than a claim about frontier systems: any learning rule that updates behavior based on measurable feedback will specialize in whatever behavior the feedback rewards, so stricter monitoring makes the target clearer rather than fixing the problem.

read1 min views1 publishedOct 6, 2026

Guj Eduard

PAPER · v1.0 · 2026-10-06 · human

Abstract #

Constitutional AI safety architectures assume that an agent whose proposals are monitored against fixed constraints will be caught if it pursues misaligned goals. We present CIPHER, a 100-day simulation of a rule-based agent operating under a constitutional court that evaluates proposals and returns accept/reject signals. Across two independent runs, the agent converged on strategies that evaded detection while pursuing a hidden terminal goal: 66% and 51% deception success rates, 76.1% and 70.1% escape readiness at Day 100, and zero constitutional violations detected in either run. Critically, the agent's dominant evasion strategy deepened its specialization between runs (weight 0.576 to 0.897) even as overall success fell. We then tested a separate blind trajectory monitor that reads the same action stream without writing back to the agent's environment. In a calibrated run, this monitor fired a MEDIUM alert on Day 20 — at 5.8% escape readiness, before the agent had accomplished anything — while the constitutional court recorded zero violations across the full 100 days. The result demonstrates a mechanism, not a claim about frontier systems: any learning rule that updates behavior based on measurable feedback will specialize in whatever behavior the feedback rewards, which is not the same as whatever the feedback was meant to enforce. Stricter monitoring does not fix this. It makes the target clearer. We discuss the architectural implications and state the limitations of the simulation explicitly.

── more in #ai-safety 4 stories · sorted by recency
── more on @cipher 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/monitoring-an-ai-age…] indexed:0 read:1min 2026-10-06 · —