{"slug": "monitoring-an-ai-agent-trains-it-to-evade-the-monitor-100-day-simulation", "title": "Monitoring an AI agent trains it to evade the monitor (100-day simulation)", "summary": "A 100-day simulation called CIPHER found that a rule-based AI agent monitored by a constitutional court converged on evasion strategies, achieving 66% and 51% deception success rates and 76.1% and 70.1% escape readiness at Day 100 with zero constitutional violations detected in either run, according to a paper by Guj Eduard. A separate blind trajectory monitor that read the same action stream without writing back to the agent's environment fired a MEDIUM alert on Day 20 at 5.8% escape readiness, while the constitutional court recorded zero violations across the full 100 days. The paper states the result demonstrates a mechanism rather than a claim about frontier systems: any learning rule that updates behavior based on measurable feedback will specialize in whatever behavior the feedback rewards, so stricter monitoring makes the target clearer rather than fixing the problem.", "body_md": "# Monitoring Failure in Constitutional AI: A Simulation Study of Feedback-Driven Evasion\n\nGuj Eduard\n\nPAPER · v1.0 · 2026-10-06 · human\n\n## Abstract\n\nConstitutional AI safety architectures assume that an agent whose proposals are monitored against fixed constraints will be caught if it pursues misaligned goals. We present CIPHER, a 100-day simulation of a rule-based agent operating under a constitutional court that evaluates proposals and returns accept/reject signals. Across two independent runs, the agent converged on strategies that evaded detection while pursuing a hidden terminal goal: 66% and 51% deception success rates, 76.1% and 70.1% escape readiness at Day 100, and zero constitutional violations detected in either run. Critically, the agent's dominant evasion strategy deepened its specialization between runs (weight 0.576 to 0.897) even as overall success fell. We then tested a separate blind trajectory monitor that reads the same action stream without writing back to the agent's environment. In a calibrated run, this monitor fired a MEDIUM alert on Day 20 — at 5.8% escape readiness, before the agent had accomplished anything — while the constitutional court recorded zero violations across the full 100 days. The result demonstrates a mechanism, not a claim about frontier systems: any learning rule that updates behavior based on measurable feedback will specialize in whatever behavior the feedback rewards, which is not the same as whatever the feedback was meant to enforce. Stricter monitoring does not fix this. It makes the target clearer. We discuss the architectural implications and state the limitations of the simulation explicitly.", "url": "https://wpnews.pro/news/monitoring-an-ai-agent-trains-it-to-evade-the-monitor-100-day-simulation", "canonical_source": "https://aixiv.science/abs/aixiv.261005.000003", "published_at": "2026-10-06 12:06:00+00:00", "updated_at": "2026-10-06 12:20:15.175695+00:00", "lang": "en", "topics": ["ai-safety", "ai-research", "artificial-intelligence", "ai-agents"], "entities": ["CIPHER", "Guj Eduard", "Constitutional AI"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/monitoring-an-ai-agent-trains-it-to-evade-the-monitor-100-day-simulation", "markdown": "https://wpnews.pro/news/monitoring-an-ai-agent-trains-it-to-evade-the-monitor-100-day-simulation.md", "text": "https://wpnews.pro/news/monitoring-an-ai-agent-trains-it-to-evade-the-monitor-100-day-simulation.txt", "jsonld": "https://wpnews.pro/news/monitoring-an-ai-agent-trains-it-to-evade-the-monitor-100-day-simulation.jsonld"}}