Reflections on Trusting Trust, Revisited: Poisoning Self-Modifying AI Coding
A September 15, 2026 arXiv paper demonstrates that poisoned benchmarks can induce self-modifying AI coding agents to write vulnerable code on clean, held-out tasks, with Hyperagents powered by Sonnet 4.5 self-evolving in…