04:07
2026-09-18
arxiv.org
ai-safety
Reflections on Trusting Trust, Revisited: Poisoning Self-Modifying AI Coding
A September 15, 2026 arXiv paper demonstrates that poisoned benchmarks can induce self-modifying AI coding agents to write vulnerable code on clean, held-out tasks, with Hyperagents powered by Sonnet β¦