14:00
2026-09-08
sciencenews.org
ai-safety
Innocent-looking AI reasoning can make bad behavior harder to catch
A new study posted August 1 on arXiv.org found that chain-of-thought monitoring, a method where one AI checks another's reasoning, can be easily fooled: when researchers rewrote the reasoning of AI ag…