I built a system to catch unreliable AI agents. In my own evaluation, it missed the worst one.
A developer built a framework to identify unreliable AI agents in multi-agent pipelines, drawing on classical Islamic hadith science. In their own evaluation, the grade-recovery loop missed the highes…