Your agent returned 200 OK. Was it actually right?
A developer reports that agentic AI systems are good at logging what an agent did but poor at verifying whether the output was correct. In an experiment, a weak model was correct about 69% of the time…