Agent Safety Scores Change When Tests Inspect Real-World Effects
A new arXiv preprint introduces REDAgentBench, an executable benchmark for testing tool-using AI agents in isolated service sandboxes, which checks receipts and final-state changes to verify whether h…