04:00
2026-09-16
arxiv.org
ai-safety
BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents
Researchers introduced Blindspot, a benchmark for trajectory-level safety calibration of long-horizon tool-using LLM agents, containing 22 attack families and 35 scenarios across seven domains that yi…