UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents Researchers introduced UndoBench, a benchmark spanning 36 base workflows and 36 fault-injected variants, to separate baseline planning competence from operational fault recovery in tool-using AI agents. The benchmark addresses the tendency of widely used agent benchmarks to evaluate only nominal task completion, conflating the two capabilities. Tool-using AI agents are increasingly deployed across enterprise software systems, yet widely used benchmarks primarily evaluate nominal task completion, conflating baseline planning competence with operational fault recovery. We introduce UndoBench, a benchmark spanning 36 base workflows and 36 fau