Microsoft ThinkingBox Exposes the AI Agent Reliability Gap
Microsoft's new open-source benchmark ThinkingBox reveals a significant reliability gap in AI agents, finding that the best-performing model succeeded on only 25.25% of tasks across 20 consecutive run…