14:21
2026-08-05
arxiv.org
artificial-intelligence
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks 2025
Researchers introduced MCP-Bench, a benchmark for evaluating large language models (LLMs) on realistic multi-step tasks requiring tool use, built on the Model Context Protocol (MCP) and connecting to …