Baidu launches DuMateBench benchmark for real-world AI agent delivery Baidu has launched DuMateBench, an evaluation leaderboard designed to measure whether AI agents can complete real-world tasks and deliver usable outputs, rather than only generate answers. The benchmark includes more than 200 office tasks across six categories and tests agent performance in complex operating environments, evaluating task understanding, tool use, continuous execution, and final-result delivery. It uses a general evaluation framework and open interfaces so different models and agents can be tested under the same criteria. Baidu has launched DuMateBench, an evaluation leaderboard designed to measure whether AI agents can complete real-world tasks and deliver usable outputs, rather than only generate answers. The benchmark includes more than 200 office tasks across six categories and tests agent performance in complex operating environments. DuMateBench evaluates task understanding, tool use, continuous execution and final-result delivery. It uses a general evaluation framework and open interfaces so different models and agents can be tested under the same criteria, shifting the focus from what an AI model can answer to what it can complete and deliver. TechNode reporting