{"slug": "dumatebench-evaluating-autonomous-agents-in-complex-real-world-workflows", "title": "DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows", "summary": "Researchers introduced DuMateBench, a benchmark of 200 real-session tasks reconstructed from anonymized user sessions on a production agent platform, to evaluate autonomous agents in complex workflows with real-world environmental complexity. Testing five agent frameworks paired with four LLMs revealed substantial gaps in strict task completion, with performance jointly shaped by both the LLM and the framework.", "body_md": "arXiv:2608.26546v1 Announce Type: new\nAbstract: Autonomous agents are increasingly adopted to complete complex, multi-tool workflows in real-world settings. However, existing benchmarks typically separate tasks by application or capability and evaluate agents in environments that are cleaner and more stable than those encountered in practice. We introduce DuMateBench, a real-session benchmark reconstructed from anonymized and privacy-screened user sessions collected from a large-scale production agent platform. Each task preserves the relevant pre-solution interaction history, persistent configurations, and workspace state, and is then validated through human verification. The resulting benchmark comprises 200 tasks spanning 8 broad scenarios and 17 fine-grained capability categories, with most tasks requiring multiple capability coordination. We execute these tasks in isolated Docker containers injected with three forms of real-world environmental complexity: Insufficient, Unstable, and Noisy, and assess performance using a hybrid deterministic and LLM-as-Judge evaluation protocol. Experiments across five representative autonomous-agent frameworks paired with four state-of-the-art LLMs reveal substantial gaps in strict task completion. Complementary robustness, efficiency, and diagnostic analyses further show that performance under environmental perturbations is jointly shaped by the capabilities of the LLM and the surrounding agent framework. The code and data are publicly available at https://dumatebench.com/.", "url": "https://wpnews.pro/news/dumatebench-evaluating-autonomous-agents-in-complex-real-world-workflows", "canonical_source": "https://www.machinebrief.com/news/dumatebench-evaluating-autonomous-agents-in-complex-real-wor-i40c", "published_at": "2026-08-28 04:00:00+00:00", "updated_at": "2026-08-28 05:18:49.010973+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-agents", "ai-research", "ai-infrastructure"], "entities": ["DuMateBench", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/dumatebench-evaluating-autonomous-agents-in-complex-real-world-workflows", "markdown": "https://wpnews.pro/news/dumatebench-evaluating-autonomous-agents-in-complex-real-world-workflows.md", "text": "https://wpnews.pro/news/dumatebench-evaluating-autonomous-agents-in-complex-real-world-workflows.txt", "jsonld": "https://wpnews.pro/news/dumatebench-evaluating-autonomous-agents-in-complex-real-world-workflows.jsonld"}}