cd /news/artificial-intelligence/dumatebench-evaluating-autonomous-ag… · home topics artificial-intelligence article
[ARTICLE · art-113856] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows

Researchers introduced DuMateBench, a benchmark of 200 real-session tasks reconstructed from anonymized user sessions on a production agent platform, to evaluate autonomous agents in complex workflows with real-world environmental complexity. Testing five agent frameworks paired with four LLMs revealed substantial gaps in strict task completion, with performance jointly shaped by both the LLM and the framework.

read1 min views1 publishedAug 28, 2026

arXiv:2608.26546v1 Announce Type: new Abstract: Autonomous agents are increasingly adopted to complete complex, multi-tool workflows in real-world settings. However, existing benchmarks typically separate tasks by application or capability and evaluate agents in environments that are cleaner and more stable than those encountered in practice. We introduce DuMateBench, a real-session benchmark reconstructed from anonymized and privacy-screened user sessions collected from a large-scale production agent platform. Each task preserves the relevant pre-solution interaction history, persistent configurations, and workspace state, and is then validated through human verification. The resulting benchmark comprises 200 tasks spanning 8 broad scenarios and 17 fine-grained capability categories, with most tasks requiring multiple capability coordination. We execute these tasks in isolated Docker containers injected with three forms of real-world environmental complexity: Insufficient, Unstable, and Noisy, and assess performance using a hybrid deterministic and LLM-as-Judge evaluation protocol. Experiments across five representative autonomous-agent frameworks paired with four state-of-the-art LLMs reveal substantial gaps in strict task completion. Complementary robustness, efficiency, and diagnostic analyses further show that performance under environmental perturbations is jointly shaped by the capabilities of the LLM and the surrounding agent framework. The code and data are publicly available at https://dumatebench.com/.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @dumatebench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/dumatebench-evaluati…] indexed:0 read:1min 2026-08-28 ·