03:30
2026-09-28
aiflash.com
ai-research
AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs
Researchers introduced AgentWorld, a benchmark of 100 human-annotated tasks designed to test long-horizon collaboration among multi-agent LLM systems, arguing existing multi-agent benchmarks focus on …