AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs Researchers introduced AgentWorld, a benchmark of 100 human-annotated tasks designed to test long-horizon collaboration among multi-agent LLM systems, arguing existing multi-agent benchmarks focus on competitive settings, interactions under 20 steps, or aggregate individual performance rather than isolating genuine collaboration. The benchmark targets the gap in measuring collaborative capability of LLM-based agents. Existing multi-agent benchmarks primarily test in competitive settings, short-horizon interactions under 20 steps, or simply aggregate individual performance, failing to isolate and highlight genuine collaboration capabilities of LLM-based agents. We introduce AgentWorld, a benchmark of 100 human-an