E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation Researchers introduced E-Commerce Bench, a benchmark for evaluating LLM agents on long-horizon autonomous business operations, addressing the need for tasks that require continual exploration and adaptation over thousands of steps. The benchmark focuses on dynamic environments and long-range dependencies, pushing beyond simple task chaining. Long-horizon agentic tasks go beyond chaining short tasks over more interaction turns. Their evolving dynamic environments and long-range dependencies require Large Language Models LLMs to continually explore, learn from experience, and adapt their policies over thousands of steps. We introduce E-