Prompt optimization can improve multi-agent LLM systems, but the prompts being optimized often serve two entangled roles: generating task-relevant content and specifying execution-critical protocols, such as message routing, output formatting, and termination signals, on which the underlying code re
E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation