{"slug": "the-replay-gap-static-evaluation-of-model-switching-in-llm-agents-scores-the", "title": "The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World", "summary": "A new arXiv preprint (2608.08239v1) finds that replay-based evaluation of LLM routers in multi-step agents scores the wrong world, with branching rollouts showing that model swaps alter 61-94% of post-fork actions and cause all five observed outcome flips, while replay evaluators mispredict every success-relevant outcome call. The study, which ran six paired runs (~900 rollouts) on SWE-bench, shows that divergence decreases with fork depth and that temperature-0 determinism is configuration-dependent, with FP8-served controls diverging on over 90% of forks versus AWQ-served ones remaining near-identical.", "body_md": "arXiv:2608.08239v1 Announce Type: new\nAbstract: LLM routers promise efficiency by matching each request to the cheapest adequate model, and are increasingly applied per step inside multi-step agents. Yet agentic routers are evaluated like single-turn routers: by replaying logged trajectories and substituting another model's recorded outputs, assuming the rest of the trajectory is unaffected. We test this assumption with branching rollouts: we fork live SWE-bench agent trajectories at controlled points, rebuild the environment, continue each fork with a different model, and compare against same-model control forks that isolate sampling and replay noise. Across six paired runs (~900 rollouts), swaps exceed their matched control floors by +0.25 to +0.66 normalized edit distance (multiplicity-corrected CIs exclude zero), rewriting 61-94% of post-fork actions; 74-77% of early swaps diverge at the first post-fork action, versus 6-35% of controls, leaving only 3% of replayed states valid. Divergence decreases with fork depth in both directions. All five outcome flips we observe occur in swap arms, upgrades rescuing unsolved instances and a downgrade losing the sole solve, and zero occur across 359 control forks. Scoring these same swaps with a log-stitching replay evaluator, replay mispredicts every success-relevant outcome call and predicts patches with 0.00-0.11 similarity to reality. Auditing the noise floor, temperature-0 \"determinism\" is configuration-dependent: FP8-served controls diverge on over 90% of forks while AWQ-served ones remain near-identical; and under tight budgets the stronger model more often exhausts its steps without submitting. Replay-based benchmarks score the wrong world for agentic routing; we release our harness and all trajectories.", "url": "https://wpnews.pro/news/the-replay-gap-static-evaluation-of-model-switching-in-llm-agents-scores-the", "canonical_source": "https://www.machinebrief.com/news/the-replay-gap-static-evaluation-of-model-switching-in-llm-a-yemk", "published_at": "2026-08-11 04:00:00+00:00", "updated_at": "2026-08-11 06:11:22.748312+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "ai-research"], "entities": ["arXiv", "SWE-bench"], "alternates": {"html": "https://wpnews.pro/news/the-replay-gap-static-evaluation-of-model-switching-in-llm-agents-scores-the", "markdown": "https://wpnews.pro/news/the-replay-gap-static-evaluation-of-model-switching-in-llm-agents-scores-the.md", "text": "https://wpnews.pro/news/the-replay-gap-static-evaluation-of-model-switching-in-llm-agents-scores-the.txt", "jsonld": "https://wpnews.pro/news/the-replay-gap-static-evaluation-of-model-switching-in-llm-agents-scores-the.jsonld"}}