{"slug": "agent-eval-costs-stop-runs-you-can-already-predict", "title": "Agent Eval Costs: Stop Runs You Can Already Predict", "summary": "A paper posted September 2, 2026, \"EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction\" (Shi, Sun, Dong, Wan, Lo, Gu), proposes halting agent benchmark runs whose outcome is already predictable, using two gradient-boosted classifiers that score partial trajectories at every step. The authors tabulate full-pass costs priced in June 2026 of roughly $715 to $935 on SWE-bench Verified (500 tasks), $122 to $1,305 on GAIA (165 tasks), and $641 to $2,270 on SWE-bench Multimodal (517 tasks), noting per-pass cost varies by close to an order of magnitude across frontier models on the same benchmark. The approach targets per-task step cost rather than task count, which prior benchmark distillation work reduces.", "body_md": "# Agent Eval Costs: Stop Runs You Can Already Predict\n\nFull-pass agent benchmarks now cost hundreds to thousands of dollars per run. Early outcome prediction halts runs whose result is already evident, and the architecture around it decides whether your rankings survive.\n\n## Table of Contents\n\nMost teams that ship agents discover the same thing around the third month: the eval suite that was supposed to make iteration safe has become the thing that makes iteration slow. A single pass over a realistic agentic benchmark, with a frontier model driving a multi-step scaffold, is no longer a coffee-break job. You queue it, you wait, you look at the invoice, and then you start quietly skipping it for changes that “shouldn’t matter.” That is exactly how regressions reach production.\n\nThe decision in front of you is not whether to evaluate agents. It is how to keep evaluation frequent enough to be a real gate when each full run costs real money and each trajectory burns dozens of model calls. The usual answer, shrinking the task set, only solves half the problem. Every task you keep still runs to completion, including the many that were visibly doomed or visibly solved long before the final step.\n\nA paper posted on September 2, 2026, “EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction” (Shi, Sun, Dong, Wan, Lo, Gu), makes that second half explicit and gives us a concrete, measurable pattern to reason about. It is worth treating as evidence for an architectural idea rather than as a product to adopt, because the idea is what transfers to your stack.\n\n## Why Eval Cost Is Now an Architecture Problem\n\nThe paper’s framing is blunt: one pass of a frontier model over an agentic benchmark can cost hundreds to thousands of dollars, and that cost recurs on every iteration of the development cycle. The authors tabulate a full pass with a single open scaffold across three frontier models, priced in June 2026. On SWE-bench Verified (500 tasks) the range is roughly $715 to $935 depending on the model. On GAIA (165 tasks) one model costs $1,305 while another costs $122. On SWE-bench Multimodal (517 tasks) the range runs from $641 to $2,270.\n\nTwo things follow from those numbers. First, the cost per evaluation pass varies by close to an order of magnitude across models on the same benchmark, so your eval budget is coupled to your model choice in a way most capacity plans ignore. Second, if you run the pass per candidate change, per model upgrade, per prompt revision, and per scaffold tweak, the multiplication happens in your CI, not in a spreadsheet anyone reviews.\n\nPrior work on benchmark distillation attacks the task count. Keep fewer, more informative tasks. That helps, but it leaves the per-task cost unchanged. EarlyEval attacks the other axis: within each task, stop paying for steps after the outcome is already knowable.\n\n## The Pattern: Predict, Then Halt\n\nThe mechanism is simple enough to reimplement. A pair of gradient-boosted classifiers, one predicting eventual success and one predicting eventual failure, is applied to the partial trajectory at every step. If either classifier’s calibrated confidence crosses its threshold, the run is halted and the predicted outcome is recorded as the score. Otherwise the run continues to completion and records its true outcome.\n\nThe features are worth reading closely, because they tell you what “evident from intermediate behavior” actually means. The paper groups them into behavioral signals (cumulative counts of actions, tool calls, edits, and test runs; timing of first edit, first test, first submission; stalling and repeated-action indicators; streaks with no edits; submitting without testing; error and test status), textual signals (TF-IDF over the task prompt, the action history, and the environment feedback, reduced with truncated SVD), and reference signals (properties of the gold solution and overlap between the partial trajectory and that solution). The authors say the approach leans primarily on the reference-free behavioral signals, which matters a great deal for enterprise use, since most of your internal tasks do not have a gold patch.\n\nTraining data is historical: more than 21,000 trajectories from dozens of distinct agents across SWE-bench Verified, TerminalBench and Toolathlon. The evaluation protocol is leave-one-agent-out, meaning the predictor is tested on an agent it never saw. That is the honest test for our purposes, because the thing you evaluate tomorrow is, by definition, a configuration you have not evaluated before.\n\nThe headline results across the three benchmarks: 13% to 26% fewer agent steps, up to 44.1% fewer input tokens and up to 29.4% fewer output tokens, prediction accuracy of 89% to 97%, and per-agent resolve rates shifting by only one to two percentage points on average. On SWE-bench Verified specifically, about 35% of runs were halted at roughly 95% prediction accuracy, and the ranking of the 16 agents held at a Spearman correlation of 0.991, with only three adjacent agents swapping by a single rank. On the other two benchmarks the paper reports rankings preserved at a correlation of at least 0.959.\n\nRead that carefully, though. The savings are meaningful, not magical. A 13% to 26% reduction in steps is a solid cost line item, not a 10x. And the authors themselves note that the threshold is a knob trading how early you stop against how reliably the prediction matches the truth.\n\n## Common Failure Modes\n\nThe pattern is sound, but it introduces failure modes that a plain full-pass eval does not have, and you should design for them up front rather than discover them in a postmortem.\n\nThe first is silent score corruption. A wrong early-stop prediction is recorded as the task score. That is why the paper reports per-agent resolve-rate deviation as a primary metric: errors do not crash anything, they just nudge your number. A one to two point average shift is fine for ranking candidates, and potentially not fine if you are comparing against a contractual or regulatory threshold that sits at a specific number.\n\nThe second is distribution shift in the predictor. The classifiers are trained per benchmark on historical trajectories. If your agent architecture changes in a way that alters its behavioral signature, for example a new planner that edits late but correctly, the “no edits so far means probably failing” signal gets miscalibrated exactly when you most need the eval to be trustworthy. The leave-one-agent-out results suggest reasonable robustness to new agents on these public benchmarks, but they are not a guarantee for your internal task distribution.\n\nThe third is asymmetric usefulness. The paper’s own breakdown indicates the failure predictor carries much of the savings on some benchmarks, while success prediction is weaker there. Practically, that means “stop the hopeless runs” is the dependable win, and “declare victory early” is the riskier half. If you can only trust one direction, trust that one.\n\nThe fourth is threshold drift. Thresholds are fixed in advance. If they are set once and forgotten while the agent, the model, and the task mix all move, the effective precision of your early stops decays without any alert firing.\n\n## Architecture Impact\n\n**What changes in system design?**\nThe eval runner stops being a dumb loop that executes every trajectory to completion and becomes a streaming consumer of its own trajectories. You need a per-step feature extractor that can run online, a predictor service (or in-process model) queried each step, and a control path that can terminate a run and record a synthetic outcome with provenance. Historical trajectories become a first-class training asset, which means the trace store you built for observability now feeds your eval economics.\n\n**What new failure mode appears?**\nSilent score bias. Early-stopped outcomes are predictions, not measurements, so an uncalibrated or stale predictor shifts your benchmark numbers without any run failing. The second-order risk is a feedback loop: if you later train the predictor on scores that were themselves early-stop predictions, errors compound. Label provenance (measured versus predicted) has to be stored per run.\n\n**What enterprise teams should evaluate:**\n\n- Platform / MLOps: whether your trace store retains full per-step trajectories with outcome labels and scaffold and model identifiers, since that is the training set for any predictor.\n- Eval / QA owners: a held-out audit sample that always runs to completion, to continuously measure early-stop precision and resolve-rate deviation against ground truth.\n- Risk and compliance: whether any eval number is used as a formal threshold or attestation, in which case predicted outcomes should not count toward it without disclosure.\n- Finance / FinOps: attribution of eval spend per pipeline and per model, since the cost per pass differs several-fold across models on the same benchmark.\n\n**Cost / latency / governance / reliability implications:**\nCost: the paper reports 13% to 26% fewer steps and up to 44.1% fewer input tokens, which on a $700 to $900 full pass is a saving in the low hundreds of dollars per run, multiplied by every iteration you run. Latency: halted runs free your eval queue sooner, which improves iteration time, though you add a small per-step prediction overhead that the authors describe as negligible. Governance: predicted scores need provenance labels so that audit trails distinguish measured from inferred outcomes. Reliability: a one to two percentage point average shift in resolve rate is acceptable for ranking, and worth testing before you rely on it for pass/fail gating.\n\n## Decision Framework\n\nNot every team should build this. The pattern pays off when three conditions hold together: your full-pass eval is expensive enough that you are already rationing it, you have a meaningful history of labeled trajectories for the benchmark you care about, and your use of the score is comparative (which candidate is better) rather than absolute (does this clear a fixed bar).\n\nIf your suite is small and cheap, skip it and spend the effort on task quality. If your evals are run rarely and cost is not a constraint, the complexity is not justified. If you have plenty of trajectories and your bottleneck is iteration speed, this is a good fit. If you lack history, start by instrumenting and collecting, because the predictor cannot be built without it, and the collection itself is valuable for observability.\n\nA useful middle path before any ML is rule-based early termination on behaviors that are unambiguous in your own system: a run that has exceeded a step budget with no state change, a run looping on the same failing action, a run that hit a hard error class. These do not need a classifier, and they capture part of the failure-side savings at near-zero risk.\n\n## Implementation Guide\n\nStart with data, not models. Make sure every eval run emits a per-step trace with enough structure to reconstruct cumulative counts of actions, tool calls, edits, tests, and errors, plus the final outcome and the identity of the scaffold and base model. If your trace schema is already aligned with OpenTelemetry GenAI conventions, you are most of the way there. Keep these traces for every eval you run, including the expensive ones, because they are the only way you will ever be able to cheapen the next one. A few thousand labeled trajectories per benchmark across several agent variants is a realistic starting corpus.\n\nBuild the smallest useful version next: a failure-side predictor only, with a high confidence threshold, running in shadow mode. Shadow mode means it scores every step and records what it would have halted, but the run continues to completion. After a few hundred runs you will know empirically how often the “would halt as failing” call was correct, and you can set a threshold from measured precision instead of intuition. Only after that, enable actual halting, and keep a fixed fraction of runs, say one in ten, exempt from early stopping as a permanent audit sample. That audit sample is what turns a one-time calibration into a continuously verified control.\n\nWhat to avoid: do not enable success-side early stopping first, because declaring a pass early is the riskier error and it inflates your headline metric. Do not use the predicted scores for any absolute threshold or attestation without labeling them as predicted. And do not train the predictor on scores from early-stopped runs, which will compound bias; train only on runs that completed.\n\nHow to know it is working: track three numbers on the audit sample, namely early-stop precision, the average absolute shift in resolve rate versus full-pass truth, and rank correlation between early-stopped and full-pass leaderboards for the same candidate set. The paper reports correlations above 0.95 and average shifts of one to two points on public benchmarks, which is a reasonable target to hold yourself to, but your own numbers on your own tasks are the only ones that matter. Also track saved spend per week, so the program justifies itself.\n\nThe 6 to 12 month maturity path looks like this. In the first quarter you have instrumented traces and a shadow predictor. By the second you are halting hopeless runs in CI with an audit sample and have per-pipeline eval cost dashboards. By the third or fourth quarter you retrain predictors on a schedule keyed to scaffold and model changes, route cheap predicted evals to the inner development loop and reserve full passes for release gates, and treat eval spend as a budgeted, attributed line item with the same discipline as inference spend. Teams that get here can afford to run evals on every meaningful change, which is the whole point.\n\n## Sources\n\nEnterprise AI Architecture\n\n## Want more enterprise AI architecture breakdowns?\n\nSubscribe to SuperML.", "url": "https://wpnews.pro/news/agent-eval-costs-stop-runs-you-can-already-predict", "canonical_source": "https://superml.dev/agent-eval-cost-early-stopping-2026", "published_at": "2026-09-30 20:19:10.466855+00:00", "updated_at": "2026-09-30 20:19:12.615304+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-research", "mlops"], "entities": ["EarlyEval", "Shi", "Sun", "Dong", "Wan", "Lo", "Gu", "SWE-bench Verified"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/agent-eval-costs-stop-runs-you-can-already-predict", "markdown": "https://wpnews.pro/news/agent-eval-costs-stop-runs-you-can-already-predict.md", "text": "https://wpnews.pro/news/agent-eval-costs-stop-runs-you-can-already-predict.txt", "jsonld": "https://wpnews.pro/news/agent-eval-costs-stop-runs-you-can-already-predict.jsonld"}}