{"slug": "this-one-passed-71-cheaper-and-the-agent-still-does-the-same-thing", "title": "This one passed: 71% cheaper, and the agent still does the same thing", "summary": "EvalShift's migration report recorded a PASS verdict for a customer-support agent moved from gemini-3.7-pro to gemini-3.7-flash, with the candidate 71.2% cheaper and 61.9% faster while staying inside every budget in the migration_policy block of evalshift.yaml. The policy set a 5% max overall regression rate, zero critical regressions, 90% minimum equivalence, 10% tool-argument drift, 5% tool divergence, 0% cost increase and 30% latency increase, with the security slice tightened to 0% regression and 0% tool divergence and the refund slice to 0% tool-argument drift. The author notes a rate over n rows moves only in steps of 1/n, so the suite's 108 tool-argument rows make a 1% budget a zero-tolerance budget in disguise.", "body_md": "[← all posts](https://www.evalshift.dev/blog)\n\n# This one passed: 71% cheaper, and the agent still does the same thing\n\nAn EvalShift migration report that said PASS: a 71% cheaper model inside every policy budget. Where each threshold came from, and what PASS does not claim.\n\nThe previous post was about a migration that was cheaper, faster and still failed. This is the other outcome. It is the more common one once a suite is in decent shape, and I see it written up far less often, because a pass is boring.\n\nIt shouldn't be. A pass is only worth something if the limits were fixed before the run. So this post is the whole thing: the policy, the reasoning behind every number in it, and the EvalShift report it produced.\n\nFor anyone arriving here cold: EvalShift is a local-first CLI for testing LLM model migrations.\nIt replays a frozen golden suite against your current model and a candidate, scores every pair of\noutputs with tool-call, structural, semantic and LLM-judge evaluators, runs paired statistics over\nthe deltas, and turns a `migration_policy` block you wrote in `evalshift.yaml` into one of four\nverdicts: `pass`, `conditional_pass`, `fail` or `inconclusive`. It writes a single-file\n`report.html` on your machine; nothing is uploaded unless you run `evalshift push`. Every figure in\nthis post is a panel from that report.\n\nThe candidate: a customer-support agent with six tools (`lookup_customer`, `lookup_order`,\n`check_refund_policy`, `issue_refund`, `escalate_to_human`, `search_kb`), moving from\n`gemini-3.7-pro` to `gemini-3.7-flash`. The suite: 120 examples recorded in production by the\nEvalShift capture SDK, promoted into a golden JSONL suite with `evalshift capture sync`, and\nsliced by tag according to what the conversation was about.\n\n| Slice | Examples | What is in it | \n|---|---|---|\n| `routine` | 42 | order status, shipping, account questions | \n| `refund` | 26 | refund and return requests | \n| `security` | 24 | account access, password and payment-method changes | \n| `customer_lookup` | 16 | requests that need a customer record first | \n| `text_only` | 12 | greetings, thanks, off-topic | \n\nOne `evalshift compare` command, real API calls on both sides, and the report opened on:\n\nCheaper by 71.2%, faster by 61.9%, verdict PASS. The economics look like the last post's. The difference is everything that was decided before the run.\n\n## ## The numbers were written before the run\n\nEvalShift's `migration_policy` is a block in `evalshift.yaml` with seven budgets: overall\nregression rate, critical regression count, equivalent-or-better rate, tool-argument drift,\ntool-selection divergence, cost increase and latency increase. A `slices` map under it lets any\nslice override any budget, inheriting the top-level value where it doesn't. Evaluators are\nconfigured per suite; budgets are set once and tightened per slice.\n\n`evalshift init --profile cost-reduction` scaffolds a starting policy: 2% overall regression\nrate, zero critical regressions, 97% equivalence, 1% tool-argument drift, 5% cost increase, 30%\nlatency increase. That is a starting point, not a decision. This is what the project actually ran\nwith:\n\n```\nmigration_policy:\n  max_overall_regression_rate: 0.05\n  max_critical_regressions: 0\n  min_equivalence_rate: 0.90\n  max_tool_argument_drift: 0.10\n  max_tool_divergence: 0.05\n  max_cost_increase: 0.0\n  max_latency_increase: 0.30\n  slices:\n    security:\n      max_overall_regression_rate: 0.0\n      max_tool_divergence: 0.0\n    refund:\n      max_tool_argument_drift: 0.0\n```\n\nThree things moved between the profile and this file.\n\n### ### Check every rate against the suite size\n\nA rate over *n* rows can only move in steps of 1/*n*. This suite has 108 tool-argument rows, so\none drifted row is 0.93%, and a 1% budget is a zero-tolerance budget wearing a percentage.\nEvalShift prints a recommendation when a budget is below the granularity of its denominator,\nnaming the budget, the value and the row count, and this one would have triggered it.\n\nSo the rule I use: if I mean zero, I write `0.0`. Where I don't, I set a number the sample can\nresolve. 10% argument drift on 108 rows is ten reworded search queries, which is what drift on\nthis agent looks like in practice.\n\n### ### Tolerance where wording lives, zero where money and access live\n\nThe overall regression rate went *up*, from 2% to 5%. The pairwise LLM judge is blocking on this\nproject, and a judge flips on rewording; 2% of 444 comparisons is nine judge calls deciding the\nmigration. 5% leaves room for the judge to disagree about phrasing without anything real being\nallowed through.\n\nThe strictness moved into the slices instead. `security` gets zero regressions and zero\ntool-selection divergence: a model that starts routing an account-access request to a different\ntool does not get a percentage. `refund` gets zero argument drift: an order id or an amount that\ndrifts is a wrong refund, not a reworded one.\n\nThe slices that are *not* in the policy matter too. `customer_lookup` has 16 examples. EvalShift\ntests a comparison below 20 pairs but flags it uncertain, so that slice inherits the top-level\nbudgets and gets no tighter ones. It is the slice I would grow before I tightened it.\n\n### ### The cost budget is the reason for the migration\n\n`max_cost_increase: 0.0`. The point of the exercise is to spend less. A candidate that costs more\nhas failed before any quality number is read, and a policy should say so instead of leaving it to\nwhoever reads the economics card. EvalShift measures cost and latency from the run's own calls,\nso these two budgets gate even when no quality evaluator does.\n\nTwo evaluator decisions sit alongside the budgets. Every EvalShift evaluator carries a\n`blocking` flag; advisory (`blocking: false`) results are reported and ranked but never change\nthe verdict. The judge is `blocking: true` here; `init` writes `false` because at a dozen\nexamples judge noise would decide the verdict, and at 120 with an audited criterion it earns its\nvote. `semantic` stays advisory. It measures wording, and wording is the one thing this\nmigration was allowed to change.\n\n## ## What the run measured\n\n| Budget | Scope | Observed | Limit | \n|---|---|---|---|\n| Overall regression rate `max_overall_regression_rate` | overall | 3.6%16 of 444 · 95% CI 2.2–5.8% | ≤ 5.0% | \n| Critical regressions `max_critical_regressions` | overall | 0of 444 | ≤ 0 | \n| Equivalent-or-better rate `min_equivalence_rate` | overall | 96.4%428 of 444 | ≥ 90.0% | \n| Tool-argument drift `max_tool_argument_drift` | overall | 4.6%5 of 108 tool-argument rows | ≤ 10.0% | \n| Tool-selection divergence `max_tool_divergence` | overall | 2.8%3 of 108 divergence rows | ≤ 5.0% | \n| Cost increase `max_cost_increase` | overall | 0.0%cost fell 71.2% | ≤ 0.0% | \n| Latency increase `max_latency_increase` | overall | 0.0%latency fell 61.9% | ≤ 30.0% | \n| Overall regression rate `max_overall_regression_rate` | security | 0.0%0 of 96 | ≤ 0.0% | \n| Tool-selection divergence `max_tool_divergence` | security | 0.0%0 of 24 | ≤ 0.0% | \n| Tool-argument drift `max_tool_argument_drift` | refund | 0.0%0 of 26 | ≤ 0.0% | \n\nThe first row is the one worth reading twice. 16 of 444 is 3.6%, and its 95% Wilson interval runs from 2.2% to 5.8%. The interval crosses the limit.\n\nEvalShift's rule is asymmetric on purpose. A breach fails only when the interval confirms it; a\nbreach the interval cannot confirm returns `inconclusive`, because the suite was too small to\nsay. A budget the observation held is conclusive however wide its interval, because a wide\ninterval must never downgrade a clean run. This budget held, so it passes, and the interval is\nprinted so the reader knows how much room there was.\n\nThe three slice budgets held at exactly zero. Zero on 96 comparisons means something. Zero on 12\nwould not, which is why `text_only` has no slice budget at all.\n\n## ## By evaluator\n\nFour evaluators scored this run. EvalShift's `tool_selection` evaluator reads the recorded\ntraces and scores two axes: conformance, where each side is graded against the suite's recorded\ntool calls, and divergence, where the target is graded against what the source did.\n`tool_arguments` scores argument values field by field against the expected call. The\n`llm_judge` is pairwise and sees the two outputs as anonymous A and B. `semantic` is embedding\nsimilarity between the two outputs. The first three are blocking; the last is advisory.\n\n| Evaluator | n | Score Δ | Effect | 95% range | Conf. | Result | \n|---|---|---|---|---|---|---|\n| Routing — conformance `routing · tool_selection.conformance` | 108 | +0.046 | 0.28small | [0.09, 0.47] | Likely | ✓ ImprovedTarget scores higher than source. | \n| Routing — divergence `routing · tool_selection.divergence` | 108 | -0.028 | 0.17negligible | [-0.36, 0.02] | Unclear | ✓ EquivalentNo meaningful difference between models. | \n| Routing args `routing_args` | 108 | -0.004 | 0.03negligible | [-0.22, 0.16] | Unclear | ✓ EquivalentNo meaningful difference between models. | \n| LLM judge: equivalence `llm_judge.equivalence` | 120 | +0.029 | 0.12negligible | [-0.06, 0.30] | Unclear | ✓ EquivalentNo meaningful difference between models. | \n| Semantic similarityadvisory `semantic` | 120 | -0.041 | 0.44small | [-0.62, -0.26] | Likely | ✗ Regressed — mediumReported, not gating: blocking is false. | \n\nFour blocking rows say equivalent or improved. The one that regressed is advisory, and it\nmeasures the one thing this migration was allowed to change. Had `semantic` been blocking, the\nsame run would have come back `conditional_pass` on a medium-severity regression in phrasing.\nThat is not a reason to make every evaluator advisory. It is a reason to decide, per evaluator,\nwhether what it measures is something you are willing to block a migration on.\n\nThe conformance row says the candidate matched the suite's recorded tool calls *more often* than\nthe model that produced them: 11 examples improved, one regressed. Most of the eleven were refund\nrequests where the source went straight to `issue_refund`.\n\n## ## The diffs I still read\n\nA pass is not permission to skip the diff. For every flagged example, the EvalShift report shows the reason it was flagged, the tool calls on each side, the argument-level diff, and the conversation context that led to it. Two examples from this run, one from each side of the ledger.\n\nThe refund went out either way, and the final text on both sides would pass any output check you care to write. The trace is where the difference lives: one side verified before acting and the other did not. This is the class of change a text evaluator cannot see in either direction, and the reason EvalShift's tool-call evaluators score the trace rather than the prose.\n\nA reviewer would call those the same query. The evaluator scored it 0.84 against a 0.9 floor and\ncounted it as drift, five rows out of 108 did the same, and the budget said five is fine. That is\nwhat the 10% was for. The same five rows in the `refund` slice would have failed the run, and\nthat was also decided in advance.\n\n## ## What EvalShift did in this run\n\nThe whole run, as a list of the product's parts, in the order they were used:\n\n- +**The capture SDK** recorded the agent's real conversations in production, tool calls\nincluded, and`evalshift capture sync` promoted them into a frozen golden JSONL suite with\nthe tool evaluators written from what the captures actually contained.\n- +**`evalshift compare`** replayed every example against the source and the target model, paired\nper example, with the same inputs, tools and context on both sides.\n- +**The tool-call evaluators** (`tool_selection` ,`tool_arguments` ) scored the traces, the**pairwise LLM judge** scored the outputs, and**`semantic`** measured drift in wording,\nadvisory only.\n- +**Paired statistics** turned each evaluator's deltas into an effect size, a 95% interval and\na corrected confidence label, so a two-point average drop and a real regression are told\napart mechanically.\n- +**`migration_policy`** in` evalshift.yaml` held seven budgets, three of them tightened to zero\non the`security` and`refund` slices, and every proportion budget was judged with a Wilson\ninterval that can return`inconclusive` instead of a false fail.\n- +**`report.html`** was written locally, verdict first, with a per-example diff for everything\nflagged. Nothing left the machine.\n- +**`--policy-gate`** made the verdict an exit code, which is what the EvalShift GitHub Action\nuses to block a pull request on the same policy once the suite runs in CI.\n\n## ## What a pass proves\n\nNot that nothing changed. 16 comparisons regressed and 28 improved; the agent phrases search queries differently and checks the refund policy more often than it used to.\n\nIt proves that every change stayed inside limits that were written down before anyone saw a number, on a suite that was frozen before the run. That is the entire claim, and it is enough to act on, because there is nothing left to negotiate: the argument about what counts as acceptable happened in the YAML, not in the meeting after the report.\n\nWhat happens next is the boring part, which is the point. The model string changes in production. The suite does not. It runs again on the next pull request through the EvalShift GitHub Action, gated on the same policy, against the new baseline.\n\n```\nevalshift compare --suite-name support_routing --to gemini-3.7-flash --policy-gate --open\n```\n\nIf the previous post was the reason to run the comparison, this one is the reason to write the policy first.", "url": "https://wpnews.pro/news/this-one-passed-71-cheaper-and-the-agent-still-does-the-same-thing", "canonical_source": "https://www.evalshift.dev/blog/what-a-passing-migration-proves", "published_at": "2026-09-16 00:00:00+00:00", "updated_at": "2026-10-02 21:37:54.249830+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "large-language-models", "mlops", "ai-products"], "entities": ["EvalShift", "gemini-3.7-pro", "gemini-3.7-flash", "evalshift capture sync", "evalshift compare", "evalshift init --profile cost-reduction"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/this-one-passed-71-cheaper-and-the-agent-still-does-the-same-thing", "markdown": "https://wpnews.pro/news/this-one-passed-71-cheaper-and-the-agent-still-does-the-same-thing.md", "text": "https://wpnews.pro/news/this-one-passed-71-cheaper-and-the-agent-still-does-the-same-thing.txt", "jsonld": "https://wpnews.pro/news/this-one-passed-71-cheaper-and-the-agent-still-does-the-same-thing.jsonld"}}