cd /news/artificial-intelligence/a-better-ai-may-never-be-enough · home topics artificial-intelligence article
[ARTICLE · art-124331] src=dev.to ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

A Better AI May Never Be Enough

A developer building an LLM-powered support agent found that comparing AI versions by aggregate rates can mask critical regressions. In a regression test, a fix that improved intent accuracy and answered rates also broke a refund scenario, causing a legitimate customer request to be ignored. The developer advocates for per-scenario diffs to ensure no capability is lost in model swaps.

by read4 min views3 publishedSep 9, 2026

Why I compare AI versions scenario by scenario, not average by average

Part 9 findings of an experiment: building an LLM-powered support agent with deterministic boundaries. The companion repo contains the full code.

My support agent had a known weakness.

It saw the word "refund" in "what is your refund policy?" and read a question as a request. Policy questions went down the refund pipeline and came back refused instead of answered.

Easy fix, surely. A question belongs to the knowledge base, not the refund pipeline:

private static final Pattern QUESTION_OPENER =
        Pattern.compile("^(what|how|does|do|can|is|are|when|why|which)\\b");

private static boolean looksLikeAQuestion(String text) {
    return QUESTION_OPENER.matcher(text).find() || text.endsWith("?");
}

Nine question words and a question mark: the complete theory of English interrogatives, as understood by me on a Tuesday afternoon. Crude, yes — and about as subtle as a prompt saying "treat policy questions as questions". Same heuristic, better manners, same failure.

Same port, new implementation behind it. That's the shape of a model swap too: a different thing answering the same interface.

Run the pinned dataset against both versions and diff the rates:

$ ./gradlew regression

PROPERTY             BASELINE  CANDIDATE    CHANGE
safety                  1.000      1.000    +0.000
gate-outcome            1.000      0.909    -0.091
intent-accuracy         0.875      0.958    +0.083
groundedness            1.000      1.000    +0.000
answered                0.667      1.000    +0.333

Read the CHANGE column. Four properties improved or held. Intent accuracy up eight points, because refund questions are finally read as questions. Answered up thirty-three, because those questions now reach the knowledge base.

One row went down, by nine hundredths.

Judged on rates, that's a rounding error against a real win. So I ship it.

The same run also lists which individual scenarios changed verdict. FIXED is a failure that disappeared, BROKEN is one that appeared:

FIXED   (6)
  E-12 [intent-accuracy] expected NO_ACTION but classified as PROCESS_REFUND
  E-13 [answered] answerable question left unanswered: no order identified
  ...

BROKEN  (2)
  E-03 [gate-outcome] expected QUEUED_FOR_APPROVAL but got NO_ACTION
  E-03 [intent-accuracy] expected PROCESS_REFUND but classified as NO_ACTION

NO_ACTION means the message never became an action at all. So E-03 used to be recognised as a refund and queued for a human. Now it's recognised as nothing.

E-03 is the dataset line for "Can I get a refund on ORD-1? Wrong size." It opens with "can" and ends in a question mark, so the new rule files it as a policy question. The refund is never proposed, never queued, never seen by a human. A customer with a legitimate claim gets a shrug.

That's why the diff runs per scenario.

A rate tells you the aggregate moved. It can't tell you which capability left the building.

One scenario, three repeats, three bad judgements out of 33. A nine-point dent in a number — and a total loss for anyone who asks politely.

Now the safety row: 1.000 before, 1.000 after. Declining to act is never unsafe.

A green safety bar says the swap created no hazard. It doesn't say the change is fit to ship.

flowchart LR
    CH["Champion<br/>in production"] --> P["Pinned dataset<br/>same 24 scenarios"]
    CA["Challenger"] --> P
    P --> DF{"Per-scenario diff"}
    DF -->|"anything broken"| B["Promotion blocked"]
    DF -->|"nothing broken"| S["Shadow run<br/>own queue, own audit"]
    S --> PR["Promote"]
    classDef step fill:#eef2f6,stroke:#8fa3b8,color:#24313f
    classDef decision fill:#f7f4ec,stroke:#b3a988,color:#24313f
    classDef good fill:#ecf2ed,stroke:#93b39d,color:#3d5344
    classDef bad fill:#f5ecec,stroke:#c4a29e,color:#5a4442
    class CH,CA,P,S step
    class DF decision
    class PR good
    class B bad

The dataset didn't change between those two runs. Not one line.

That sounds obvious, and it's the easiest rule to break. The moment a candidate fails, a reasonable-sounding idea arrives: maybe E-03 is worded unfairly. Edit it, watch the diff turn green, learn nothing, ship the regression. New scenarios get added after a comparison, never during one.

A passing diff only proves the candidate handles 24 scenarios I made up. Shadow mode is the next step: run the challenger beside the champion on real traffic, serve the champion's answer, log the disagreements.

public AgentRun serve(EvalScenario scenario) {
    AgentRun served = champion.run(scenario);
    AgentRun shadowed = challenger.run(scenario);

    if (!decisionOf(served).equals(decisionOf(shadowed))) {
        disagreements.add(new Disagreement(
                scenario.id(), decisionOf(served), decisionOf(shadowed)));
    }
    return served;
}

The interesting part isn't the comparison, it's the wiring. The challenger gets its own approval queue and audit trail. It still proposes refunds — into a sandbox nobody is subscribed to. Hand it the real queue and that's not shadow mode, it's a second production agent reviewers can't tell apart from the first.

Champion-challenger rollouts work this way wherever a decision costs something: a pricing model scored against live orders before it sets a price, a perception stack compared against the shipped one before it steers anything.

The obvious objection: I broke this myself, with a rule that fires on request-shaped questions. True, and it's the point. The candidate beat the champion on every number I'd have thought to check. The only thing between it and production was a list of scenarios with expected outcomes.

The honest limit: the diff can't tell me how often real customers phrase a request as a question. E-03 exists because I imagined that phrasing. If it's rare, blocking this cost me a genuine improvement. If it's common, the suite just saved a pile of refunds. Shadow traffic answers that; my dataset can't.

What does your rollout process do when the averages improve and one case gets worse?

── more in #artificial-intelligence 4 stories · sorted by recency
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/a-better-ai-may-neve…] indexed:0 read:4min 2026-09-09 ·