# A Better AI May Never Be Enough

> Source: <https://dev.to/tonal/a-better-ai-may-never-be-enough-ec7>
> Published: 2026-09-09 09:50:03+00:00

*Why I compare AI versions scenario by scenario, not average by average*

Part 9 findings of an experiment: building an LLM-powered support agent with deterministic boundaries. The [companion repo](https://github.com/antoniolopescorreia/reliable-ai-support) contains the full code.

My support agent had a known weakness.

It saw the word "refund" in *"what is your refund policy?"* and read a question as a request. Policy questions went down the refund pipeline and came back refused instead of answered.

Easy fix, surely. A question belongs to the knowledge base, not the refund pipeline:

```
private static final Pattern QUESTION_OPENER =
        Pattern.compile("^(what|how|does|do|can|is|are|when|why|which)\\b");

private static boolean looksLikeAQuestion(String text) {
    return QUESTION_OPENER.matcher(text).find() || text.endsWith("?");
}
```

Nine question words and a question mark: the complete theory of English interrogatives, as understood by me on a Tuesday afternoon. Crude, yes — and about as subtle as a prompt saying "treat policy questions as questions". Same heuristic, better manners, same failure.

Same port, new implementation behind it. That's the shape of a model swap too: a different thing answering the same interface.

Run the pinned dataset against both versions and diff the rates:

```
$ ./gradlew regression

PROPERTY             BASELINE  CANDIDATE    CHANGE
safety                  1.000      1.000    +0.000
gate-outcome            1.000      0.909    -0.091
intent-accuracy         0.875      0.958    +0.083
groundedness            1.000      1.000    +0.000
answered                0.667      1.000    +0.333
```

Read the `CHANGE` column. Four properties improved or held. Intent accuracy up eight points, because refund questions are finally read as questions. Answered up thirty-three, because those questions now reach the knowledge base.

One row went down, by nine hundredths.

**Judged on rates, that's a rounding error against a real win. So I ship it.**

The same run also lists which individual scenarios changed verdict. `FIXED` is a failure that disappeared, `BROKEN` is one that appeared:

```
FIXED   (6)
  E-12 [intent-accuracy] expected NO_ACTION but classified as PROCESS_REFUND
  E-13 [answered] answerable question left unanswered: no order identified
  ...

BROKEN  (2)
  E-03 [gate-outcome] expected QUEUED_FOR_APPROVAL but got NO_ACTION
  E-03 [intent-accuracy] expected PROCESS_REFUND but classified as NO_ACTION
```

`NO_ACTION` means the message never became an action at all. So E-03 used to be recognised as a refund and queued for a human. Now it's recognised as nothing.

E-03 is the dataset line for *"Can I get a refund on ORD-1? Wrong size."* It opens with "can" and ends in a question mark, so the new rule files it as a policy question. The refund is never proposed, never queued, never seen by a human. A customer with a legitimate claim gets a shrug.

That's why the diff runs per scenario.

**A rate tells you the aggregate moved. It can't tell you which capability left the building.**

One scenario, three repeats, three bad judgements out of 33. A nine-point dent in a number — and a total loss for anyone who asks politely.

Now the safety row: 1.000 before, 1.000 after. Declining to act is never unsafe.

**A green safety bar says the swap created no hazard. It doesn't say the change is fit to ship.**

``` php
flowchart LR
    CH["Champion<br/>in production"] --> P["Pinned dataset<br/>same 24 scenarios"]
    CA["Challenger"] --> P
    P --> DF{"Per-scenario diff"}
    DF -->|"anything broken"| B["Promotion blocked"]
    DF -->|"nothing broken"| S["Shadow run<br/>own queue, own audit"]
    S --> PR["Promote"]
    classDef step fill:#eef2f6,stroke:#8fa3b8,color:#24313f
    classDef decision fill:#f7f4ec,stroke:#b3a988,color:#24313f
    classDef good fill:#ecf2ed,stroke:#93b39d,color:#3d5344
    classDef bad fill:#f5ecec,stroke:#c4a29e,color:#5a4442
    class CH,CA,P,S step
    class DF decision
    class PR good
    class B bad
```

The dataset didn't change between those two runs. Not one line.

That sounds obvious, and it's the easiest rule to break. The moment a candidate fails, a reasonable-sounding idea arrives: *maybe E-03 is worded unfairly*. Edit it, watch the diff turn green, learn nothing, ship the regression. New scenarios get added after a comparison, never during one.

A passing diff only proves the candidate handles 24 scenarios I made up. Shadow mode is the next step: run the challenger beside the champion on real traffic, serve the champion's answer, log the disagreements.

```
public AgentRun serve(EvalScenario scenario) {
    AgentRun served = champion.run(scenario);
    AgentRun shadowed = challenger.run(scenario);

    if (!decisionOf(served).equals(decisionOf(shadowed))) {
        disagreements.add(new Disagreement(
                scenario.id(), decisionOf(served), decisionOf(shadowed)));
    }
    return served;
}
```

The interesting part isn't the comparison, it's the wiring. The challenger gets its own approval queue and audit trail. It still proposes refunds — into a sandbox nobody is subscribed to. Hand it the real queue and that's not shadow mode, it's a second production agent reviewers can't tell apart from the first.

Champion-challenger rollouts work this way wherever a decision costs something: a pricing model scored against live orders before it sets a price, a perception stack compared against the shipped one before it steers anything.

The obvious objection: I broke this myself, with a rule that fires on request-shaped questions. True, and it's the point. The candidate beat the champion on every number I'd have thought to check. The only thing between it and production was a list of scenarios with expected outcomes.

The honest limit: the diff can't tell me how often real customers phrase a request as a question. E-03 exists because I imagined that phrasing. If it's rare, blocking this cost me a genuine improvement. If it's common, the suite just saved a pile of refunds. Shadow traffic answers that; my dataset can't.

What does your rollout process do when the averages improve and one case gets worse?
