# When 99.63% Accuracy Wasn't Enough

> Source: <https://dev.to/tanmay_devare_45/when-9963-accuracy-wasnt-enough-3e3h>
> Published: 2026-10-11 19:29:44+00:00

AI agents make hundreds of decisions that look intelligent but are often surprisingly bounded.

Which tool should run next?

Should the workflow retry?

Should this request escalate?

Should the agent continue or hand control back?

The first time an AI system encounters these decisions, using a general-purpose model can make sense.

The thousandth time is a different question.

Ink is built around a simple thesis:

**Models handle novelty. Ink turns proven behavior into software.**

But “proven” matters.

We do not want a system that sees a high classifier score, decides it is probably right, and silently starts replacing model calls.

We want local execution to be something a behavior has to earn.

So we ran a controlled experiment to answer a narrower question:

Can a repeated agent decision accumulate enough evidence to earn local serving authority, while similar decisions are still rejected when the evidence is not strong enough?

The answer was yes.

And the most interesting part was not the decision Ink accepted.

It was the one Ink refused.

This was a controlled internal evaluation, not customer production traffic.

Before running the benchmark, we froze the protocol, code revision, random seeds, dataset identities, model checkpoint identity and qualification configuration. The evaluation used the frozen `phase22e_v1` protocol.

We evaluated six bounded decision workloads representing different operational risk levels.

For this case study, two matter most:

**Tool Routing**

An agent chooses which execution tool or action should run next.

Because a wrong tool invocation can have side effects, we assigned the site a 1% local error budget. That means local behavior needed evidence supporting at least 99% accuracy before it could earn authority.

**Policy Disposition**

A bounded policy decision with the same 1% error budget and 99% required accuracy.

The important rule was established before looking at the final result:

**point accuracy alone was not enough.**

A candidate could only become active when the lower confidence bound on its verified performance cleared the site's required accuracy.

That distinction produced two very different outcomes.

The tool-routing workload accumulated 751 qualification observations.

There were two errors.

Observed accuracy was:

**99.73%**

That number looked good, but Ink did not make the decision based on the point estimate.

The 95% Wilson lower confidence bound was:

**99.03%**

The site's requirement was:

**99.00%**

For the first time, the evidence floor actually cleared the site's configured requirement.

The decision earned `ACTIVE` authority.

This is the transition Ink is designed to manage:

**model-dependent behavior → evidence → qualified local software**

It was no longer enough to say, “the local candidate seems accurate.”

The system had accumulated enough evidence to say:

**this bounded portion of behavior has earned the right to execute locally.**

In the subsequent operational serving window, Ink served 250 tool-routing requests locally.

Observed errors:

**0**

Observed point accuracy:

**100%**

The 95% monitoring lower bound for that smaller 250-request window was 98.49%, which is an important reminder that “0 observed errors” is not equivalent to mathematical certainty.

The result should therefore not be described as “guaranteed 100% accurate.”

What we can say is simpler:

**250 post-activation requests were served locally and no error was observed in that window.**

That distinction matters to us.

Ink is supposed to turn evidence into authority without turning statistics into marketing fiction.

We also compared the representation-assisted compilation path against the classical local baseline.

For the 1,000-request tool-routing workload:

|  | Classical local path | Ink compiler path | 
|---|---|---|
| Host calls | 871 | 750 | 
| Local serves | 129 | 250 | 
| False local serves | 0 | 0 | 

Ink moved **121 additional requests** from the Host path to local execution in this controlled workload, without increasing the observed false-serve count in the tool-routing arm.

That number is intentionally modest.

We are not claiming that Ink removes 80% of an AI stack.

We are not claiming universal savings.

We are saying something narrower:

In this bounded tool-routing workload, additional production-like behavioral evidence allowed 121 decisions that previously required the Host to execute through a qualified local Fast Path instead.

That is the Behavior JIT working as intended.

Now consider the policy-disposition workload.

The local candidate had 272 verified observations.

It made one error.

**99.63%**

The site's required accuracy:

At first glance, this looks like an obvious pass.

99.63% is greater than 99%.

A normal classifier deployment process might stop there.

Ink did not.

The 95% Wilson lower confidence bound was only:

**97.95%**

That meant the available sample did not provide enough evidence to establish the required 99% floor.

So Ink denied serving authority.

The candidate stayed in `EVALUATING`.

The Host remained responsible for the decision.

This result is arguably more important than the tool-routing win.

Because this is the difference between prediction and permission.

Most machine-learning systems ask:

What does the model predict?

Ink asks a second question:

Has this behavior earned the right to act without the model?

Those are not the same problem.

A candidate can have:

high confidence,

high point accuracy,

a strong model,

or an impressive benchmark score

and still lack enough evidence to deserve production authority.

In the policy-disposition case, Ink estimated that with the same one observed error, another **291 verified correct observations** would be required before the Wilson lower bound reached the site's 99% requirement.

Until then, the correct action was not to optimize harder.

It was to keep using the Host.

That is a core product principle:

**A failed qualification is not a failure of Ink. Serving an unqualified decision would be.**

The wider experiment reinforced the same point.

Different workloads ended in different states.

Some had sufficient evidence.

Some had high observed quality but insufficient evidence.

Some failed the quality requirement.

Some had no useful local coverage.

The controller distinguished these cases rather than turning every candidate into a production optimization.

That is important because the easiest way to make an optimization system look impressive is to weaken the conditions under which it is allowed to optimize.

We want the opposite.

Ink should be useful precisely because it is willing to say:

**not yet**

or:

**not this workload**

**send this one to the model.**

This study has limits.

It was an internal controlled benchmark.

It was not traffic from a paying design partner.

The experiment did not perform a real external Host API trial, so the **121 Host calls avoided** metric should not be translated into a dollar amount or a measured remote latency saving from this study. The frozen report explicitly records the real Host trial as not performed.

It also does not prove that every agent has repeatable decisions.

It does not prove that every bounded decision will qualify.

It does not establish a universal local-coverage percentage.

And it does not mean that qualification makes a behavior permanently correct.

Production behavior can change.

Outcomes can drift.

Policies can change.

That is why qualification is only part of the Ink lifecycle.

The full loop is:

**Observe → Qualify → Compile → Serve → Verify**

If the evidence stops supporting local authority, the model needs to take over again.

The conventional way to improve an AI application's economics is to make inference cheaper.

Use a smaller model.

Route between models.

Cache responses.

Optimize prompts.

Fine-tune.

All of those approaches can be useful.

Ink asks a different question:

What if part of the workload has stopped being an inference problem?

If a production system has made a bounded decision hundreds or thousands of times, and independently verified outcomes consistently support the same behavior, continuing to rent that behavior from a general-purpose model may eventually stop making sense.

The production history itself has value.

It represents behavioral knowledge specific to that application.

Ink's job is to determine when enough of that knowledge exists to turn part of it into software.

Not before.

The experiment reinforced three ideas behind Ink.

**1. Repetition is not enough.**

A decision needs a meaningful verifier and enough evidence to support its risk requirement.

**2. Accuracy is not authority.**

99.63% observed accuracy can still be insufficient evidence.

**3. The goal is not maximum local coverage.**

The goal is defensible local coverage.

If only part of a workload earns authority, only that part should become software.

Everything else stays with the model.

This experiment established that the authority lifecycle behaves the way we want under a frozen controlled protocol.

The next standard is harder:

**real production decisions, real outcome signals, and a real external Host.**

That is what we are now looking for.

We are working with teams that have high-volume bounded AI decisions such as:

tool selection,

routing,

retry and recovery,

escalation,

approval gates,

and agent dispatch.

The first step is a Decision Audit.

Give us one repeated model decision and the outcome signal that tells you whether it was right.

We will tell you whether any part of it has a defensible path toward becoming local software.

If it does not, that answer is useful too.
