{"slug": "when-99-63-accuracy-wasn-t-enough", "title": "When 99.63% Accuracy Wasn't Enough", "summary": "A developer built Ink, a system that lets AI agents earn local execution authority for repeated bounded decisions only when the lower confidence bound on verified accuracy clears a site's required threshold. In a frozen internal evaluation, a tool-routing workload with 751 qualification observations and two errors (99.73% observed accuracy) cleared a 99% requirement with a 95% Wilson lower bound of 99.03%, earning ACTIVE authority; 250 subsequent requests were then served locally with zero observed errors, while a policy-disposition workload was refused. On a 1,000-request tool-routing workload, the Ink compiler path moved 121 additional requests from the host path to local serving versus the classical local baseline.", "body_md": "AI agents make hundreds of decisions that look intelligent but are often surprisingly bounded.\n\nWhich tool should run next?\n\nShould the workflow retry?\n\nShould this request escalate?\n\nShould the agent continue or hand control back?\n\nThe first time an AI system encounters these decisions, using a general-purpose model can make sense.\n\nThe thousandth time is a different question.\n\nInk is built around a simple thesis:\n\n**Models handle novelty. Ink turns proven behavior into software.**\n\nBut “proven” matters.\n\nWe do not want a system that sees a high classifier score, decides it is probably right, and silently starts replacing model calls.\n\nWe want local execution to be something a behavior has to earn.\n\nSo we ran a controlled experiment to answer a narrower question:\n\nCan a repeated agent decision accumulate enough evidence to earn local serving authority, while similar decisions are still rejected when the evidence is not strong enough?\n\nThe answer was yes.\n\nAnd the most interesting part was not the decision Ink accepted.\n\nIt was the one Ink refused.\n\nThis was a controlled internal evaluation, not customer production traffic.\n\nBefore running the benchmark, we froze the protocol, code revision, random seeds, dataset identities, model checkpoint identity and qualification configuration. The evaluation used the frozen `phase22e_v1` protocol.\n\nWe evaluated six bounded decision workloads representing different operational risk levels.\n\nFor this case study, two matter most:\n\n**Tool Routing**\n\nAn agent chooses which execution tool or action should run next.\n\nBecause a wrong tool invocation can have side effects, we assigned the site a 1% local error budget. That means local behavior needed evidence supporting at least 99% accuracy before it could earn authority.\n\n**Policy Disposition**\n\nA bounded policy decision with the same 1% error budget and 99% required accuracy.\n\nThe important rule was established before looking at the final result:\n\n**point accuracy alone was not enough.**\n\nA candidate could only become active when the lower confidence bound on its verified performance cleared the site's required accuracy.\n\nThat distinction produced two very different outcomes.\n\nThe tool-routing workload accumulated 751 qualification observations.\n\nThere were two errors.\n\nObserved accuracy was:\n\n**99.73%**\n\nThat number looked good, but Ink did not make the decision based on the point estimate.\n\nThe 95% Wilson lower confidence bound was:\n\n**99.03%**\n\nThe site's requirement was:\n\n**99.00%**\n\nFor the first time, the evidence floor actually cleared the site's configured requirement.\n\nThe decision earned `ACTIVE` authority.\n\nThis is the transition Ink is designed to manage:\n\n**model-dependent behavior → evidence → qualified local software**\n\nIt was no longer enough to say, “the local candidate seems accurate.”\n\nThe system had accumulated enough evidence to say:\n\n**this bounded portion of behavior has earned the right to execute locally.**\n\nIn the subsequent operational serving window, Ink served 250 tool-routing requests locally.\n\nObserved errors:\n\n**0**\n\nObserved point accuracy:\n\n**100%**\n\nThe 95% monitoring lower bound for that smaller 250-request window was 98.49%, which is an important reminder that “0 observed errors” is not equivalent to mathematical certainty.\n\nThe result should therefore not be described as “guaranteed 100% accurate.”\n\nWhat we can say is simpler:\n\n**250 post-activation requests were served locally and no error was observed in that window.**\n\nThat distinction matters to us.\n\nInk is supposed to turn evidence into authority without turning statistics into marketing fiction.\n\nWe also compared the representation-assisted compilation path against the classical local baseline.\n\nFor the 1,000-request tool-routing workload:\n\n|  | Classical local path | Ink compiler path | \n|---|---|---|\n| Host calls | 871 | 750 | \n| Local serves | 129 | 250 | \n| False local serves | 0 | 0 | \n\nInk moved **121 additional requests** from the Host path to local execution in this controlled workload, without increasing the observed false-serve count in the tool-routing arm.\n\nThat number is intentionally modest.\n\nWe are not claiming that Ink removes 80% of an AI stack.\n\nWe are not claiming universal savings.\n\nWe are saying something narrower:\n\nIn this bounded tool-routing workload, additional production-like behavioral evidence allowed 121 decisions that previously required the Host to execute through a qualified local Fast Path instead.\n\nThat is the Behavior JIT working as intended.\n\nNow consider the policy-disposition workload.\n\nThe local candidate had 272 verified observations.\n\nIt made one error.\n\n**99.63%**\n\nThe site's required accuracy:\n\nAt first glance, this looks like an obvious pass.\n\n99.63% is greater than 99%.\n\nA normal classifier deployment process might stop there.\n\nInk did not.\n\nThe 95% Wilson lower confidence bound was only:\n\n**97.95%**\n\nThat meant the available sample did not provide enough evidence to establish the required 99% floor.\n\nSo Ink denied serving authority.\n\nThe candidate stayed in `EVALUATING`.\n\nThe Host remained responsible for the decision.\n\nThis result is arguably more important than the tool-routing win.\n\nBecause this is the difference between prediction and permission.\n\nMost machine-learning systems ask:\n\nWhat does the model predict?\n\nInk asks a second question:\n\nHas this behavior earned the right to act without the model?\n\nThose are not the same problem.\n\nA candidate can have:\n\nhigh confidence,\n\nhigh point accuracy,\n\na strong model,\n\nor an impressive benchmark score\n\nand still lack enough evidence to deserve production authority.\n\nIn the policy-disposition case, Ink estimated that with the same one observed error, another **291 verified correct observations** would be required before the Wilson lower bound reached the site's 99% requirement.\n\nUntil then, the correct action was not to optimize harder.\n\nIt was to keep using the Host.\n\nThat is a core product principle:\n\n**A failed qualification is not a failure of Ink. Serving an unqualified decision would be.**\n\nThe wider experiment reinforced the same point.\n\nDifferent workloads ended in different states.\n\nSome had sufficient evidence.\n\nSome had high observed quality but insufficient evidence.\n\nSome failed the quality requirement.\n\nSome had no useful local coverage.\n\nThe controller distinguished these cases rather than turning every candidate into a production optimization.\n\nThat is important because the easiest way to make an optimization system look impressive is to weaken the conditions under which it is allowed to optimize.\n\nWe want the opposite.\n\nInk should be useful precisely because it is willing to say:\n\n**not yet**\n\nor:\n\n**not this workload**\n\n**send this one to the model.**\n\nThis study has limits.\n\nIt was an internal controlled benchmark.\n\nIt was not traffic from a paying design partner.\n\nThe experiment did not perform a real external Host API trial, so the **121 Host calls avoided** metric should not be translated into a dollar amount or a measured remote latency saving from this study. The frozen report explicitly records the real Host trial as not performed.\n\nIt also does not prove that every agent has repeatable decisions.\n\nIt does not prove that every bounded decision will qualify.\n\nIt does not establish a universal local-coverage percentage.\n\nAnd it does not mean that qualification makes a behavior permanently correct.\n\nProduction behavior can change.\n\nOutcomes can drift.\n\nPolicies can change.\n\nThat is why qualification is only part of the Ink lifecycle.\n\nThe full loop is:\n\n**Observe → Qualify → Compile → Serve → Verify**\n\nIf the evidence stops supporting local authority, the model needs to take over again.\n\nThe conventional way to improve an AI application's economics is to make inference cheaper.\n\nUse a smaller model.\n\nRoute between models.\n\nCache responses.\n\nOptimize prompts.\n\nFine-tune.\n\nAll of those approaches can be useful.\n\nInk asks a different question:\n\nWhat if part of the workload has stopped being an inference problem?\n\nIf a production system has made a bounded decision hundreds or thousands of times, and independently verified outcomes consistently support the same behavior, continuing to rent that behavior from a general-purpose model may eventually stop making sense.\n\nThe production history itself has value.\n\nIt represents behavioral knowledge specific to that application.\n\nInk's job is to determine when enough of that knowledge exists to turn part of it into software.\n\nNot before.\n\nThe experiment reinforced three ideas behind Ink.\n\n**1. Repetition is not enough.**\n\nA decision needs a meaningful verifier and enough evidence to support its risk requirement.\n\n**2. Accuracy is not authority.**\n\n99.63% observed accuracy can still be insufficient evidence.\n\n**3. The goal is not maximum local coverage.**\n\nThe goal is defensible local coverage.\n\nIf only part of a workload earns authority, only that part should become software.\n\nEverything else stays with the model.\n\nThis experiment established that the authority lifecycle behaves the way we want under a frozen controlled protocol.\n\nThe next standard is harder:\n\n**real production decisions, real outcome signals, and a real external Host.**\n\nThat is what we are now looking for.\n\nWe are working with teams that have high-volume bounded AI decisions such as:\n\ntool selection,\n\nrouting,\n\nretry and recovery,\n\nescalation,\n\napproval gates,\n\nand agent dispatch.\n\nThe first step is a Decision Audit.\n\nGive us one repeated model decision and the outcome signal that tells you whether it was right.\n\nWe will tell you whether any part of it has a defensible path toward becoming local software.\n\nIf it does not, that answer is useful too.", "url": "https://wpnews.pro/news/when-99-63-accuracy-wasn-t-enough", "canonical_source": "https://dev.to/tanmay_devare_45/when-9963-accuracy-wasnt-enough-3e3h", "published_at": "2026-10-11 19:29:44+00:00", "updated_at": "2026-10-11 19:32:11.070751+00:00", "lang": "en", "topics": ["ai-agents", "machine-learning", "ai-infrastructure"], "entities": ["Ink"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/when-99-63-accuracy-wasn-t-enough", "markdown": "https://wpnews.pro/news/when-99-63-accuracy-wasn-t-enough.md", "text": "https://wpnews.pro/news/when-99-63-accuracy-wasn-t-enough.txt", "jsonld": "https://wpnews.pro/news/when-99-63-accuracy-wasn-t-enough.jsonld"}}