cd /news/artificial-intelligence/the-checking-problem-what-must-be-tr… · home topics artificial-intelligence article
[ARTICLE · art-84188] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

The Checking Problem: What must be true before AI ships in a regulated firm

A new arXiv paper (arXiv:2607.28666v1) finds that only 32 of 72 AI configurations (56.1%) in regulated financial services meet a production bar requiring sustained accuracy, reproducibility, verifiable attribution, and informative confidence signals, based on 5,093 scored output elements across six document-heavy workflows. The paper, which measures why enterprise AI programmes stall, estimates that requiring tools to cite sources and state confidence reduces human review burden from 100% to 49% while holding error tolerance in 17 of 20 configurations, whereas adding a self-verification pass costs 2.3 times latency and fails to hold error tolerance.

read1 min views1 publishedAug 3, 2026

arXiv:2607.28666v1 Announce Type: new Abstract: Enterprise AI programmes stall at a rate that is widely quoted and poorly explained. This paper measures the mechanism. Six document-heavy workflows of the kind performed daily in regulated financial services were run across four model families and three tool configurations, three times each, producing 5,093 scored output elements across 72 configurations. Each configuration was assessed twice: against a demonstration bar, being a single correct run on a single case, and against a production bar requiring sustained accuracy, reproducibility across repeats, verifiable attribution, and a confidence signal that carries information. 57 of 72 configurations cleared the demonstration bar and 32 cleared the production bar, a survival rate of 56.1%. The paper then computes the review burden each configuration imposes, estimated out of sample rather than with hindsight. A tool that states no confidence requires review of 100% of its output, because it offers a reviewer no basis for triage. Requiring the tool to cite its sources and state a confidence reduces that to 49% while holding the residual error tolerance in 17 of 20 configurations. Adding a self-verification pass costs 2.3 times the latency of the plain configuration, reaches 44%, and is the only configuration that fails to hold the error tolerance. The practical implication is that the value of an AI workflow is set less by how often it is right than by how much of it a human must still check, and that the second property is measurable and rarely measured.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-checking-problem…] indexed:0 read:1min 2026-08-03 ·