{"slug": "the-checking-problem-what-must-be-true-before-ai-ships-in-a-regulated-firm", "title": "The Checking Problem: What must be true before AI ships in a regulated firm", "summary": "A new arXiv paper (arXiv:2607.28666v1) finds that only 32 of 72 AI configurations (56.1%) in regulated financial services meet a production bar requiring sustained accuracy, reproducibility, verifiable attribution, and informative confidence signals, based on 5,093 scored output elements across six document-heavy workflows. The paper, which measures why enterprise AI programmes stall, estimates that requiring tools to cite sources and state confidence reduces human review burden from 100% to 49% while holding error tolerance in 17 of 20 configurations, whereas adding a self-verification pass costs 2.3 times latency and fails to hold error tolerance.", "body_md": "arXiv:2607.28666v1 Announce Type: new\nAbstract: Enterprise AI programmes stall at a rate that is widely quoted and poorly explained. This paper measures the mechanism. Six document-heavy workflows of the kind performed daily in regulated financial services were run across four model families and three tool configurations, three times each, producing 5,093 scored output elements across 72 configurations. Each configuration was assessed twice: against a demonstration bar, being a single correct run on a single case, and against a production bar requiring sustained accuracy, reproducibility across repeats, verifiable attribution, and a confidence signal that carries information. 57 of 72 configurations cleared the demonstration bar and 32 cleared the production bar, a survival rate of 56.1%. The paper then computes the review burden each configuration imposes, estimated out of sample rather than with hindsight. A tool that states no confidence requires review of 100% of its output, because it offers a reviewer no basis for triage. Requiring the tool to cite its sources and state a confidence reduces that to 49% while holding the residual error tolerance in 17 of 20 configurations. Adding a self-verification pass costs 2.3 times the latency of the plain configuration, reaches 44%, and is the only configuration that fails to hold the error tolerance. The practical implication is that the value of an AI workflow is set less by how often it is right than by how much of it a human must still check, and that the second property is measurable and rarely measured.", "url": "https://wpnews.pro/news/the-checking-problem-what-must-be-true-before-ai-ships-in-a-regulated-firm", "canonical_source": "https://arxiv.org/abs/2607.28666", "published_at": "2026-08-03 04:00:00+00:00", "updated_at": "2026-08-03 04:04:25.520959+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-research", "ai-products"], "entities": ["arXiv"], "alternates": {"html": "https://wpnews.pro/news/the-checking-problem-what-must-be-true-before-ai-ships-in-a-regulated-firm", "markdown": "https://wpnews.pro/news/the-checking-problem-what-must-be-true-before-ai-ships-in-a-regulated-firm.md", "text": "https://wpnews.pro/news/the-checking-problem-what-must-be-true-before-ai-ships-in-a-regulated-firm.txt", "jsonld": "https://wpnews.pro/news/the-checking-problem-what-must-be-true-before-ai-ships-in-a-regulated-firm.jsonld"}}