{"slug": "ece-achieves-97-8-accuracy-on-answered-claims", "title": "ECE achieves 97.8% accuracy on answered claims", "summary": "A tool-using fact-checking agent achieved 97.8% selective accuracy on answered claims by adding an \"uncertain\" abstention verdict, deferring just 6 of 95 cases, while overall accuracy remained at 91.6%, according to a preprint on arXiv. The Evidence Chain Evaluation framework enables the agent to abstain on uncertain claims, acting as a targeted safety valve for weak evidence rather than a general calibration fix.", "body_md": "[arXiv](https://arxiv.org/abs/2607.18240)\n\n### ECE achieves 97.8% accuracy on answered claims\n\nWhich summary reads better? Pick one — models revealed after.Both summaries are AI-generated.\n\nAdding an \"uncertain\" abstention verdict to a tool-using fact-checking agent pushed selective accuracy on answered claims to 97.8% by deferring just 6 of 95 cases—almost all concentrated in weak-evidence settings—while overall accuracy stayed at 91.6%. Notably, this abstention gate didn't improve aggregate calibration metrics (ECE, Brier, AURC), so if you're routing verification agents in production, treat abstention as a targeted safety valve for epistemically thin evidence rather than a general confidence-calibration fix.\n\nImplementing the Evidence Chain Evaluation framework for tool-using agents yields a 97.8% selective accuracy on fact-checking tasks by enabling the agent to abstain on uncertain claims instead of forcing a binary verdict. For production systems, this provides a reliable safety-valve mechanism that filters out weak or inconsistent source evidence by deferring low-reliability queries, even though it does not improve overall aggregate calibration metrics like the Brier score. This allows you to deploy highly reliable automated content verification workflows where the vast majority of claims are handled with near-perfect precision while ambiguous edge cases are safely escalated.", "url": "https://wpnews.pro/news/ece-achieves-97-8-accuracy-on-answered-claims", "canonical_source": "https://www.snipvote.com/story/cmrvr5e1i00046l3wlqagmm4k", "published_at": "2026-07-22 07:23:20.305357+00:00", "updated_at": "2026-07-22 07:23:22.144713+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-agents", "ai-research"], "entities": ["arXiv"], "alternates": {"html": "https://wpnews.pro/news/ece-achieves-97-8-accuracy-on-answered-claims", "markdown": "https://wpnews.pro/news/ece-achieves-97-8-accuracy-on-answered-claims.md", "text": "https://wpnews.pro/news/ece-achieves-97-8-accuracy-on-answered-claims.txt", "jsonld": "https://wpnews.pro/news/ece-achieves-97-8-accuracy-on-answered-claims.jsonld"}}