cd /news/ai-research/certified-against-which-oracle-execu… · home topics ai-research article
[ARTICLE · art-137857] src=machinebrief.com ↗ pub= topic=ai-research verified=true sentiment=· neutral

Certified Against Which Oracle? Execution Labels Set the Reported Risk of Conformal Abstention for Text-to-SQL

A preregistered intervention on the Spider-Realistic text-to-SQL benchmark found that swapping the shipped single-database oracle for the benchmark's distilled multi-instance test suite raised a conformal abstention certificate's held-out risk 2.73 to 10.23 points above the risk its own labels report, across four SQL-specialist checkpoints and two split schemes. Under blinded labels from two SQL experts, a certificate calibrated at a nominal 0.10 carried 20.0 and 17.2 points of risk on two checkpoints, and every execution-consistency score looked better under the labels of the oracle that built its clusters in 16 of 16 combinations. The authors conclude a certificate should be reported with both oracles and that a consistency score should be evaluated under an oracle that did not build it.

by read2 min views1 publishedSep 23, 2026

arXiv:2609.25938v1 Announce Type: new Abstract: A conformal abstention certificate for text-to-SQL is only as truthful as the correctness labels it is calibrated on. The uncertainty pipelines that read confidence off execution consistency take those labels from the single database a benchmark ships, an oracle known to be lenient. We run a preregistered intervention on Spider-Realistic, swapping that database for the benchmark's distilled multi-instance test suite. Across four SQL-specialist checkpoints and two split schemes, the swap raises the certificate's held-out risk 2.73 to 10.23 points above the risk its own labels report. Neither oracle reports the risk experts assign. Under blinded labels from two SQL experts, a certificate calibrated at a nominal 0.10 carries 20.0 and 17.2 points of risk on two checkpoints. The stricter oracle errs in both directions: most of the answers it rejects are not judged wrong, and some of those it accepts are. An AI-assigned census of what it rejects finds a semantic error in a quarter to a third of them, depending on the population. It attributes most of the rest to underspecified questions, synthetic instances or suspected reference-query defects, a flag supported by a preregistered blinded expert audit. The oracle also decides how a confidence score is judged. Every execution-consistency score looks better under the labels of the oracle that built its clusters, in 16 of 16 combinations. Under expert labels, building such a score on suite clusters instead of shipped-database clusters raises its area under the ROC curve (AUROC) by 6.96 points on one checkpoint and 1.53 on the other. On the second, the expert interval excludes the 8.3 points the suite labels report. A certificate should be reported with both oracles, and an oracle-relative difference read as semantic risk only after the benchmark is audited. A consistency score should be evaluated under an oracle that did not build it.

── more in #ai-research 4 stories · sorted by recency
── more on @spider-realistic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/certified-against-wh…] indexed:0 read:2min 2026-09-23 ·