cd /news/artificial-intelligence/verification-without-sufficiency-per… · home topics artificial-intelligence article
[ARTICLE · art-85577] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Verification Without Sufficiency: Per-Chunk Filtering Fails on Multi-Hop RAG, and Decomposition Repairs It

A new arXiv preprint (2608.00585v1) shows that per-chunk verification fails on multi-hop retrieval-augmented generation, with entailment scoring reaching only 0.643, 0.523, and 0.560 AUC on HotpotQA, 2WikiMultihopQA, and MuSiQue, respectively, versus 0.951 on single-hop SQuAD. The authors demonstrate that conditioning verification on decomposed sub-questions repairs the failure, lifting entailment on a later hop from 0.546 to 0.840 (+0.355, bootstrap interval [0.331, 0.382]) using MuSiQue's gold decomposition, and that an off-the-shelf Qwen2.5-7B decomposer captures 31% of that ceiling.

read1 min views1 publishedAug 4, 2026

arXiv:2608.00585v1 Announce Type: new Abstract: Verification for retrieval-augmented generation usually scores each retrieved chunk and drops the ones that fail. We show this cannot work for multi-hop questions, and show what does. Per-chunk scoring assumes one chunk is a sufficient premise for the answer. Multi-hop questions are built so that none is, and the paragraph carrying the answer is the one the question does not name. Entailment scoring reaches 0.643, 0.523 and 0.560 AUC on HotpotQA, 2WikiMultihopQA and MuSiQue, against 0.951 on single-hop SQuAD. Seven controls rule out model capacity, premise length, hypothesis template, decision threshold, retriever, answer-matching criterion and prompt. End to end across three datasets, three generator sizes and two prompts, per-chunk gating is significantly worse than not filtering at all in every cell, and its penalty grows with generator capability. The repair is to condition verification on the decomposed sub-question rather than the original query. Using MuSiQue's gold decomposition, entailment on a later hop rises from 0.546, which is chance, to 0.840, a paired lift of +0.355 with a bootstrap interval of [0.331, 0.382]. An off-the-shelf Qwen2.5-7B decomposer, given the question and the top retrieved paragraph, reaches 0.637 and captures 31% of that ceiling; decomposing without retrieval reaches 0.533, below the original question. Iterative retrieval systems already produce such decompositions and discard them before verifying.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/verification-without…] indexed:0 read:1min 2026-08-04 ·