cd /news/artificial-intelligence/vld-rag-agentic-vision-language-retr… · home topics artificial-intelligence article
[ARTICLE · art-78151] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

VLD-RAG: Agentic Vision-Language Retrieval-Augmented Generation for Long, Visually-Rich Multi-Page Documents

Researchers present VLD-RAG, an agentic multimodal retrieval-augmented generation framework for question answering over visually-rich long documents, achieving improved evidence-page retrieval and end-task performance on LongDocURL and MMLongBench-Doc benchmarks. The framework uses a hybrid retrieval strategy combining keyword-based sparse search with dense semantic queries and a verifier-guided agent workflow to coordinate retrieval, answer, and validation agents.

read1 min views1 publishedJul 29, 2026

arXiv:2607.24748v1 Announce Type: cross Abstract: Visually-rich documents such as reports, slides, and manuals often distribute the evidence needed to answer a question across multiple pages, mixing text with layout cues, tables, charts, and figures. This work studies multimodal retrieval-augmented generation for question answering over such visually-rich long documents, where retrieval must select evidence pages that include both textual and visual signals. We present VLD-RAG, an agentic multimodal RAG framework for multi-page evidence retrieval and cross-page reasoning over long documents. VLD-RAG builds a page-preserving multimodal index that stores parsed text, page-level metadata, and dense visual representations, and uses a hybrid retrieval strategy that combines keyword-based sparse search with dense semantic queries to identify candidate sources and evidence pages. A verifier-guided agent workflow coordinates a Retrieval Agent, Answer Agent, and Validation Agent to broaden evidence coverage, detect missing citations, and refine retrieval requests when needed. We evaluate retrieval with Top-1 and Top-5 evidence-page accuracy and generation with generalized accuracy, and show that VLD-RAG improves both evidence-page retrieval and end-task question answering on visually-rich long-document benchmarks, including LongDocURL and MMLongBench-Doc, outperforming previous vision-based retrieval baselines. These findings highlight that coordinated agent verification and multimodal hybrid retrieval are crucial for reliable grounding when correct answers depend on evidence scattered across pages.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @vld-rag 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/vld-rag-agentic-visi…] indexed:0 read:1min 2026-07-29 ·