{"slug": "vld-rag-agentic-vision-language-retrieval-augmented-generation-for-long-visually", "title": "VLD-RAG: Agentic Vision-Language Retrieval-Augmented Generation for Long, Visually-Rich Multi-Page Documents", "summary": "Researchers present VLD-RAG, an agentic multimodal retrieval-augmented generation framework for question answering over visually-rich long documents, achieving improved evidence-page retrieval and end-task performance on LongDocURL and MMLongBench-Doc benchmarks. The framework uses a hybrid retrieval strategy combining keyword-based sparse search with dense semantic queries and a verifier-guided agent workflow to coordinate retrieval, answer, and validation agents.", "body_md": "arXiv:2607.24748v1 Announce Type: cross\nAbstract: Visually-rich documents such as reports, slides, and manuals often distribute the evidence needed to answer a question across multiple pages, mixing text with layout cues, tables, charts, and figures. This work studies multimodal retrieval-augmented generation for question answering over such visually-rich long documents, where retrieval must select evidence pages that include both textual and visual signals. We present VLD-RAG, an agentic multimodal RAG framework for multi-page evidence retrieval and cross-page reasoning over long documents. VLD-RAG builds a page-preserving multimodal index that stores parsed text, page-level metadata, and dense visual representations, and uses a hybrid retrieval strategy that combines keyword-based sparse search with dense semantic queries to identify candidate sources and evidence pages. A verifier-guided agent workflow coordinates a Retrieval Agent, Answer Agent, and Validation Agent to broaden evidence coverage, detect missing citations, and refine retrieval requests when needed. We evaluate retrieval with Top-1 and Top-5 evidence-page accuracy and generation with generalized accuracy, and show that VLD-RAG improves both evidence-page retrieval and end-task question answering on visually-rich long-document benchmarks, including LongDocURL and MMLongBench-Doc, outperforming previous vision-based retrieval baselines. These findings highlight that coordinated agent verification and multimodal hybrid retrieval are crucial for reliable grounding when correct answers depend on evidence scattered across pages.", "url": "https://wpnews.pro/news/vld-rag-agentic-vision-language-retrieval-augmented-generation-for-long-visually", "canonical_source": "https://www.machinebrief.com/news/vld-rag-agentic-vision-language-retrieval-augmented-generati-434t", "published_at": "2026-07-29 04:00:00+00:00", "updated_at": "2026-07-29 06:00:04.810469+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research", "ai-agents"], "entities": ["VLD-RAG", "LongDocURL", "MMLongBench-Doc"], "alternates": {"html": "https://wpnews.pro/news/vld-rag-agentic-vision-language-retrieval-augmented-generation-for-long-visually", "markdown": "https://wpnews.pro/news/vld-rag-agentic-vision-language-retrieval-augmented-generation-for-long-visually.md", "text": "https://wpnews.pro/news/vld-rag-agentic-vision-language-retrieval-augmented-generation-for-long-visually.txt", "jsonld": "https://wpnews.pro/news/vld-rag-agentic-vision-language-retrieval-augmented-generation-for-long-visually.jsonld"}}