{"slug": "less-is-more-graph-free-multimodal-rag-via-multi-signal-late-fusion", "title": "Less Is More: Graph-free Multimodal RAG via Multi-signal Late Fusion", "summary": "TrioRAG, a graph-free multimodal retrieval-augmented generation framework, matches or outperforms graph-based systems across three benchmarks while cutting total cost and speeding per-query inference by 1.6 to 2.3 times, according to an arXiv paper (2609.19417v1). TrioRAG retrieves independently over a shared multi-vector index of page text and page images using three signals — the question, the anchor image, and a VLM-enhanced query — then combines results through late fusion. The paper also introduces AutoQA, a multimodal automotive benchmark grounded in noisy web-sourced images, where image retrieval reaches only 19.3% document-level recall while text-derived signals, especially the VLM-enhanced query, keep retrieval robust.", "body_md": "arXiv:2609.19417v1 Announce Type: new \nAbstract: Graph-based retrieval-augmented generation (RAG) is widely used for multimodal, cross-document question answering. However, building corpus-level graphs is expensive, slow to query, and difficult to maintain. We present TrioRAG, a graph-free multimodal framework that integrates evidence from three complementary signals: the question, the anchor image, and a VLM-enhanced query generated from both. Each signal retrieves independently over a shared multi-vector index of page text and page images, and the results are combined through late fusion. Further, we introduce AutoQA, a multimodal automotive benchmark whose questions are grounded in noisy, web-sourced images rather than clean document-sourced figures. Its questions require reasoning across manuals. We position it as a model-curated testbed rather than a human-validated gold standard. Across three benchmarks, TrioRAG matches or outperforms graph-based systems while reducing total cost and accelerating per-query inference by 1.6-2.3 times. By construction, AutoQA grounds its questions in out-of-corpus web images. In this setting image retrieval reaches only 19.3% document-level recall, while text-derived signals, especially the VLM-enhanced query, keep retrieval robust.", "url": "https://wpnews.pro/news/less-is-more-graph-free-multimodal-rag-via-multi-signal-late-fusion", "canonical_source": "https://arxiv.org/abs/2609.19417", "published_at": "2026-09-18 04:00:00+00:00", "updated_at": "2026-09-18 04:26:24.114583+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "generative-ai", "computer-vision", "natural-language-processing"], "entities": ["TrioRAG", "AutoQA", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/less-is-more-graph-free-multimodal-rag-via-multi-signal-late-fusion", "markdown": "https://wpnews.pro/news/less-is-more-graph-free-multimodal-rag-via-multi-signal-late-fusion.md", "text": "https://wpnews.pro/news/less-is-more-graph-free-multimodal-rag-via-multi-signal-late-fusion.txt", "jsonld": "https://wpnews.pro/news/less-is-more-graph-free-multimodal-rag-via-multi-signal-late-fusion.jsonld"}}