{"slug": "multimodal-colrag-tf-triple-filtered-retrieval-for-complex-pdfs", "title": "Multimodal CoLRAG-TF: Triple-Filtered Retrieval for Complex PDFs", "summary": "Researchers present Multimodal CoLRAG-TF, a retrieval-augmented generation architecture that integrates dense text embeddings, BM25 keyword matching, knowledge-graph triple filtering, and image-based similarity to handle complex PDFs. Evaluated on a 457-pair benchmark from Japanese disaster lesson PDFs, the system achieves a Retrieval Recall of 0.9909 and a 71.6% improvement in multi-hop answer similarity over single-hop queries. The approach demonstrates that triple-filtered multimodal fusion is essential for structured reasoning over noisy, heterogeneous documents.", "body_md": "arXiv:2607.20517v1 Announce Type: new\nAbstract: Retrieval-augmented generation (RAG) over heterogeneous PDF collections remains challenging due to multimodal content, domain-specific terminology, and the need for multi-hop reasoning across dispersed evidence. We present Multimodal CoLRAG-TF, a four-axis fusion architecture that integrates dense text embeddings, BM25 keyword matching, knowledge-graph triple filtering, and image-based similarity for robust retrieval over complex documents. Our system constructs a multimodal index of 2,403 blocks extracted from 43 Japanese disaster lesson PDFs, supported by a hybrid OCR pipeline and LLM-based caption generation. To enhance compositional reasoning, we extract 11,414 OpenIE triples and index them with FAISS, enabling sub-second triple lookup and hierarchical propagation of relevance signals. A HippoRAG2-inspired coarse-to-fine retriever (volume $\\to$ chapter $\\to$ block) narrows the search space before final fusion scoring. Bayesian optimization over fusion weights reveals that the triple axis must dominate ($\\alpha_\\text{triple} = 0.44$) to counteract lexical bias and sustain multi-hop retrieval quality. Evaluated on a 457-pair benchmark, Multimodal CoLRAG-TF achieves a Retrieval Recall of 0.9909 and a 71.6$\\%$ improvement in multi-hop answer similarity over single-hop queries. An image-to-lesson pipeline using a vision LLM further demonstrates the applicability of the approach to visual inputs. These results show that triple-filtered multimodal fusion is essential for structured reasoning over noisy, heterogeneous PDFs and provides a general framework applicable beyond the disaster domain.", "url": "https://wpnews.pro/news/multimodal-colrag-tf-triple-filtered-retrieval-for-complex-pdfs", "canonical_source": "https://arxiv.org/abs/2607.20517", "published_at": "2026-07-24 04:00:00+00:00", "updated_at": "2026-07-24 04:11:55.507367+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "natural-language-processing", "ai-research"], "entities": ["Multimodal CoLRAG-TF", "FAISS", "HippoRAG2", "OpenIE", "BM25"], "alternates": {"html": "https://wpnews.pro/news/multimodal-colrag-tf-triple-filtered-retrieval-for-complex-pdfs", "markdown": "https://wpnews.pro/news/multimodal-colrag-tf-triple-filtered-retrieval-for-complex-pdfs.md", "text": "https://wpnews.pro/news/multimodal-colrag-tf-triple-filtered-retrieval-for-complex-pdfs.txt", "jsonld": "https://wpnews.pro/news/multimodal-colrag-tf-triple-filtered-retrieval-for-complex-pdfs.jsonld"}}