{"slug": "hc-rag-evidence-centric-retrieval-augmented-generation-over-heterogeneous", "title": "HC-RAG: Evidence-Centric Retrieval-Augmented Generation over Heterogeneous Financial Filings", "summary": "Researchers propose HC-RAG, a hierarchical cross-modal retrieval-augmented generation framework for evidence-centric financial question answering, which organizes filings into a typed financial evidence graph and routes evidence by query intent. On the new Multi-Doc-2025 benchmark of 2,327 expert-verified QA pairs from 179 SEC 10-K filings, HC-RAG outperforms GraphRAG by 10.9 F1 points, and it beats RAPTOR by 6.6 F1 points on DocFinQA.", "body_md": "arXiv:2608.12335v1 Announce Type: new\nAbstract: Financial question answering over annual reports requires more than retrieving semantically similar passages. It often involves identifying relevant companies and fiscal years, locating standardized filing sections, collecting textual and tabular evidence, and checking answers against the original documents. Existing RAG systems, however, usually flatten long filings into unordered chunks, pay limited attention to the typed structure of financial reports, and use fixed text-table fusion strategies without considering query intent. To address these limitations, we propose \\textbf{HC-RAG}, a hierarchical cross-modal retrieval-augmented generation framework for evidence-centric financial QA. HC-RAG organizes filings into a typed financial evidence graph with documents, sections, text units, table units, and metadata nodes. It retrieves evidence through document-section-unit paths, aligns textual and tabular evidence in a shared retrieval space, and routes evidence according to four semantic intents: calculation, trend, fact, and comparison. We further introduce \\textbf{Multi-Doc-2025}, a benchmark containing 2,327 expert-verified QA pairs from 179 SEC 10-K filings of 87 S\\&P 500 companies across fiscal years 2022--2024, with labels for intent, difficulty, and structural evidence attributes. Experiments on public financial QA benchmarks and Multi-Doc-2025 show that HC-RAG improves both answer quality and evidence localization, especially in long-document, table-related, and cross-document settings. HC-RAG outperforms RAPTOR by 6.6 F1 points on DocFinQA and GraphRAG by 10.9 F1 points on Multi-Doc-2025. Evidence-level analysis and ablation studies show that the improvements mainly come from more accurate section localization, table grounding, cross-document evidence aggregation, and intent-aware text-table routing.", "url": "https://wpnews.pro/news/hc-rag-evidence-centric-retrieval-augmented-generation-over-heterogeneous", "canonical_source": "https://arxiv.org/abs/2608.12335", "published_at": "2026-08-14 04:00:00+00:00", "updated_at": "2026-08-14 04:12:46.219654+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "natural-language-processing", "generative-ai"], "entities": ["HC-RAG", "Multi-Doc-2025", "RAPTOR", "GraphRAG", "DocFinQA", "SEC", "S&P 500"], "alternates": {"html": "https://wpnews.pro/news/hc-rag-evidence-centric-retrieval-augmented-generation-over-heterogeneous", "markdown": "https://wpnews.pro/news/hc-rag-evidence-centric-retrieval-augmented-generation-over-heterogeneous.md", "text": "https://wpnews.pro/news/hc-rag-evidence-centric-retrieval-augmented-generation-over-heterogeneous.txt", "jsonld": "https://wpnews.pro/news/hc-rag-evidence-centric-retrieval-augmented-generation-over-heterogeneous.jsonld"}}