{"slug": "tree-navigation-without-llm-summaries-a-matched-cost-study-of-hierarchical-for", "title": "Tree Navigation Without LLM Summaries: A Matched-Cost Study of Hierarchical Retrieval for Long-Document QA", "summary": "A new arXiv paper (arXiv:2610.06902v1) introduces NavTree, a leaves-only hierarchical retriever that builds a deterministic balanced segment tree over document chunks with zero language-model calls at indexing and uses the tree purely as a navigation scaffold via a hybrid lexical-and-dense frontier walk. On a matched-cost evaluation against flat retrievers and an extractive re-implementation of RAPTOR, NavTree was the strongest matched-cost hierarchical retriever in the evaluated grid and tied the strongest flat baseline, and on long-document multi-hop QA it was the only hierarchical method that significantly beat BM25 on a class-vs-class basis. A matched-reader replication of the published abstractive RAPTOR variant, given strong cluster summaries, still lost to NavTree at every multi-chunk budget at zero indexing cost, and the ranking carried across stronger and open-weight readers, a stronger encoder, and a full factorial isolating leaves-only emission as the structural lever.", "body_md": "arXiv:2610.06902v1 Announce Type: new \nAbstract: Retrieval-augmented generation grounds language models in external context, but for long documents flat top-$k$ retrieval can cluster on a single region and miss complementary evidence. RAPTOR-style summary trees address this by recursively clustering chunks and using a language model to summarize each cluster at indexing time, then ranking summary nodes alongside raw chunks at query time. We show the main benefit of summary trees in long-document QA can come from navigation rather than the generated summary content. We introduce NavTree, a leaves-only retriever that builds a deterministic balanced segment tree over chunks (zero language-model calls at indexing) and uses the tree purely as a navigation scaffold: a hybrid lexical-and-dense frontier walk, anchored on top retrieved leaves, descends from the root and emits only leaf chunks to the reader. On a matched-cost evaluation against flat retrievers and an extractive re-implementation of RAPTOR, NavTree is the strongest matched-cost hierarchical retriever in our evaluated grid and ties the strongest flat baseline. On long-document multi-hop QA, it is the only hierarchical method that significantly beats BM25 on a class-vs-class basis, corroborated by a reader-free retrieval-recall check. A matched-reader replication of the published abstractive RAPTOR variant, given strong cluster summaries, still loses to NavTree at every multi-chunk budget, at zero indexing cost. The ranking carries across stronger and open-weight readers, a stronger encoder, and a full factorial that isolates leaves-only emission as the structural lever.", "url": "https://wpnews.pro/news/tree-navigation-without-llm-summaries-a-matched-cost-study-of-hierarchical-for", "canonical_source": "https://arxiv.org/abs/2610.06902", "published_at": "2026-10-07 04:00:00+00:00", "updated_at": "2026-10-07 04:19:24.779460+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research", "natural-language-processing"], "entities": ["NavTree", "RAPTOR", "arXiv", "BM25"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/tree-navigation-without-llm-summaries-a-matched-cost-study-of-hierarchical-for", "markdown": "https://wpnews.pro/news/tree-navigation-without-llm-summaries-a-matched-cost-study-of-hierarchical-for.md", "text": "https://wpnews.pro/news/tree-navigation-without-llm-summaries-a-matched-cost-study-of-hierarchical-for.txt", "jsonld": "https://wpnews.pro/news/tree-navigation-without-llm-summaries-a-matched-cost-study-of-hierarchical-for.jsonld"}}