cd /news/artificial-intelligence/tree-navigation-without-llm-summarie… · home › topics › artificial-intelligence › article
[ARTICLE · art-146566] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Tree Navigation Without LLM Summaries: A Matched-Cost Study of Hierarchical Retrieval for Long-Document QA

A new arXiv paper (arXiv:2610.06902v1) introduces NavTree, a leaves-only hierarchical retriever that builds a deterministic balanced segment tree over document chunks with zero language-model calls at indexing and uses the tree purely as a navigation scaffold via a hybrid lexical-and-dense frontier walk. On a matched-cost evaluation against flat retrievers and an extractive re-implementation of RAPTOR, NavTree was the strongest matched-cost hierarchical retriever in the evaluated grid and tied the strongest flat baseline, and on long-document multi-hop QA it was the only hierarchical method that significantly beat BM25 on a class-vs-class basis. A matched-reader replication of the published abstractive RAPTOR variant, given strong cluster summaries, still lost to NavTree at every multi-chunk budget at zero indexing cost, and the ranking carried across stronger and open-weight readers, a stronger encoder, and a full factorial isolating leaves-only emission as the structural lever.

by read1 min views1 publishedOct 7, 2026

arXiv:2610.06902v1 Announce Type: new Abstract: Retrieval-augmented generation grounds language models in external context, but for long documents flat top-$k$ retrieval can cluster on a single region and miss complementary evidence. RAPTOR-style summary trees address this by recursively clustering chunks and using a language model to summarize each cluster at indexing time, then ranking summary nodes alongside raw chunks at query time. We show the main benefit of summary trees in long-document QA can come from navigation rather than the generated summary content. We introduce NavTree, a leaves-only retriever that builds a deterministic balanced segment tree over chunks (zero language-model calls at indexing) and uses the tree purely as a navigation scaffold: a hybrid lexical-and-dense frontier walk, anchored on top retrieved leaves, descends from the root and emits only leaf chunks to the reader. On a matched-cost evaluation against flat retrievers and an extractive re-implementation of RAPTOR, NavTree is the strongest matched-cost hierarchical retriever in our evaluated grid and ties the strongest flat baseline. On long-document multi-hop QA, it is the only hierarchical method that significantly beats BM25 on a class-vs-class basis, corroborated by a reader-free retrieval-recall check. A matched-reader replication of the published abstractive RAPTOR variant, given strong cluster summaries, still loses to NavTree at every multi-chunk budget, at zero indexing cost. The ranking carries across stronger and open-weight readers, a stronger encoder, and a full factorial that isolates leaves-only emission as the structural lever.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @navtree 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/tree-navigation-with…] indexed:0 read:1min 2026-10-07 · —