cd /news/artificial-intelligence/show-hn-hubmesh-multi-hop-rag-retrie… · home topics artificial-intelligence article
[ARTICLE · art-86997] src=github.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Show HN: Hubmesh – Multi-hop RAG retrieval with zero LLM calls in the query path

Hubmesh, a new Python library, improves multi-hop RAG retrieval by adding a centrality-aware planner over existing vector databases, eliminating LLM calls in the query path. The library uses multi-component seed selection and budget-aware context packing to enhance retrieval quality, as demonstrated in its Show HN launch.

read6 min views1 publishedAug 5, 2026
Show HN: Hubmesh – Multi-hop RAG retrieval with zero LLM calls in the query path
Image: source

Centrality-aware GraphRAG retrieval planner. Drop-in layer over any vector DB.

hubmesh

is a Python library that improves multi-hop RAG quality on top of an existing vector database. You don't replace your infrastructure — you add a smart planner between your vector DB and your LLM.

Naive vector retrieval ("embed query, get top-k by cosine similarity") fails on multi-hop questions like "Where was the founder of the company that acquired Slack born?" The correct answer requires retrieving entities along a reasoning path, not the single most similar item.

GraphRAG and HippoRAG showed that running a small Personalized PageRank over a knowledge graph at query time can substantially improve multi-hop retrieval. hubmesh

extends that line with two contributions:

Multi-component seed selection. Instead of picking PPR seeds by raw query similarity (which picks wrong-community seeds at high feature overlap), seeds are chosen by a multi-component score combining query relevance, structural fit, and coverage diversity.Budget-aware context packing. Once relevant entities are scored, pack them into the LLM's context window with explicit coverage and redundancy control rather than just truncating top-k.

The multi-component scoring pattern is adapted from the NNSI framework (Naidu Dsk, ICOMP'25 — to appear) for SDN topology optimization, repurposed here for retrieval planning.

from hubmesh import Planner
from hubmesh.adapters import InMemoryStore

embed = ...   # callable: text -> np.ndarray
docs = [...]  # list of Document or strings or dicts

store = InMemoryStore.from_documents(docs, embed=embed)
planner = Planner(store=store, embed=embed)
result = planner.retrieve(query="...", top_k=10, budget_tokens=4000)
python
from hubmesh import Planner
from hubmesh.adapters import QdrantStore

store = QdrantStore.from_documents(docs)                          # in-memory
store = QdrantStore.from_documents(docs, path="./qdrant_data")    # on-disk
store = QdrantStore.from_documents(docs, url="http://localhost:6333")  # remote

planner = Planner(store=store, embed=embed)
result = planner.retrieve(query="...", top_k=10)
python
from hubmesh.adapters import ChromaStore

store = ChromaStore.from_documents(docs)                          # ephemeral
store = ChromaStore.from_documents(docs, persist_directory="./chroma_data")
store = ChromaStore.from_documents(docs, host="localhost", port=8000)
python
from hubmesh.kg import build_entity_kg
import spacy

nlp = spacy.load("en_core_web_sm")
kg = build_entity_kg(docs, nlp=nlp)

planner = Planner(store=store, kg=kg, nlp=nlp)
result = planner.retrieve(query="Where was the founder of the company that bought Slack born?",
                          top_k=10, budget_tokens=4000)

for path in result.reasoning:
    print(f"  score={path.score:.3f}  {' → '.join(path.node_ids)}")
python
from hubmesh.kg_llm import build_entity_kg_llm
from hubmesh.entity_linker import EmbeddingLinker, make_st_embedder

def llm(prompt):  # provider-agnostic — bring your own
    return your_llm_call(prompt)

kg = build_entity_kg_llm(docs, llm=llm, cache_path="kg_cache.json")

kg = build_entity_kg_llm(docs, llm=llm, cache_path="kg_cache.json",
                         linker=EmbeddingLinker(embed=make_st_embedder()))

planner = Planner(store=store, kg=kg)
python
from hubmesh.kg import build_entity_kg
from hubmesh.entity_linker import EmbeddingLinker, make_st_embedder

linker = EmbeddingLinker(embed=make_st_embedder(), threshold=0.82)
kg = build_entity_kg(docs, linker=linker)
r1 = planner.retrieve(query=question, top_k=5)

r2 = planner.retrieve(
    query=question, top_k=5,
    seed_entities=["Nimbus Analytics"],           # merged with the query's own seeds
    exclude_docs=[s.doc.id for s in r1.sources],  # don't re-retrieve consumed docs
)

Seed mentions resolve through the alias index, so free-text entity names work. The query path stays deterministic and LLM-free — the planning intelligence lives in the caller.

pip install "hubmesh[mcp]"
python -m spacy download en_core_web_sm
{"mcpServers": {"hubmesh": {"command": "hubmesh-mcp"}}}

Exposes the planner as deterministic operator tools over stdio — index_corpus

, retrieve

(seed-steerable, as above), resolve_entities

, entity_neighbors

, path_between

, get_document

, graph_stats

, list_corpora

. Your agent is the solver: it decomposes the question, reads each hop, and aims the next one; the server answers in milliseconds with zero LLM calls. Corpora persist as plain JSON/NPZ under ~/.hubmesh/corpora

.

The server warms up models and persisted corpora in the background at launch (~5-10s on first run), so tool calls stay fast from the start — relevant for strict-timeout connector clients (Perplexity, etc.).

For web-based connector clients, serve SSE natively — no gateway process needed:

hubmesh-mcp --transport sse --port 8000 --allow-tunnel
ngrok http 8000     # paste https://<your-url>/sse into the connector

Tunnel field notes (from a live Perplexity integration): ngrok works (free tier included); cloudflared quick tunnels buffer SSE bodies and hang tool calls; supergateway is unnecessary here and crashes on reconnect. --allow-tunnel

accepts the tunnel's forwarded Host header — without it, proxied requests get 421 Misdirected Request.

Full field report — setup, error decoder, a 9/9 test battery run through Perplexity, and two findings about reasoning-model behaviour — in docs/perplexity.md.

from hubmesh import chunk_by_sentences, chunk_documents

chunks = chunk_documents(
    [{"id": "doc1", "text": long_text}, ...],
    strategy="sentences", target_tokens=200,
)
pip install hubmesh                   # core
pip install "hubmesh[qdrant]"         # Qdrant adapter
pip install "hubmesh[chroma]"         # Chroma adapter
pip install "hubmesh[kg]"             # entity-linked KG (spaCy)
pip install "hubmesh[linker]"         # embedding-based entity linker
pip install "hubmesh[all]"            # everything
python -m spacy download en_core_web_sm   # required for KG mode
query → first-pass ANN  → induced subgraph → multi-component scoring
                              ↓                        ↓
                       community anchoring → Personalized PageRank
                              ↓                        ↓
                              └─────→ ranking → budget-aware packing → context

Each layer is independently testable and replaceable. Adapters wrap your existing vector DB so you don't have to migrate.

Headline: on multi-hop QA, hubmesh's KG mode beats both naive cosine retrieval and a HippoRAG-style PPR-only ablation that uses the same KG, at every hop depth.

Benchmark Setting recall@10 vs naive
HotpotQA dev, N=7405 (full)
KG mode +5.90 pts
HotpotQA dev, N=500 KG mode +5.0 pts
MuSiQue dev, N=300, 2-hop KG mode +6.0 pts
MuSiQue dev, N=300, 3-hop KG mode +3.2 pts
MuSiQue dev, N=300, 4-hop KG mode +5.0 pts

All rows measured with v0.4.0 defaults (alias-indexed seeds + NNSI-KG convergence; ablation JSONs committed in benchmarks/

). Disclosed: convergence trades top-rank precision for depth recall — recall@2 is −0.75 pts vs naive on full dev (dips ≤0.5 at smaller n); if you retrieve with top_k=2

, set use_convergence=False

. Multi-seed queries cost ~1.5–1.8× (still zero LLM tokens, deterministic).

vs PPR-only ablation on the same KG: +29.8 pts on HotpotQA at N=500 (measured on v0.2.0) — the multi-component scoring is doing the work, not just "having a graph."

On the full N=7405 HotpotQA dev: hubmesh hits 75.2% supporting-fact recall@10 vs naive cosine's 69.3% (+4.21 pts at recall@5; recall@2 −0.75, disclosed above).

Latency: ~22 ms mean / 26 ms p95 per query on a 7K-node KG (after PPR matrix caching); ~3 s/query at the 66K-paragraph full-dev scale with v0.4 convergence on.

See BENCHMARKS.md for the full methodology, ablations, per-hop breakdown, and notes on what this proves and doesn't.

Reproduce:

python benchmarks/run_hotpotqa.py --n 500 --kg
python benchmarks/run_musique.py  --n 300 --kg
python benchmarks/profile_query.py        # latency profile

Pre-alpha (v0.4.0). Core algorithms implemented and validated; adapters for in-memory, Qdrant, and Chroma; entity-linked KG with both spaCy NER and LLM-based extraction (both linker-aware); alias-indexed entity resolution; NNSI-KG scoring (multi-source convergence default-on, hub-discounted PPR opt-in); agent-driven iterative multi-hop via seed_entities

/ exclude_docs

; MCP operator server (hubmesh-mcp

, native SSE) with JSON/NPZ corpus persistence; document chunking; reasoning-path explanation; PPR-cache latency optimisation. Pinecone / pgvector / Weaviate adapters and additional multi-hop benchmarks are tracked as good first issues.

The multi-component scoring pattern is adapted from the Network Node Significance Index (NNSI) framework introduced in Naidu Dsk, "A Framework for Improving Network Topology Based on Graph Theory in Software-Defined Networking", 26th International Conference on Internet Computing & IoT (ICOMP'25), Las Vegas, July 2025 — proceedings to appear. Repurposed here from SDN topology optimization to retrieval planning.

MIT

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @hubmesh 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/show-hn-hubmesh-mult…] indexed:0 read:6min 2026-08-05 ·