cd /news/ai-agents/graph-rag-where-it-actually-breaks · home › topics › ai-agents › article
[ARTICLE · art-142181] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Graph RAG: where it actually breaks

A developer building a Graph RAG system over corporate entity data found that the architecture's failures cluster in two places: query routing between graph traversal and vector search, and the gradual degradation of entity resolution as source data changes. In one case, a query about a Korean chipmaker's supplier network kept returning news articles even though the supplier relationships were already in the graph, because the router never sent the query there; a hand-tuned keyword-based routing function using relational signals such as "supplier" and "subsidiary" cut the miss rate more than any other fix. The developer also reports that entity merges go stale after events like acquisitions, splitting one company across two nodes, and now treats entity quality as continuously degrading rather than a one-time state, with periodic audits of recently active merged pairs.

by read3 min views2 publishedSep 30, 2026

Neo4j with a working schema: two days. Cypher traversal for the relationships I needed: another day or two, once I knew what I was querying. The graph structure, once committed, stayed mostly stable.

That's not where things kept breaking.

When someone asks "what companies are exposed to semiconductor export restrictions?" the system needs to figure out whether to do graph traversal, vector search, or some combination. Getting this wrong doesn't produce an obvious error. It returns something that cites real sources and sounds confident. It just misses context that was sitting in the graph the whole time.

I had one case that stayed with me. A query about a Korean chipmaker's supplier network kept returning news results. The answer was plausible — recent articles about the company, relevant details. It took longer than I'd like to admit to notice that the supplier relationships were already in the graph. The system just wasn't routing that query there. Once I found the gap, the fix was two lines of code. Finding it was the work.

The routing logic I ended up with uses the query's dependency structure as a signal:

def route_query(query: str, graph_searcher, vector_searcher) -> list:
    relational_signals = ["supplier", "subsidiary", "controlled by", 
                          "exposed to", "connected to", "supply chain"]

    needs_graph = any(s in query.lower() for s in relational_signals)

    if needs_graph:
        entity_ids = graph_searcher.find_related(query)
        return vector_searcher.search_scoped(query, entity_ids)
    else:
        return vector_searcher.search(query)

The signal list is hand-tuned and the scoping logic has edge cases. But it reduced the miss rate more than anything else I tried.

The second place things break is entity quality, and this one is more insidious because it's gradual.

The initial entity resolution pass establishes a set of merged entities. After that, the source data keeps changing. A company I had cleanly resolved was acquired six months into the project. Post-acquisition news started using a new parent company name alongside the old subsidiary name. The graph started accumulating both as separate nodes for the same underlying entity.

Graph traversal from the old node gave one picture; traversal from the new node gave a different one. Neither looked wrong in isolation.

MATCH (e:Entity {name: "Old Name"})-[:SUPPLIES]->(s:Entity)
RETURN s.name, s.country

-- If the same entity now also exists as "New Name" post-acquisition,
-- those relationships are on a different node entirely

I treat entity quality as something that degrades continuously now, not a state you establish once. Running a periodic audit of merged pairs helps — checking whether the alias-to-entity mapping still makes sense for recently active entities.

from datetime import date, timedelta

def stale_merges(registry: dict, days_threshold: int = 90) -> list:
    cutoff = date.today() - timedelta(days=days_threshold)
    stale = []
    for entity_id, meta in registry["entities"].items():
        last_verified = meta.get("last_verified")
        if last_verified:
            verified_date = date.fromisoformat(last_verified)
            if verified_date < cutoff and meta.get("alias_count", 0) > 3:
                stale.append(entity_id)
    return stale

The honest version: figuring out which merged pairs have actually gone stale is still mostly manual review. The detection is automatable. What to do with a flagged pair isn't something the code decides.

I build er-api, a multilingual entity resolution service for Korean, Japanese, Chinese, and English corporate data. More at hannune.ai.

── more in #ai-agents 4 stories · sorted by recency
── more on @er-api 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/graph-rag-where-it-a…] indexed:0 read:3min 2026-09-30 · —