Neo4j with a working schema: two days. Cypher traversal for the relationships I needed: another day or two, once I knew what I was querying. The graph structure, once committed, stayed mostly stable.
That's not where things kept breaking.
When someone asks "what companies are exposed to semiconductor export restrictions?" the system needs to figure out whether to do graph traversal, vector search, or some combination. Getting this wrong doesn't produce an obvious error. It returns something that cites real sources and sounds confident. It just misses context that was sitting in the graph the whole time.
I had one case that stayed with me. A query about a Korean chipmaker's supplier network kept returning news results. The answer was plausible — recent articles about the company, relevant details. It took longer than I'd like to admit to notice that the supplier relationships were already in the graph. The system just wasn't routing that query there. Once I found the gap, the fix was two lines of code. Finding it was the work.
The routing logic I ended up with uses the query's dependency structure as a signal:
def route_query(query: str, graph_searcher, vector_searcher) -> list:
relational_signals = ["supplier", "subsidiary", "controlled by",
"exposed to", "connected to", "supply chain"]
needs_graph = any(s in query.lower() for s in relational_signals)
if needs_graph:
entity_ids = graph_searcher.find_related(query)
return vector_searcher.search_scoped(query, entity_ids)
else:
return vector_searcher.search(query)
The signal list is hand-tuned and the scoping logic has edge cases. But it reduced the miss rate more than anything else I tried.
The second place things break is entity quality, and this one is more insidious because it's gradual.
The initial entity resolution pass establishes a set of merged entities. After that, the source data keeps changing. A company I had cleanly resolved was acquired six months into the project. Post-acquisition news started using a new parent company name alongside the old subsidiary name. The graph started accumulating both as separate nodes for the same underlying entity.
Graph traversal from the old node gave one picture; traversal from the new node gave a different one. Neither looked wrong in isolation.
MATCH (e:Entity {name: "Old Name"})-[:SUPPLIES]->(s:Entity)
RETURN s.name, s.country
-- If the same entity now also exists as "New Name" post-acquisition,
-- those relationships are on a different node entirely
I treat entity quality as something that degrades continuously now, not a state you establish once. Running a periodic audit of merged pairs helps — checking whether the alias-to-entity mapping still makes sense for recently active entities.
from datetime import date, timedelta
def stale_merges(registry: dict, days_threshold: int = 90) -> list:
cutoff = date.today() - timedelta(days=days_threshold)
stale = []
for entity_id, meta in registry["entities"].items():
last_verified = meta.get("last_verified")
if last_verified:
verified_date = date.fromisoformat(last_verified)
if verified_date < cutoff and meta.get("alias_count", 0) > 3:
stale.append(entity_id)
return stale
The honest version: figuring out which merged pairs have actually gone stale is still mostly manual review. The detection is automatable. What to do with a flagged pair isn't something the code decides.
I build er-api, a multilingual entity resolution service for Korean, Japanese, Chinese, and English corporate data. More at hannune.ai.