{"slug": "stop-looking-for-the-best-way-to-retrieve-context-for-ai-agents-build-a-router", "title": "Stop Looking for the Best Way to Retrieve Context for AI Agents. Build a Router", "summary": "Zenith engineers describe building a retrieval router for AI agents after two production failures: a vector index could not report which version of a document an agent had read when it wrote a pricing tier, and could not trace which published pieces cited a changed source. Citing a token-matched RAG-versus-GraphRAG evaluation on MultiHop-RAG where a plain baseline scored 69.33% versus GraphRAG's 69.01% with Llama 3.1-8B, they argue retrieval should be routed by task shape—passages, exact terms, relationships, live state, and open-ended search—rather than by a single best method.", "body_md": "At Zenith, we build agents that write technical content for engineering-led products, where documentation, APIs, and product specifications change continuously. One of our agents pulled a pricing tier from a product’s documentation and wrote it into a piece. The piece shipped. Three weeks later the source documentation changed, and someone asked what exactly the agent had seen when it wrote it.\n\nThe vector index could not answer. It held the latest embedding of the page and nothing about which version the agent had read or when. Retrieval had worked exactly as designed. The chunk said what the page said. The page no longer said it.\n\nA different question, same corpus: which published pieces cited the page that changed? The index returned three plausible paragraphs about the product and none of the pieces, because nothing in a chunk records that piece P cited source S at version V.\n\nBoth failures get filed under “retrieval quality.” Neither is fixed by a better embedding model, a bigger context window, or a reranker. The first question needed the state of a document at a point in time, with a timestamp attached. The second needed a relationship path. Passages were never going to answer either.\n\nThis is where the GraphRAG pitch usually enters, and it deserves more skepticism than it gets. In [a systematic RAG-versus-GraphRAG evaluation](https://openreview.net/pdf?id=NOK8g6AxRI) on MultiHop-RAG, a token-matched RAG baseline scored 69.33% overall accuracy with Llama 3.1-8B against GraphRAG’s 69.01%. With Llama 3.1-70B the numbers were 71.01% and 71.17%. Give plain retrieval the same token budget and the multi-hop advantage mostly evaporates. The graph is not the answer to every hard question. It is the answer to one specific shape of question.\n\nContext retrieval for AI agents is the step that decides what evidence enters the model’s window before it acts, and it is the retrieval half of [context engineering](https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents). The useful framing is not “which retrieval method is best.” It is “what kind of evidence does this question need,” and the answer varies question by question inside a single agent session. Enterprise context is scattered across documents, tickets, CRMs, databases, APIs, and permission boundaries. A production agent needs several retrieval paths and a router that picks among them.\n\n## **TL;DR**\n\n- **Route retrieval by task shape, not by a single “best” method.** Passages, exact terms, relationships, current state, and unknown paths are five different kinds of evidence. Each has one retrieval path that fits and several that do not.\n- **Hybrid retrieval is the production default for documents.** Vector search alone misses IDs, acronyms, and internal names. Add BM25 before you add anything exotic.\n- **GraphRAG earns its place only when tuned baselines fail on relationships.** Token-matched RAG tied GraphRAG on MultiHop-RAG. Build the graph when relationship paths are the evidence, and after you have measured that hybrid retrieval, reranking, and query decomposition still miss them.\n- **Live state comes from the system of record, with an observed-at timestamp.** No index, vector or graph, should be the source of truth for a ticket status. Agentic search is the fallback for open-ended investigation, and it needs budgets, permissions, and trace evaluation to be safe.\n\n## **Why RAG Fails for AI Agents: “Retrieval Quality” Is the Wrong Bug Report**\n\nThe LLM needs private, relevant, permission-safe context to answer an enterprise question, and that context rarely lives in one clean document set. Everyone agrees on that. The problems start with how most teams first solve it: chunk everything, embed everything, retrieve the top-k, and call the result “the knowledge base.”\n\nChunk-based vector retrieval ranks passages independently. It has no notion that a user belongs to an account, that an account depends on a vendor, or that a vendor has an open exception. Every chunk competes on similarity to the query and nothing else. That is a fine model for “what does the refund policy say about annual contracts” and a broken one for “which accounts are exposed if this vendor goes down.”\n\nContext windows do not rescue this. They are [finite and unevenly attended](https://arxiv.org/abs/2307.03172); dumping every candidate chunk into the prompt raises cost and dilutes attention on the chunks that matter. More context is not more evidence.\n\nFreshness breaks it from the other direction. A policy document stays valid for months. A ticket status or an inventory count changes minute by minute. An index that refreshes nightly is a fine source for the first and a liability for the second, and most pipelines treat both the same way.\n\nThe symptoms are recognizable:\n\n- The agent retrieves plausible chunks that do not answer the question.\n- Answers miss exact IDs, acronyms, or internal terms that a keyword search would have caught.\n- Multi-hop questions turn into manual follow-up searches by the user.\n- Nobody can explain which source, relationship path, or system-of-record state produced the answer, or when it was observed.\n\nEach symptom points at a different missing capability, and no single method supplies all of them. Vector retrieval does not traverse relationships. GraphRAG should not stand in for a live query against fast-changing state. Live queries beat an offline index on currency but are not guaranteed fresh, because replicas lag and caches have TTLs. Agentic search finds paths nobody modeled, at the price of latency, tokens, and nondeterminism.\n\nFive kinds of evidence. One retriever returns one of them well.\n\n## **Five Task Shapes for AI Agent Retrieval, Not One Maturity Ladder**\n\nStart from the question the user asks. Five shapes cover most of what an enterprise agent sees:\n\n- **Direct lookup.** “What does this policy say?” Vector retrieval works when the answer sits in a semantically similar passage.\n- **Exact-term recall.** “Which incidents mention SSO, SCIM, or EUC-17?” Exact strings and semantic meaning both matter, which makes it a hybrid retrieval problem.\n- **Relationship-constrained retrieval.** “Which vendors support products affected by this outage?” Documented entity relationships determine the answer. This is the GraphRAG shape, once tuned baselines have failed on it.\n- **Freshness-sensitive state.** “What is the current status of ticket INC-4821?” This belongs in an authorized query to the system of record, returned with its source, observed-at time, and known consistency limits.\n- **Unknown-path investigation.** “Why did expansion revenue decline in EMEA?” Nobody knows the retrieval path up front. Agentic iterative search is the only path that does not require one.\n\nThese are not stages. A team does not graduate from vector to hybrid to graph to agentic. A single agent fields all five shapes in one afternoon, and it should add a retrieval path only when a task shape demands it.\n\n### **AI Agent Retrieval Approaches Compared**\n\nNot every valid retrieval path is RAG. Some questions should go to tools, APIs, SQL, or systems of record, and calling those “retrieval” is correct even though nothing gets embedded. The table below compares the five paths on what they are good for and what they cost to run.\n\n## **Path 1: Vector RAG for Direct Document Lookup**\n\nVector RAG chunks text, embeds each chunk as a dense vector, and retrieves the nearest semantic neighbors at query time. It is the pattern from the [original RAG paper](https://arxiv.org/abs/2005.11401), and it remains the right first move for an unstructured document corpus.\n\nIt fits direct factual lookups, early prototypes, and internal knowledge assistants where the answer usually lives in one or a few nearby passages. Setup is low, retrieval over large collections is efficient, and it handles users who ask in different words than the source material used. Metadata filters and rerankers bolt on cleanly.\n\nIt fails in predictable ways. Chunks that are too small or poorly segmented lose their surrounding context. Passages that sound relevant come back without the evidence the question needs. And nothing in the design models relationships between entities, so whole-corpus synthesis and multi-hop questions fail unless hybrid retrieval, reranking, query decomposition, and enough matched context supply the connected evidence by other means.\n\nMove on when the agent keeps returning semantically similar but factually disconnected chunks, or when relationship paths rather than passages turn out to be the evidence after you have tuned the baseline.\n\n**Verdict:** The default starting point and still the right path for direct lookup. Most teams leave it too early, before they have added keyword matching.\n\n## **Path 2: Hybrid Retrieval (Vector + BM25) for Production Knowledge Assistants**\n\nHybrid retrieval runs dense vector search and sparse keyword search in the same query and fuses the results. [Weaviate’s hybrid search](https://weaviate.io/developers/weaviate/search/hybrid) is a representative implementation: one query, two rankers, one merged list. The keyword side is usually [BM25](https://doi.org/10.1561/1500000019), which catches exactly the things embeddings blur: ticket IDs, SKUs, acronyms, legal clause numbers, internal project names.\n\n### **Contextual Retrieval: Adding Document Context Before Embedding**\n\n[Contextual Retrieval](https://www.anthropic.com/news/contextual-retrieval), is a preprocessing step that pairs well with it. Before embedding each chunk and before building the BM25 index, you prepend a short, document-specific description of where the chunk sits and what it is about. A chunk that says “the exception was approved for 90 days” becomes a chunk that says “from the Q3 vendor risk review for Vendor B; the exception was approved for 90 days,” and both the vector index and the keyword index get the benefit.\n\nThis is the production default for documentation assistants, and for policy, support, compliance, or product knowledge bases full of acronyms and identifiers. It is the right upgrade for a team that needs better recall but has no measured relationship-traversal failures. Retrieval stays bounded: one vector pass, one keyword pass, fusion, and an optional rerank. Many mature documentation Q&A systems need nothing more.\n\nIts limits are the limits of chunks. Hybrid still treats the passage as the retrieval unit. It does not model organizational hierarchies, supplier networks, ownership chains, or causal dependencies. It needs tuning across chunking, keyword weighting, metadata, and reranking, and contextualized chunks add token cost at indexing time.\n\nMove on when users need explanations of the form “this account is affected because product A depends on vendor B,” and the same relationship failures keep appearing in evaluation after the baseline is tuned.\n\n**Verdict:** If you run exactly one retrieval path over documents, run this one.\n\nAdd BM25 before you add a graph. Many “GraphRAG problems” turn out to be missing-keyword problems.\n\nHybrid retrieval closes the exact-term gap. The relationship gap is a different kind of problem.\n\n## **Path 3: GraphRAG for Multi-Hop and Relationship-Constrained Retrieval**\n\nGraphRAG is a broad label for retrieval that uses entities and relationships as context rather than, or in addition to, passages. [Microsoft GraphRAG](https://microsoft.github.io/graphrag/) is the most influential implementation: it extracts entities and relationships from documents, builds a graph, summarizes clusters into community reports, and supports [local search](https://microsoft.github.io/graphrag/query/overview/) over entity-centered neighborhoods and global search over the community reports.\n\n### **GraphRAG vs. Vector RAG: Key Differences**\n\nThe two approaches differ on more than retrieval unit:\n\nGraphRAG makes sense when relationships are the evidence rather than metadata on the evidence: multi-hop questions across accounts, assets, incidents, vendors, risks, controls, or dependencies, where tuned baselines still miss the connected answer. It also fits corpus-level synthesis, such as themes across a year of incident reports, provided the documented relationships are stable enough to extract and refresh. When it fits, it makes relationships explicit and improves explainability, because the answer can cite entities, edges, and the sources behind them.\n\nThe costs are real and widely underestimated. Full extraction, summarization, and refresh pipelines cost far more than a vector index. Relationship quality is capped by extraction quality, schema design, and source quality. Incremental updates, permissions on traversal, and derived summaries each add operational surface. And none of it suits fast-changing transactional state.\n\nZenith gets relationship questions constantly: which published pieces cite this source, which claims depend on this product version, which reviewer approved the piece that used the now-wrong number. We never built a graph for them. Every event an agent emits carries explicit references, document_id, source_hash, policy_version, parent_event_id, and the edge is derived when a query needs it. Co-occurrence in the same task window is a hint. An explicit reference is evidence. For our workload the derived query is cheaper than a graph that would be stale by the next run.\n\n### **When GraphRAG Does Not Beat Tuned RAG**\n\nDo not assume the graph beats tuned retrieval on multi-hop questions. The MultiHop-RAG numbers above are the caution: 69.33% for token-matched RAG against 69.01% for GraphRAG at 8B, and 71.01% against 71.17% at 70B. Vendor-cited studies report much larger gaps. [Neo4j points to](https://neo4j.com/blog/agentic-ai/vector-rag-vs-graphrag/) a 2026 evaluation by the U.K.’s National Innovation Centre for Data in which GraphRAG answered 65.3% of 510 complex questions against 28.9% for vector-only RAG. The gap between those two results is the baseline. The NICD setup compares against vector-only retrieval; the MultiHop-RAG evaluation gives RAG the same token budget as the graph. Before attributing a gain to graph structure, compare against hybrid retrieval, reranking, query decomposition, and a baseline that gets the same token budget. Build the graph when explicit relationships measurably improve the evidence set, and not before.\n\nMove on when the relationship you need is current system-of-record state rather than extracted document knowledge, or when the graph refresh cycle cannot meet the freshness requirement.\n\n**Verdict:** Powerful for the relationship shape, and a tax everywhere else. Earn it with an evaluation that shows the baseline failing.\n\nBuild the graph when relationship paths are the evidence. Not when GraphRAG is the trend.\n\nEven a perfect graph of documented relationships cannot tell you what is true right now.\n\n## **Path 4: Live-System Queries and Tool Calls for Current Operational State**\n\nDirect live-system queries give the agent [authorized tools](https://manveerc.substack.com/p/mcp-vs-cli-ai-agents) to query the system of record at request time: CRMs, ERPs, ticketing, observability, inventory, internal databases, through APIs, SQL, or approved connectors. Nothing is embedded. Nothing is indexed. The agent asks the source.\n\nThis is the only correct path when the user asks for current status, ownership, inventory, entitlement, balance, incident severity, or ticket progress, and the source-system query can meet the freshness SLO. The authoritative answer already exists in a structured system, and copying it into a secondary index only adds a place for it to go stale. When the connector passes user identity, roles, or scopes through, the system of record enforces its own access model, which is the one you trust anyway.\n\nThe catch is in the word “live.” Live does not mean guaranteed fresh. APIs and SQL tools return lagged replicas, cached responses, partial pagination, and eventually consistent results. A tool that returns a status should return the source, the observed-at timestamp, and the known consistency limits alongside it, so the model can reason about staleness instead of assuming none. Beyond that, the agent has to pick the right tool and formulate the right query, rate limits shape response time, and structured systems have no answer for semantic search over long unstructured text.\n\nDocumented knowledge is not automatically the vector shape. The number in the opening story lived in a documentation page, which looks like Path 1. But the page changed faster than the index refreshed, and someone later needed to know what the agent had seen, which makes it Path 4 wearing a document’s clothing. The question that decides the route is not “is this a document” but “does this fact change faster than my index does, and will anyone ask afterward what the agent saw.” At Zenith the answer to the second half is always yes, so every retrieved source now carries two timestamps: when the fact was true in the world, and when our system recorded it. A current-state query returns the corrected value. The second timestamp is what lets us return the value the agent actually had. <!-- link “two timestamps” to the ClickHouse post when it is live -->\n\nMove on when the answer needs current state and historical documents, policies, or relationship-aware context in the same breath.\n\n**Verdict:** Mandatory for mutable state. The failure mode is not using it, and letting an index answer a question only the source can.\n\nA vector index is a search tool, not a state store.\n\nFour paths, each for a known shape. The fifth is for the questions whose shape nobody knows yet.\n\n## **Path 5: Agentic RAG and Iterative Search for Open-Ended Investigation**\n\nAgentic iterative search lets the model search, inspect what came back, decide what is missing, and issue follow-up searches or tool calls until it has a reasoning chain. Most vendor material calls this agentic RAG. The label matters less than the loop. It can combine vector search, keyword search, graph traversal, web search, SQL, and APIs in whatever order the investigation demands, and it does not require every useful path to be modeled in advance.\n\nThat makes it the right path for investigative work where the retrieval path is unknown up front: deep research, due diligence, root-cause analysis, exploratory business questions like the EMEA revenue decline. It assumes latency is acceptable because the user values completeness over speed.\n\nThe costs stack per hop. Each iteration adds latency and tokens. Each hop forces a query-selection decision, an evidence-preservation decision, and a stopping decision, which is precisely the work an explicit graph front-loads into indexing. And without step, token, tool-call, and tool-scope budgets, the loop can wander, drift from the goal, or over-search indefinitely.\n\nMove on when the same multi-hop path is requested often enough to deserve a graph or a dedicated tool, or when the workflow needs predictable, governed retrieval rather than exploration.\n\n**Verdict:** The right fallback for the unknown shape, and a bad default for the known ones. If you see the same investigation twice, model it.\n\n## **How to Choose an AI Agent Retrieval Path: Start From the Question, Not the Technology**\n\nPut the five shapes next to real questions and the routing mostly writes itself:\n\n### **Five Questions to Ask Before Routing a Query**\n\nWhere it does not write itself, five questions settle it:\n\n- **Is this documented knowledge or current system-of-record state?** Extracted or documented relationships go to the graph. Authoritative operational state goes to a live query.\n- **How fresh must the answer be?** For mutable facts, prefer the system of record so the agent can cite the exact timestamp of what it observed.\n- **What is the cost and latency tolerance?** Vector and hybrid run one bounded pass. GraphRAG adds construction and refresh work, plus query-time traversal or report synthesis. Agentic search keeps calling tools until it hits a stopping condition.\n- **What permissions apply before retrieval?** The router has to respect access controls so users retrieve only the documents and relationships they are authorized to see.\n- **How confident is the router?** Low-confidence routes should fall back to the safer path, ask a clarifying question, or run multiple retrievers under a strict budget.\n\nIn code, the router is smaller than the argument around it:\n\n``` python\ndef retrieve(query, user):\n    shape, confidence = classify(query)   # lookup | exact_term | relationship | live_state | unknown\n    scope = authorize(user)               # applied before retrieval, never after\n\n    if shape == \"lookup\":\n        evidence = vector.search(query, scope)\n    elif shape == \"exact_term\":\n        evidence = hybrid.search(query, scope)            # vector + BM25, fused\n    elif shape == \"relationship\":\n        evidence = graph.traverse(query, scope)\n    elif shape == \"live_state\":\n        evidence = tools.query(query, scope)              # returns source + observed_at\n    else:\n        evidence = agentic.search(query, scope, budget=Budget(steps=8, tool_calls=20))\n\n    if confidence < 0.6:\n        evidence = hybrid.search(query, scope) + evidence  # fall back to the safer path\n\n    return evidence  # every item carries source, observed_at, and the scope it was retrieved under\n```\n\nThe shape classifier is the only new component. Everything else is a retriever you already have, called with the user’s scope attached and asked to say where and when it got what it returned.\n\nThe router is the architecture. The retrievers are plumbing.\n\n## **Production Requirements for AI Agent Context Retrieval**\n\nA routed retrieval stack has a recognizable shape. It is also where context engineering stops being a prompt concern and becomes infrastructure. An authenticated query enters a classifier or router. Authorization and scoping controls run before any retrieval. The request fans out to a vector/BM25 retriever, a graph traversal layer, a live-system connector, or some combination. Results pass through a reranker or context compressor, into the LLM, and out through logging and evaluation. Four of those boxes are where production systems break.\n\n### **Permission-Aware Retrieval and Access Control**\n\nRetrieval must respect document-level, row-level, tenant-level, and role-based permissions, and it must do so before unauthorized evidence enters retrieval, traversal, summarization, logs, or generation. Post-filtering is acceptable only when an independent enforcement boundary backs it and it has been tested against leakage. Graph traversal, relationship paths, and derived summaries need their own checks, because a community report can reveal a restricted relationship without ever quoting a restricted document. Pre-filter metadata has to stay current; stale permission metadata is an open door. Live connectors should [pass user identity, roles, or scopes](https://manveerc.substack.com/p/mcp-supply-chain-attack-vector) into the system of record wherever it supports that.\n\n### **Tenant Isolation for Enterprise Retrieval**\n\nPick the boundary that matches the threat model. Physical or database-level isolation separates customers or environments. Logical collection scoping partitions data inside an already-authorized boundary. Shared indexes work only with careful metadata partitioning and pre-retrieval filters, and they are the option most likely to leak under pressure.\n\n### **Freshness, Deletion, and Observed-at Timestamps**\n\nPrivacy and retention requirements mean targeted deletion for vectors, nodes, edges, documents, and derived summaries, not just source documents. Fast-changing state should stay in the system of record unless a separate indexed copy is explicitly justified. Any tool that returns a mutable fact should return the source, the observed-at timestamp, and the known consistency limits, and the team should define freshness SLOs, because live APIs and SQL do not promise strict consistency. Store the embedding model, model version, vector dimension, and source record ID with every embedding. Otherwise a model upgrade silently mixes incompatible vector spaces in one index, and retrieval quality drops without a single error in the logs.\n\n### **Logging and Retrieval Evaluation**\n\nLog router decisions, retrieved context IDs, permission-filter outcomes, latency, token cost, observed-at times, and answer-quality signals. Do not log restricted source content into a system with broader access than the source. Then evaluate retrieval on whether it supplied the right evidence, not on whether the answer sounded plausible: context precision and context recall in the [Ragas](https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/context_precision/) sense, and groundedness of the answer against the retrieved context, which [Azure AI Foundry](https://learn.microsoft.com/en-us/azure/ai-foundry/concepts/evaluation-metrics-built-in) ships as a built-in metric.\n\nAt Zenith, every trajectory where an output missed the technical accuracy bar on the first pass becomes an evaluation fixture, with the exact sources, tool outputs, and intermediate decisions from the original run. We stopped inventing synthetic examples for failures we had already observed. Reviewing those slices by task shape is what makes the router improve.\n\nTwo habits keep the stack honest. Track token economics on both ends, because GraphRAG spends at indexing time and agentic search spends at query time, and the two bills look nothing alike. And review failed queries by task shape, so the router improves instead of the team reaching for the next retrieval method. The two mistakes this whole piece has been circling are the opposite habits: building a graph because GraphRAG is popular rather than because relationships are the measured bottleneck, and using agentic search as a substitute for a predictable path you could have built.\n\n## **Vector RAG vs. GraphRAG Was Never the Question**\n\nBack to the piece with the wrong number. The retrieval was never bad. The chunk was a faithful copy of the page on the day the agent read it. What the question needed was the state of a page at a moment in time, with a timestamp, and a path back to every piece that cited it. A vector index was asked for both, answered confidently, and could supply neither, not because the embeddings were weak but because nobody asked what kind of evidence the question required.\n\nVector and hybrid retrieval cover passages and exact terms well. GraphRAG covers documented relationships, once you have proven a tuned baseline cannot. The system of record covers current state and nothing else should. Agentic search covers the unknown path at a price worth paying only when the path is truly unknown.\n\nSo before the next retrieval migration, measure the failures you have: missing evidence, stale answers, latency, permission leaks, hallucinations. Each one points at a task shape, and each task shape points at a path. Start with the simplest path that satisfies the question. Add the next one only when the failure mode is clear and the baseline has been tuned.\n\nWhich of your agent’s failures last week were retrieval failures, and which were routing failures that nobody classified?\n\nThere is no best retrieval method. There is a best retrieval method for this question.\n\n## **FAQ: Context Retrieval for AI Agents**\n\n### **What is the best way to retrieve context for AI agents?**\n\nRoute retrieval by task shape. Use vector or hybrid retrieval for document lookup, GraphRAG for relationship-heavy questions, live-system queries for current operational state, and agentic search for open-ended investigation.\n\n### **Is GraphRAG better than vector RAG?**\n\nNot by default. In a systematic evaluation on MultiHop-RAG, token-matched RAG scored 69.33% against GraphRAG’s 69.01% with Llama 3.1-8B, and 71.01% against 71.17% with 70B. GraphRAG wins when explicit relationships are the evidence and the baseline has been given the same token budget.\n\n### **When should I use GraphRAG instead of vector RAG?**\n\nUse GraphRAG when explicit relationships between entities determine the answer and tuned hybrid retrieval, reranking, query decomposition, and token-matched baselines still miss the needed evidence.\n\n### **Should GraphRAG replace live-system queries for AI agents?**\n\nNo. GraphRAG is useful for documented relationships, but current operational state should usually come from authorized live-system queries with source, observed-at time, and consistency limits.\n\n### **What is hybrid retrieval with BM25?**\n\nHybrid retrieval runs dense vector search and sparse BM25 keyword search in one query and fuses the rankings. The vector side matches meaning; BM25 matches exact strings such as IDs, acronyms, SKUs, and clause numbers that embeddings blur.", "url": "https://wpnews.pro/news/stop-looking-for-the-best-way-to-retrieve-context-for-ai-agents-build-a-router", "canonical_source": "https://manveerc.substack.com/p/context-retrieval-ai-agent", "published_at": "2026-10-09 10:16:35+00:00", "updated_at": "2026-10-09 10:22:37.522452+00:00", "lang": "en", "topics": ["ai-agents", "ai-research", "large-language-models", "ai-tools"], "entities": ["Zenith", "GraphRAG", "MultiHop-RAG", "Llama 3.1-8B", "Llama 3.1-70B", "Anthropic"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/stop-looking-for-the-best-way-to-retrieve-context-for-ai-agents-build-a-router", "markdown": "https://wpnews.pro/news/stop-looking-for-the-best-way-to-retrieve-context-for-ai-agents-build-a-router.md", "text": "https://wpnews.pro/news/stop-looking-for-the-best-way-to-retrieve-context-for-ai-agents-build-a-router.txt", "jsonld": "https://wpnews.pro/news/stop-looking-for-the-best-way-to-retrieve-context-for-ai-agents-build-a-router.jsonld"}}