When do AI agents actually matter? Benchmarking RAG vs GraphRAG vs Agentic GraphRAG on TigerGraph A developer built three question-answering pipelines β€” RAG, GraphRAG, and Agentic GraphRAG β€” on TigerGraph Savanna to test when agentic loops actually beat simpler retrieval, using ~2,900 Wikipedia Olympic articles and 100 evaluation questions. GraphRAG reached 92% accuracy at roughly a fifth of RAG's token cost (247 vs 1,573 tokens per query), while the agentic pipeline hit 100% by adding a disambiguation step that lifted multi-hop accuracy from 71% to 100%. The developer concluded agents only pay off where evidence is ambiguous and needs checking, since a well-designed graph query was sufficient and far cheaper for most questions. Everyone's building agents right now. But a better question than "can an agent do this?" is "when does an agent actually beat something simpler?" For the TigerGraph Agentic GraphRAG Hackathon, I built three question-answering pipelines side by side to find out. The dataset was ~2,900 Wikipedia articles about Olympic events, plus 100 evaluation questions with answers and 50 hidden ones. The questions came in five types: simple lookups, multi-hop questions "who won gold at this venue on this date?" , temporal ones "…at the Olympics held immediately before 2016" , aggregations "how many biathlon events had more than 73 competitors?" , and superlatives "which event had the most competitors?" . One early discovery shaped everything: about 760 of the documents were distractors β€” articles about films and companies mixed into the corpus. A plain text-retrieval system can get pulled toward them. A graph built only from Olympic event records never sees them. I parsed each event's Wikipedia infobox into a structured Event vertex covering sport, year, season, venue, date, competitor count, nations and medallists, and loaded 2,187 of them into TigerGraph Savanna . Then I wrote GSQL query endpoints for filtering, counting and lookups. The key design decision was that the model never computes answers. A question becomes a structured plan, and TigerGraph executes it. Counting 11 biathlon events and filtering by competitor count is a database job. An LLM guessing at it from a few retrieved paragraphs is how you get wrong answers. | Pipeline | Accuracy | Est. tokens/query | |---|---|---| | RAG | 18% | 1,573 | | GraphRAG | 92% | 247 | | Agentic GraphRAG | 100% | 1,295 | RAG scored 0% on aggregation and superlative questions. You simply can't count 40 events from 5 retrieved documents. It also often picked the wrong Olympic year for "before 2016" questions, because "2012" and "1992" look equally similar to a text retriever. GraphRAG took the big leap to 92% at roughly a fifth of RAG's token cost. The final 8 points came from multi-hop questions. Take "who won gold at Beijing National Stadium on 16 August 2008". A venue hosts dozens of events, so GraphRAG grabbed the first match. The agent noticed several candidates, treated that as a gap, and disambiguated using the exact date in each event's record, with a similarity-search tiebreak for true ties. That one investigative behaviour took multi-hop from 71% to 100%. That's the real lesson. Agents aren't better everywhere. For most of these questions, a well-designed graph query was enough and was far cheaper. The agentic loop paid off exactly where the evidence was ambiguous and needed checking. /gsql/v1/tokens , not the older /restpp/requesttoken . Code: https://github.com/repulsortechnologies-alt/agentic-graphrag https://github.com/repulsortechnologies-alt/agentic-graphrag