# When do AI agents actually matter? Benchmarking RAG vs GraphRAG vs Agentic GraphRAG on TigerGraph

> Source: <https://dev.to/utkarsh_varshney_0c1e8ad0/when-do-ai-agents-actually-matter-benchmarking-rag-vs-graphrag-vs-agentic-graphrag-on-tigergraph-1oke>
> Published: 2026-10-03 22:01:05+00:00

Everyone's building agents right now. But a better question than "can an agent do this?" is **"when does an agent actually beat something simpler?"** For the TigerGraph Agentic GraphRAG Hackathon, I built three question-answering pipelines side by side to find out.

The dataset was ~2,900 Wikipedia articles about Olympic events, plus 100 evaluation questions with answers and 50 hidden ones. The questions came in five types: simple lookups, multi-hop questions ("who won gold at this venue on this date?"), temporal ones ("…at the Olympics held immediately before 2016"), aggregations ("how many biathlon events had more than 73 competitors?"), and superlatives ("which event had the most competitors?").

One early discovery shaped everything: about 760 of the documents were distractors — articles about films and companies mixed into the corpus. A plain text-retrieval system can get pulled toward them. A graph built only from Olympic event records never sees them.

I parsed each event's Wikipedia infobox into a structured `Event` vertex covering sport, year, season, venue, date, competitor count, nations and medallists, and loaded 2,187 of them into **TigerGraph Savanna**. Then I wrote GSQL query endpoints for filtering, counting and lookups.

The key design decision was that the model never computes answers. A question becomes a structured plan, and TigerGraph executes it. Counting 11 biathlon events and filtering by competitor count is a database job. An LLM guessing at it from a few retrieved paragraphs is how you get wrong answers.

| Pipeline | Accuracy | Est. tokens/query | 
|---|---|---|
| RAG | 18% | 1,573 | 
| GraphRAG | 92% | 247 | 
| Agentic GraphRAG | 100% | 1,295 | 

RAG scored 0% on aggregation and superlative questions. You simply can't count 40 events from 5 retrieved documents. It also often picked the wrong Olympic year for "before 2016" questions, because "2012" and "1992" look equally similar to a text retriever.

GraphRAG took the big leap to 92% at roughly a fifth of RAG's token cost.

The final 8 points came from multi-hop questions. Take "who won gold at Beijing National Stadium on 16 August 2008". A venue hosts dozens of events, so GraphRAG grabbed the first match. The agent noticed several candidates, treated that as a gap, and disambiguated using the exact date in each event's record, with a similarity-search tiebreak for true ties. That one investigative behaviour took multi-hop from 71% to 100%.

That's the real lesson. **Agents aren't better everywhere.** For most of these questions, a well-designed graph query was enough and was far cheaper. The agentic loop paid off exactly where the evidence was ambiguous and needed checking.

`/gsql/v1/tokens`, not the older `/restpp/requesttoken`.
**Code:** [https://github.com/repulsortechnologies-alt/agentic-graphrag](https://github.com/repulsortechnologies-alt/agentic-graphrag)
