{"slug": "when-does-an-agent-beat-rag-we-benchmarked-three-pipelines-on-a-tigergraph-graph", "title": "When does an agent beat RAG? We benchmarked three pipelines on a TigerGraph knowledge graph", "summary": "A developer built three retrieval pipelines on a shared TigerGraph knowledge graph and Claude Sonnet model to compare plain RAG, GraphRAG, and an agentic GraphRAG approach on 100 Olympic-history questions. The agentic pipeline scored 99% accuracy versus 48% for plain RAG, while GraphRAG reached 69% and was the cheapest per correct answer; the agent used 1.24x RAG's tokens for 2.06x the accuracy and was 40% cheaper per correct answer. The results show lookups need no agent, aggregation breaks plain RAG entirely (0 of 21), and GraphRAG can underperform RAG on multi-hop questions (21% vs 57%).", "body_md": "*Agentic GraphRAG Hackathon by TigerGraph, Round 1. Code: [https://github.com/anirudh12032008/agentic-graphrag-tigergraph](https://github.com/anirudh12032008/agentic-graphrag-tigergraph)*\n\nAsk a RAG system \"How many biathlon events at the 2018 Winter Olympics had more than 73 competitors?\" and it will give you a confident number. It will also be wrong. The answer depends on 8 to 43 documents, and a top-8 vector search can't see all of them.\n\nWe built three pipelines on the same model and the same corpus to find out where graphs help, where agents help, and where neither is needed. The agentic pipeline scored 99% on 100 public questions. Plain RAG scored 48%.\n\nThe corpus is 2,951 English Wikipedia articles about Olympic events from 1987 to 2023, plus distractor pages. The 100 public questions come from five templates:\n\nAll three pipelines use the same LLM (`claude-sonnet-5`), the same embeddings, and the same TigerGraph Savanna instance. An LLM judge grades answers against gold labels, after a deterministic pre-check. Judge tokens are never counted against a pipeline.\n\nEvent pages start with an infobox. We parse it deterministically, with no LLM extraction, so the graph is exact, free to build and reproducible. The TigerGraph schema has `Event`, `Sport`, `Games`, `Venue`, `Athlete`, `Country`, `Document` and `Chunk` vertices. Edges include `IN_SPORT`, `AT_GAMES`, `HELD_AT`, `WON`, `PREV_EDITION` and `HAS_CHUNK`. Chunks carry an embedding, so vector search runs inside the same database as the graph.\n\nThe evaluator rejects an answer if it has no citations, if it cites a document no tool returned, or if it lists unresolved gaps with low confidence. The loop is capped at 8 steps. A circuit breaker stops the run after 2 consecutive tool errors, so a backend outage doesn't burn tokens.\n\n| pipeline | accuracy | avg tokens | avg latency (s) | tokens per correct answer | \n|---|---|---|---|---|\n| RAG | 48% | 6,466 | 7.1 | 13,471 | \n| GraphRAG | 69% | 3,815 | 5.7 | 5,529 | \n| Agentic GraphRAG | 99% | 8,001 | 8.1 | 8,082 | \n\nThe agent uses 1.24x the tokens of RAG for 2.06x the accuracy. Per correct answer it is cheaper than RAG by 40%. GraphRAG is the cheapest per correct answer, but it tops out at 69%.\n\n| question type | RAG | GraphRAG | Agentic | \n|---|---|---|---|\n| lookup (19) | 100% | 100% | 100% | \n| aggregation (21) | 0% | 67% | 100% | \n| superlative (10) | 40% | 80% | 90% | \n| temporal (22) | 36% | 86% | 100% | \n| multi_hop (28) | 57% | 21% | 100% | \n\n**Lookups don't need an agent.** Every pipeline gets them right, and RAG does it in the fewest steps. Pay for the agent only when the question needs it.\n\n**Aggregation is where RAG breaks.** It scored 0 of 21. The evidence is spread over too many documents for k=8 chunks, and no prompt fixes that. The agent runs one GSQL count and gets an exact answer.\n\n**GraphRAG can be worse than RAG.** On multi-hop questions GraphRAG scored 21% against RAG's 57%. Its fixed expansion seeds from the wrong events, and then every fact it adds is confidently about the wrong thing. The agent first links the venue, then filters by date, then reads the winner.\n\n**It flagged real ambiguity.** \"The Olympic Aquatic Centre on August 14, 2004\" matches three finals (Thorpe, Phelps and Klochkova). The venue tool returns every match with an `ambiguous: true` flag, and the agent reports all candidates with citations instead of guessing.\n\n**Its one miss was a tie in the data.** In 2008 fencing, men's épée and women's foil both had 41 competitors. The agent reported the tie, and the gold label picked one. We counted that as a miss.\n\nIt makes 2.6 LLM calls per question on average, and a typical trace has about 5 entries across orchestrator, evaluator and tool steps. On the 50 hidden questions it averaged 6,809 tokens and 7.2 seconds. The cost is real but small compared to the accuracy gain on the question types where it matters.\n\nRAG sometimes couldn't finish at all. On a few aggregation questions it used up a 4,096-token output budget reasoning over 8 chunks and produced no answer. We recorded those as failures and did not hide them.\n\nThe schema already links editions with `PREV_EDITION` and `NEXT_EDITION`. For Round 2 we plan a `ClaimVersion` vertex per sourced fact with a `SUPERSEDES` edge. That would let the agent notice when documents disagree and pick the latest or most authoritative claim, with provenance.\n\nEverything is open source: the graph schema, the GSQL queries, the three pipelines, the benchmark harness and the Streamlit dashboard. Clone the repo and run `make ingest && make bench-public` to reproduce the table above.\n\n*Corpus text is derived from English Wikipedia (CC BY-SA 4.0).*", "url": "https://wpnews.pro/news/when-does-an-agent-beat-rag-we-benchmarked-three-pipelines-on-a-tigergraph-graph", "canonical_source": "https://dev.to/anirudh12032008/when-does-an-agent-beat-rag-we-benchmarked-three-pipelines-on-a-tigergraph-knowledge-graph-2epa", "published_at": "2026-10-08 18:47:16+00:00", "updated_at": "2026-10-08 18:50:12.389875+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-research", "ai-tools"], "entities": ["TigerGraph", "Claude Sonnet", "TigerGraph Savanna", "GitHub", "Wikipedia"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/when-does-an-agent-beat-rag-we-benchmarked-three-pipelines-on-a-tigergraph-graph", "markdown": "https://wpnews.pro/news/when-does-an-agent-beat-rag-we-benchmarked-three-pipelines-on-a-tigergraph-graph.md", "text": "https://wpnews.pro/news/when-does-an-agent-beat-rag-we-benchmarked-three-pipelines-on-a-tigergraph-graph.txt", "jsonld": "https://wpnews.pro/news/when-does-an-agent-beat-rag-we-benchmarked-three-pipelines-on-a-tigergraph-graph.jsonld"}}