Is your AI agent worth its tokens? We measured it with TigerGraph A developer built OCCAM, a routed Agentic GraphRAG pipeline on TigerGraph that answers 100 Olympic-events questions with 100% accuracy at 64 tokens per question, versus 890 tokens for a full agentic GraphRAG pipeline and 3,259 for plain RAG. The team found the agent tier changed answers on only 3 of 100 questions, concluding agents are decisive on roughly 3% of queries and overhead on the rest. Title: Is your AI agent worth its tokens? We measured it with TigerGraph Tags: ai, rag, graph, python Everyone is bolting agents onto retrieval. Almost nobody asks what they cost. For the TigerGraph Agentic GraphRAG Hackathon, the guidebook states the real question: it is not whether agentic produces a better answer, but whether the extra reasoning and retrieval steps are worth the extra token cost. We built OCCAM to answer that with numbers. Named for Occam's razor: entities should not be multiplied beyond necessity, and neither should retrieval steps. The same 100 public questions about Olympic events 2,951 Wikipedia documents, 2,210 event pages plus 740 distractors , through four pipelines, with one model Gemini Flash, temperature 0 : | Pipeline | Accuracy | Tokens / question | |---|---|---| | RAG | 63% | 3,259 | | GraphRAG | 97% | 838 | | Agentic GraphRAG | 100% | 890 | | OCCAM routed | 100% | 64 | All four pipelines execute the same query plan against the same graph. Only the author of the plan differs, so any difference in the numbers comes from the reasoning strategy, not the plumbing. The agent tier changed the answer on 3 of 100 questions and broke none. That is the finding: agents are decisive on about 3% of questions and overhead on the rest. OCCAM sends each question to the cheapest tier that can answer it: A tier only hands over when it can name what it was missing, and that gap is recorded in the trace. The controller never silently skips evidence. 98 of 100 questions were answered at tier 0 with zero tokens. The graph holds Games, Events, Venues, Sports, Athletes and NOCs, with edges like HAS EVENT , AT VENUE , WON GOLD and PREV EDITION . Chunk embeddings live on a Chunk vertex as a TigerVector attribute, so a similarity search can run inside a traversal. Each agent tool is an installed GSQL query on a TigerGraph Savanna workspace 4.2.5 . Aggregation runs entirely in the database: count above walks every event of a sport at one Games and returns the count, so the model never sees the 8 to 43 documents involved. A parity script runs the same questions through the database and our reference implementation: 78 checked, 0 mismatched. We also report evidence recall next to accuracy. RAG scores 63% accuracy against 76% recall, and on superlatives its recall is 0.16: it names a plausible winner without retrieving the documents that settle it. Accuracy alone would have hidden that. artifacts/results hidden.jsonl The takeaway: the value is not in having an agent. It is in knowing which questions need one.