{"slug": "is-your-ai-agent-worth-its-tokens-we-measured-it-with-tigergraph", "title": "Is your AI agent worth its tokens? We measured it with TigerGraph", "summary": "A developer built OCCAM, a routed Agentic GraphRAG pipeline on TigerGraph that answers 100 Olympic-events questions with 100% accuracy at 64 tokens per question, versus 890 tokens for a full agentic GraphRAG pipeline and 3,259 for plain RAG. The team found the agent tier changed answers on only 3 of 100 questions, concluding agents are decisive on roughly 3% of queries and overhead on the rest.", "body_md": "Title: Is your AI agent worth its tokens? We measured it with TigerGraph\n\nTags: ai, rag, graph, python\n\nEveryone is bolting agents onto retrieval. Almost nobody asks what they cost.\n\nFor the TigerGraph Agentic GraphRAG Hackathon, the guidebook states the real question: it is not whether agentic produces a better answer, but whether the extra reasoning and retrieval steps are worth the extra token cost. We built **OCCAM** to answer that with numbers.\n\nNamed for Occam's razor: *entities should not be multiplied beyond necessity, and neither should retrieval steps.*\n\nThe same 100 public questions about Olympic events (2,951 Wikipedia documents, 2,210 event pages plus 740 distractors), through four pipelines, with one model (Gemini Flash, temperature 0):\n\n| Pipeline | Accuracy | Tokens / question | \n|---|---|---|\n| RAG | 63% | 3,259 | \n| GraphRAG | 97% | 838 | \n| Agentic GraphRAG | 100% | 890 | \n| **OCCAM (routed)** | **100%** | **64** | \n\nAll four pipelines execute the same query plan against the same graph. Only the author of the plan differs, so any difference in the numbers comes from the reasoning strategy, not the plumbing.\n\nThe agent tier changed the answer on 3 of 100 questions and broke none. That is the finding: agents are decisive on about 3% of questions and overhead on the rest.\n\nOCCAM sends each question to the cheapest tier that can answer it:\n\nA tier only hands over when it can name what it was missing, and that gap is recorded in the trace. The controller never silently skips evidence. 98 of 100 questions were answered at tier 0 with zero tokens.\n\nThe graph holds Games, Events, Venues, Sports, Athletes and NOCs, with edges like `HAS_EVENT`, `AT_VENUE`, `WON_GOLD` and `PREV_EDITION`. Chunk embeddings live on a `Chunk` vertex as a TigerVector attribute, so a similarity search can run inside a traversal.\n\nEach agent tool is an installed GSQL query on a TigerGraph Savanna workspace (4.2.5). Aggregation runs entirely in the database: `count_above` walks every event of a sport at one Games and returns the count, so the model never sees the 8 to 43 documents involved. A parity script runs the same questions through the database and our reference implementation: 78 checked, 0 mismatched.\n\nWe also report **evidence recall** next to accuracy. RAG scores 63% accuracy against 76% recall, and on superlatives its recall is 0.16: it names a plausible winner without retrieving the documents that settle it. Accuracy alone would have hidden that.\n\n`artifacts/results_hidden.jsonl`\nThe takeaway: the value is not in having an agent. It is in knowing which questions need one.", "url": "https://wpnews.pro/news/is-your-ai-agent-worth-its-tokens-we-measured-it-with-tigergraph", "canonical_source": "https://dev.to/techtush/is-your-ai-agent-worth-its-tokens-we-measured-it-with-tigergraph-54jd", "published_at": "2026-10-03 10:36:21+00:00", "updated_at": "2026-10-03 10:37:50.389338+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-research", "mlops", "ai-tools"], "entities": ["TigerGraph", "OCCAM", "Gemini Flash", "TigerGraph Savanna", "GSQL", "TigerVector", "Wikipedia"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/is-your-ai-agent-worth-its-tokens-we-measured-it-with-tigergraph", "markdown": "https://wpnews.pro/news/is-your-ai-agent-worth-its-tokens-we-measured-it-with-tigergraph.md", "text": "https://wpnews.pro/news/is-your-ai-agent-worth-its-tokens-we-measured-it-with-tigergraph.txt", "jsonld": "https://wpnews.pro/news/is-your-ai-agent-worth-its-tokens-we-measured-it-with-tigergraph.jsonld"}}