cd /news/ai-agents/is-your-ai-agent-worth-its-tokens-we… · home › topics › ai-agents › article
[ARTICLE · art-144399] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Is your AI agent worth its tokens? We measured it with TigerGraph

A developer built OCCAM, a routed Agentic GraphRAG pipeline on TigerGraph that answers 100 Olympic-events questions with 100% accuracy at 64 tokens per question, versus 890 tokens for a full agentic GraphRAG pipeline and 3,259 for plain RAG. The team found the agent tier changed answers on only 3 of 100 questions, concluding agents are decisive on roughly 3% of queries and overhead on the rest.

by read2 min views1 publishedOct 3, 2026

Title: Is your AI agent worth its tokens? We measured it with TigerGraph

Tags: ai, rag, graph, python

Everyone is bolting agents onto retrieval. Almost nobody asks what they cost.

For the TigerGraph Agentic GraphRAG Hackathon, the guidebook states the real question: it is not whether agentic produces a better answer, but whether the extra reasoning and retrieval steps are worth the extra token cost. We built OCCAM to answer that with numbers. Named for Occam's razor: entities should not be multiplied beyond necessity, and neither should retrieval steps.

The same 100 public questions about Olympic events (2,951 Wikipedia documents, 2,210 event pages plus 740 distractors), through four pipelines, with one model (Gemini Flash, temperature 0):

Pipeline Accuracy Tokens / question
RAG 63% 3,259
GraphRAG 97% 838
Agentic GraphRAG 100% 890
OCCAM (routed) 100% 64

All four pipelines execute the same query plan against the same graph. Only the author of the plan differs, so any difference in the numbers comes from the reasoning strategy, not the plumbing.

The agent tier changed the answer on 3 of 100 questions and broke none. That is the finding: agents are decisive on about 3% of questions and overhead on the rest.

OCCAM sends each question to the cheapest tier that can answer it:

A tier only hands over when it can name what it was missing, and that gap is recorded in the trace. The controller never silently skips evidence. 98 of 100 questions were answered at tier 0 with zero tokens.

The graph holds Games, Events, Venues, Sports, Athletes and NOCs, with edges like HAS_EVENT, AT_VENUE, WON_GOLD and PREV_EDITION. Chunk embeddings live on a Chunk vertex as a TigerVector attribute, so a similarity search can run inside a traversal.

Each agent tool is an installed GSQL query on a TigerGraph Savanna workspace (4.2.5). Aggregation runs entirely in the database: count_above walks every event of a sport at one Games and returns the count, so the model never sees the 8 to 43 documents involved. A parity script runs the same questions through the database and our reference implementation: 78 checked, 0 mismatched.

We also report evidence recall next to accuracy. RAG scores 63% accuracy against 76% recall, and on superlatives its recall is 0.16: it names a plausible winner without retrieving the documents that settle it. Accuracy alone would have hidden that.

artifacts/results_hidden.jsonl The takeaway: the value is not in having an agent. It is in knowing which questions need one.

── more in #ai-agents 4 stories · sorted by recency
── more on @tigergraph 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/is-your-ai-agent-wor…] indexed:0 read:2min 2026-10-03 · —