{"slug": "when-does-an-ai-agent-actually-earn-its-cost-i-measured-it-on-100-questions", "title": "When does an AI agent actually earn its cost? I measured it on 100 questions", "summary": "A developer built six retrieval pipelines over 2,951 Wikipedia articles in TigerGraph Savanna and benchmarked them on 100 questions using gemini-3.1-flash-lite, finding that an agentic architecture alone added only 3 points of exact match (67% to 70%) while the same agent paired with structured graph tools built from Wikipedia infoboxes reached 99%, a 32-point gain. Aggregation questions went from 0/21 to 21/21 once counting became a graph traversal rather than retrieval, and replacing the generative planner with two typed selection calls held the same 99% while cutting latency from 13 seconds to 1.7 seconds per question with zero generation-model calls.", "body_md": "*Built for the TigerGraph Agentic GraphRAG Hackathon. Repo and dashboard linked at the end.*\n\nRAG retrieves text. GraphRAG adds structure. Agentic GraphRAG lets a model plan\n\nits own investigation. Everyone assumes the third one is better — but by how\n\nmuch, and on which questions? That's the question this hackathon asked, so I\n\nbuilt six pipelines over one corpus and measured them against the same 100\n\nquestions, with the same generation model throughout.\n\nThe short version: **exact match went from 67% to 99%, a 32-point jump.** But\n\nthe interesting part isn't that number. It's that an agent *by itself* only got\n\nme 3 of those 32 points.\n\n2,951 Wikipedia articles in TigerGraph Savanna. 100 evaluation questions across\n\nfive types: lookup, temporal, multi-hop, superlative and aggregation. Every\n\npipeline uses `gemini-3.1-flash-lite` — a cheap model, deliberately — with local\n\nBGE embeddings and TigerGraph's native vector index. No pipeline gets a better\n\nmodel than any other.\n\nScoring is exact match against the gold answer, computed with no model in the\n\nloop. I also ran an LLM judge, and I'll come back to why I stopped trusting it.\n\nPlain RAG scored 67%. GraphRAG — entity linking plus one-hop traversal — also\n\nscored 67%. Identical. That was my first surprise.\n\nBreaking it down by question type showed why. On aggregation questions —\n\n*\"how many cycling events had more than 30 competitors?\"* — RAG got **1 out of 21**. GraphRAG got \n\nThe reason is structural, and no amount of prompt engineering touches it.\n\nAnswering that question requires *every* matching document; for one question,\n\n43 of them. Retrieval fetches the top five and counts those. The model then\n\nconfidently reports a number that is simply the size of what it was shown.\n\nMy LLM-extracted entity graph didn't help either. It had 12,000 entities with\n\nfree-text relationship labels — but \"competitors: 43\" was never a property you\n\ncould filter or count on. I had modelled the prose and not the facts.\n\nEvery one of those articles opens with an infobox: event, games, venue, date,\n\ncompetitors, nations, gold medallist. So I built a second layer in the graph\n\ndirectly from those boxes — `OlympicEvent` linked to its `Games`, `Sport` and\n\n`Venue`, plus a `PREV_GAMES` edge so \"the Olympics before 2016\" is one hop.\n\nThat ingestion makes **zero LLM calls**. It's a parser. And it turns counting\n\nfrom a retrieval problem into a graph traversal: aggregation went from 0/21 to\n\n**21/21**.\n\nMy agent plans each step and chooses between five tools: a structured graph\n\nquery, time-scoped fact lookup, entity linking, one-hop traversal, and vector\n\nsearch over chunks. It states what the evidence does and doesn't establish *in the same structured call* as the routing decision, so self-evaluation costs no\n\nHere's the comparison that I think is the real result:\n\n| Pipeline | Exact match | vs RAG | \n|---|---|---|\n| RAG | 67% | — | \n| GraphRAG | 67% | +0 | \n| Agent over text/entity tools only | 70% | **+3** | \n| Same agent + structured graph tools | 99% | **+32** | \n\nThe agentic architecture alone bought 3 points. The agent *with the right retrieval surface* bought 32. If I had only built the agent, I'd have concluded\n\nIf the agent picks the structured query on 100 out of 100 questions, does it\n\nneed to *reason* at all? I tested it: everything the planner decides already\n\nexists as a row in the graph — 47 sports, 20 Games, 316 venues, 475 event names.\n\nSo planning isn't generation, it's **selection**.\n\nI replaced the generative planner with two typed selection calls (TypeSafe's\n\nSystem One), kept the same GSQL and the same templated answer, and got:\n\n**Same 99%. 1.7 seconds per question instead of 13. Zero generation-model calls.**\n\nThat last number matters more than the speed. My free tier is 500 calls a day —\n\none benchmark run — and every measurement I took that week died on that wall,\n\nthree times in one day. A path with no generation calls can't hit a rate limit,\n\nwhich means a live demo can't fail halfway through.\n\nSo the honest answer to the hackathon's question: **agents earn their cost on open-ended questions, and they're overkill where a typed query settles it.** The\n\nA submission that only lists wins isn't worth much, so:\n\n**My LLM judge was unreliable.** It gave 4 or 5 out of 5 to **14 answers that are wrong** — mostly fluent refusals like \"the corpus does not contain this.\"\n\n**A \"fix\" of mine cost 17 points.** I changed which answer field wins, and exact\n\nmatch dropped from 99 to 82. The agent couldn't see what its tools had computed,\n\nso it recounted a truncated evidence list and answered `12` — the size of the\n\ncap — to seven counting questions.\n\n**One fix made things worse before better.** My first repair made a question\n\nfail three times out of three instead of one. A regex only matched year and\n\nseason when adjacent, so \"the Summer Olympics held immediately before 2016\"\n\nproduced no season at all.\n\n**A stale artifact nearly shipped.** My hidden-set answers were generated two\n\nminutes before a bug fix. Five counts were wrong. I only caught it by answering\n\nthe same questions with a second pipeline and scoring both against a\n\ndeterministic oracle.\n\nEvery one of those is now pinned by a regression test, because each was\n\ninvisible in normal use — the system kept answering, it just answered wrong.\n\n|  | RAG | GraphRAG | Agentic (ablation) | Agentic (full) | Selection planner | \n|---|---|---|---|---|---|\n| Exact match | 67% | 67% | 70% | **99%** | 99% | \n| Aggregation | 1/21 | 0/21 | 3/21 | **21/21** | 21/21 | \n| Tokens/question | 3,586 | 3,952 | 6,065 | 3,412 | 2,267 | \n| Seconds | 6.8 | 10.2 | 20.2 | 13.2 | **1.7** | \n\nThe agent reaches 99% on **fewer tokens than plain RAG**. A cheap model with the\n\nright architecture beats an expensive one with a lazy pipeline.\n\n**Repo:** [https://github.com/antcybersec/graph_rag](https://github.com/antcybersec/graph_rag)\n\n**Live dashboard:** [https://graphrag-c.streamlit.app/](https://graphrag-c.streamlit.app/)\n\nBuilt on TigerGraph Savanna with GSQL and its native vector index.", "url": "https://wpnews.pro/news/when-does-an-ai-agent-actually-earn-its-cost-i-measured-it-on-100-questions", "canonical_source": "https://dev.to/anant_kumar_bf65d0d3994d3/-when-does-an-ai-agent-actually-earn-its-cost-i-measured-it-on-100-questions-4j5l", "published_at": "2026-09-30 12:07:58+00:00", "updated_at": "2026-09-30 12:18:11.199665+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-research", "mlops", "ai-tools"], "entities": ["TigerGraph", "TigerGraph Savanna", "Wikipedia", "gemini-3.1-flash-lite", "BGE", "TypeSafe"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/when-does-an-ai-agent-actually-earn-its-cost-i-measured-it-on-100-questions", "markdown": "https://wpnews.pro/news/when-does-an-ai-agent-actually-earn-its-cost-i-measured-it-on-100-questions.md", "text": "https://wpnews.pro/news/when-does-an-ai-agent-actually-earn-its-cost-i-measured-it-on-100-questions.txt", "jsonld": "https://wpnews.pro/news/when-does-an-ai-agent-actually-earn-its-cost-i-measured-it-on-100-questions.jsonld"}}