cd /news/ai-agents/when-does-an-ai-agent-actually-earn-… Β· home β€Ί topics β€Ί ai-agents β€Ί article
[ARTICLE Β· art-142482] src=dev.to β†— pub= topic=ai-agents verified=true sentiment=↑ positive

When does an AI agent actually earn its cost? I measured it on 100 questions

A developer built six retrieval pipelines over 2,951 Wikipedia articles in TigerGraph Savanna and benchmarked them on 100 questions using gemini-3.1-flash-lite, finding that an agentic architecture alone added only 3 points of exact match (67% to 70%) while the same agent paired with structured graph tools built from Wikipedia infoboxes reached 99%, a 32-point gain. Aggregation questions went from 0/21 to 21/21 once counting became a graph traversal rather than retrieval, and replacing the generative planner with two typed selection calls held the same 99% while cutting latency from 13 seconds to 1.7 seconds per question with zero generation-model calls.

by read5 min views2 publishedSep 30, 2026

Built for the TigerGraph Agentic GraphRAG Hackathon. Repo and dashboard linked at the end.

RAG retrieves text. GraphRAG adds structure. Agentic GraphRAG lets a model plan

its own investigation. Everyone assumes the third one is better β€” but by how

much, and on which questions? That's the question this hackathon asked, so I

built six pipelines over one corpus and measured them against the same 100

questions, with the same generation model throughout.

The short version: exact match went from 67% to 99%, a 32-point jump. But

the interesting part isn't that number. It's that an agent by itself only got

me 3 of those 32 points.

2,951 Wikipedia articles in TigerGraph Savanna. 100 evaluation questions across

five types: lookup, temporal, multi-hop, superlative and aggregation. Every

pipeline uses gemini-3.1-flash-lite β€” a cheap model, deliberately β€” with local

BGE embeddings and TigerGraph's native vector index. No pipeline gets a better

model than any other.

Scoring is exact match against the gold answer, computed with no model in the

loop. I also ran an LLM judge, and I'll come back to why I stopped trusting it.

Plain RAG scored 67%. GraphRAG β€” entity linking plus one-hop traversal β€” also

scored 67%. Identical. That was my first surprise.

Breaking it down by question type showed why. On aggregation questions β€”

"how many cycling events had more than 30 competitors?" β€” RAG got 1 out of 21. GraphRAG got

The reason is structural, and no amount of prompt engineering touches it.

Answering that question requires every matching document; for one question,

43 of them. Retrieval fetches the top five and counts those. The model then

confidently reports a number that is simply the size of what it was shown.

My LLM-extracted entity graph didn't help either. It had 12,000 entities with

free-text relationship labels β€” but "competitors: 43" was never a property you

could filter or count on. I had modelled the prose and not the facts.

Every one of those articles opens with an infobox: event, games, venue, date,

competitors, nations, gold medallist. So I built a second layer in the graph

directly from those boxes β€” OlympicEvent linked to its Games, Sport and

Venue, plus a PREV_GAMES edge so "the Olympics before 2016" is one hop.

That ingestion makes zero LLM calls. It's a parser. And it turns counting

from a retrieval problem into a graph traversal: aggregation went from 0/21 to 21/21.

My agent plans each step and chooses between five tools: a structured graph

query, time-scoped fact lookup, entity linking, one-hop traversal, and vector

search over chunks. It states what the evidence does and doesn't establish in the same structured call as the routing decision, so self-evaluation costs no

Here's the comparison that I think is the real result:

Pipeline Exact match vs RAG
RAG 67% β€”
GraphRAG 67% +0
Agent over text/entity tools only 70% +3
Same agent + structured graph tools 99% +32

The agentic architecture alone bought 3 points. The agent with the right retrieval surface bought 32. If I had only built the agent, I'd have concluded

If the agent picks the structured query on 100 out of 100 questions, does it need to reason at all? I tested it: everything the planner decides already

exists as a row in the graph β€” 47 sports, 20 Games, 316 venues, 475 event names.

So planning isn't generation, it's selection.

I replaced the generative planner with two typed selection calls (TypeSafe's

System One), kept the same GSQL and the same templated answer, and got:

Same 99%. 1.7 seconds per question instead of 13. Zero generation-model calls.

That last number matters more than the speed. My free tier is 500 calls a day β€”

one benchmark run β€” and every measurement I took that week died on that wall,

three times in one day. A path with no generation calls can't hit a rate limit,

which means a live demo can't fail halfway through.

So the honest answer to the hackathon's question: agents earn their cost on open-ended questions, and they're overkill where a typed query settles it. The

A submission that only lists wins isn't worth much, so:

My LLM judge was unreliable. It gave 4 or 5 out of 5 to 14 answers that are wrong β€” mostly fluent refusals like "the corpus does not contain this."

A "fix" of mine cost 17 points. I changed which answer field wins, and exact

match dropped from 99 to 82. The agent couldn't see what its tools had computed,

so it recounted a truncated evidence list and answered 12 β€” the size of the

cap β€” to seven counting questions.

One fix made things worse before better. My first repair made a question

fail three times out of three instead of one. A regex only matched year and

season when adjacent, so "the Summer Olympics held immediately before 2016"

produced no season at all.

A stale artifact nearly shipped. My hidden-set answers were generated two

minutes before a bug fix. Five counts were wrong. I only caught it by answering

the same questions with a second pipeline and scoring both against a

deterministic oracle.

Every one of those is now pinned by a regression test, because each was

invisible in normal use β€” the system kept answering, it just answered wrong.

RAG GraphRAG Agentic (ablation) Agentic (full) Selection planner
Exact match 67% 67% 70% 99% 99%
Aggregation 1/21 0/21 3/21 21/21 21/21
Tokens/question 3,586 3,952 6,065 3,412 2,267
Seconds 6.8 10.2 20.2 13.2 1.7

The agent reaches 99% on fewer tokens than plain RAG. A cheap model with the

right architecture beats an expensive one with a lazy pipeline.

**Repo:** [https://github.com/antcybersec/graph_rag](https://github.com/antcybersec/graph_rag)

**Live dashboard:** [https://graphrag-c.streamlit.app/](https://graphrag-c.streamlit.app/)

Built on TigerGraph Savanna with GSQL and its native vector index.

── more in #ai-agents 4 stories Β· sorted by recency
── more on @tigergraph 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/when-does-an-ai-agen…] indexed:0 read:5min 2026-09-30 Β· β€”