# The agentic RAG pipeline that was faster and cheaper — and no more accurate than no agent at all

> Source: <https://dev.to/kultzuki/the-agentic-rag-pipeline-that-was-faster-and-cheaper-and-no-more-accurate-than-no-agent-at-all-1gbg>
> Published: 2026-10-04 10:35:45+00:00

I spent a few weeks building four retrieval architectures over the same graph database, pointing them at the same 150 questions, and measuring what each one cost.

The headline numbers:

|  | accuracy | LLM tokens/q | mean latency | tool calls/q | 
|---|---|---|---|---|
| P1 · RAG | 30% | 1,284 | 27.0 s | 1.00 | 
| P2 · GraphRAG | 89% | 1,694 | 32.2 s | 2.72 | 
| **P3 · Agentic** | **100%** | **552** | **10.2 s** | **1.95** | 
| P4 · Deterministic control | 100% | **0** | **0.1 s** | 1.57 | 

One model — **`Qwen3.8-27B`**, with its reasoning trace enabled — served every pipeline, including the agentic orchestrator's planner and verifier.

Only LLM tokens were counted; GSQL and Python were treated as zero-cost.

I built P4 as a sanity check, not a competitor: strip out the model entirely, parse the question with regexes, and run one deterministic query. I expected it to score maybe 70% and lose badly.

**It scored 100%.**

That is the most useful result in the whole benchmark, and it changes how you should read every other row.

The agentic pipeline is not better than the deterministic control. It is *more expensive* — at 552 tokens per question — for identical accuracy.

So why keep it?

Because the control is a boundary marker, not a competitor. It tells you exactly how many questions never needed a model.

And because the benchmark questions turned out to be templated, that number was **all of them** — which is precisely the caveat I'll come back to at the end.

The 150 questions come from five templates. Four of them are structured database queries wearing question marks:

| template | P1 · RAG | P2 · GraphRAG | P3 · Agentic | 
|---|---|---|---|
| lookup | 63% | 100% | 100% | 
| temporal | 27% | 55% | 100% | 
| multi_hop | 32% | 96% | 100% | 
| aggregation | **5%** | 100% | 100% | 
| superlative | **20%** | 100% | 100% | 

Plain RAG scores 63% when the answer is in one document and **5%** when the question asks how many events meet a condition.

That is not a retrieval-quality problem.

A vector retriever cannot count. It returns chunks, the model reads them, and the model is bad at counting over a set — so no amount of chunk quality fixes it.

GraphRAG fixed it by pushing aggregation into the database.

The GSQL does **`count`** with a predicate and returns a number. The model's job shrinks from "compute this" to "read this", which is the job models are actually good at.

This is the real dividing line between RAG and GraphRAG, and it has nothing to do with graph databases being fashionable.

**It's arithmetic.**

The agentic pipeline costs 67% fewer tokens than GraphRAG while being 11 points more accurate.

Agents are normally assumed to be expensive — more calls, more context, more chances to ramble.

Mine did less work.

Three design decisions did that.

The orchestrator emits **`STEP: count_above`** and nothing else.

**`sport`**, **` year`**, **` season`**, and **` threshold`** come from a regex parser.

An earlier version let the planner's parameters override the parsed ones. Because questions say "biathlon" while the graph says "Biathlon", every aggregation silently answered zero.

Slot filling is mechanical.

The *plan* is the part that varies with the question, and that's the part worth spending tokens on.

When a tool returns a definite value and the deterministic verifier agrees, the loop accepts it and stops rather than paying for another round trip.

**72 of 100 public questions ended that way.**

A planner re-issuing the same query is looping, not reasoning.

Detecting the repeat and telling it to move on is what keeps the average at ~2 tool calls instead of burning the step budget on one empty result.

The reason P3 is also the *fastest* pipeline is the same mechanism.

GraphRAG's 1,694 tokens are spent almost entirely on the model doing sums that the database could have done, and every token is also latency.

If you build on TigerGraph, some of this will save you days.

`vectorSearch()` does not work in interpreted queries
It fails with:

`GSQL-2500 Unsupported Statement | TOPK_VEC_SEARCH_FUNC`

It has to live in an *installed* query, which means one **`CREATE QUERY`** + **` INSTALL QUERY`**.

The **`USE GRAPH`** prefix also has to be repeated in every call because it does not persist between them.

**`INTERPRET QUERY (sp=STRING)`** is rejected whether you pass a dict or a query string.

So values get inlined as escaped literals.

If you build this, make sure every inlined value originates from your own parser and not from user input.

It is stricter than you'd expect:

`SELECT e FROM Event:e`` SELECT` needs a vertex-set variable.`Games:g <-IN_GAMES- Event:e`` FOREACH``MaxAccum`` coalesce``IFF`` SCHEMA_CHANGE JOB`
A bare **`ALTER ... ADD VECTOR ATTRIBUTE`** is not valid GSQL.

And if any type in your schema is declared *global*, the job must be global too.

The full list is **31 footguns** in the repo, each one of which produced a silently wrong answer before it was found.

This is the part I'd most want other people to take away.

My results summary reported **`100/100`** and **` errors: 0`** — both true — while **26 of 150 agentic answers were a whole sentence where a bare number was due**.

For example:

`"'Men's foil' has nations=29"`

instead of:

`"29"`

The cause: my answer extractor's markers were all *spaced* — **`" is "`**, **`" = "`**, **`": "`**.

The tool emitted an *unspaced* **`key=value`** tail.

All three markers missed, the function fell through to **`return line.strip()`**, and the clause became the answer.

I found it by grepping the artefacts for the *shape* of an answer, not by reading the summary.

The summary was accurate about what it measured and completely silent about what it didn't.

My scorer returned **`True`** for:

`is_correct("c", ["Chen Ding"])`

because containment had no length floor, so any single character landing inside the gold scored as correct.

A needle now has to be **4+ characters or a standalone token**, which keeps **`"29"`** matching inside **`"29 nations"`** while rejecting **`"4"`** inside **`"24"`**.

**Both defects were scored `correct` the whole time.**

Containment hid the first one, and the second had nothing to fire on because the pipelines emit corpus-exact strings.

I re-scored the entire benchmark with the hardened scorer and *nothing moved*.

That is the point.

They were latent, not active.

`"Aug 7 (prelim), Aug 10 (final)"`
The general lesson:

**Aggregate metrics are blind to shape defects.**

A summary can be entirely correct about accuracy and tell you nothing about whether your answers are in the right form to be graded.

I ended up with **74 tests** using Python's standard-library **`unittest`**.

They need neither the graph nor the model — which matters more than it sounds, because the model endpoint was down a lot during this build.

Two things, stated plainly because the results invite the wrong conclusion.

**`changed_strategy`** is **0/100**.

Every one of **211 tool calls** returned data — no question ever made the first move fail, so the re-planning machinery never fired.

The agent demonstrably runs multi-step plans and gets them right, but I never observed it recover from a bad step.

An adversarial probe set — misspelled sport, invented venue, impossible Games edition — would exercise that, and I didn't build one.

Worth noting that such a probe set probably wouldn't show the agent winning anyway.

Both the agentic and deterministic pipelines share the same regex parser and the same entity linker, so corrupting an entity breaks *both* equally.

My mental model of "the router fails, the agent recovers" was wrong about this architecture.

All three pipelines call a 27B reasoning model over a tunnel, so those **27–32 second** figures are dominated by network round-trips to a model runtime, not by retrieval.

The graph itself answers in well under a second.

The zero-token control finishes a full 150-question pass in **0.1 s average**.

If you see **32 s for a graph database**, ask what fraction of it was the model.

If I were picking this up again, the interesting work isn't squeezing accuracy — that's saturated, and the deterministic control already proves it.

It's building the adversarial probe set, because the only untested claim in the whole system is whether the agent can recover when its first move fails.

Everything else is measured.

Full source, traces, and the hidden-50 outputs:

[https://github.com/Kultzuki/TigerGraph](https://github.com/Kultzuki/TigerGraph)

Live results page (self-contained, no dependencies):

[https://kultzuki.github.io/TigerGraph/olympic-graphrag/results/report.html](https://kultzuki.github.io/TigerGraph/olympic-graphrag/results/report.html)

Setup is one dependency:

```
bash
pip install -r requirements.txt
```


