{"slug": "the-agentic-rag-pipeline-that-was-faster-and-cheaper-and-no-more-accurate-than", "title": "The agentic RAG pipeline that was faster and cheaper — and no more accurate than no agent at all", "summary": "A developer benchmarked four retrieval architectures over the same graph database against 150 questions, finding that an agentic RAG pipeline hit 100% accuracy at 552 LLM tokens per question and 10.2 seconds mean latency, while a deterministic regex-and-query control also scored 100% at zero tokens and 0.1 seconds. Plain RAG managed only 30% accuracy overall and 5% on aggregation questions, which the developer attributes to vector retrievers being unable to count, whereas GraphRAG pushed counting into the database via GSQL. The agentic pipeline used 67% fewer tokens than GraphRAG by restricting the orchestrator to emitting a step name while a regex parser handled slot filling.", "body_md": "I spent a few weeks building four retrieval architectures over the same graph database, pointing them at the same 150 questions, and measuring what each one cost.\n\nThe headline numbers:\n\n|  | accuracy | LLM tokens/q | mean latency | tool calls/q | \n|---|---|---|---|---|\n| P1 · RAG | 30% | 1,284 | 27.0 s | 1.00 | \n| P2 · GraphRAG | 89% | 1,694 | 32.2 s | 2.72 | \n| **P3 · Agentic** | **100%** | **552** | **10.2 s** | **1.95** | \n| P4 · Deterministic control | 100% | **0** | **0.1 s** | 1.57 | \n\nOne model — **`Qwen3.8-27B`**, with its reasoning trace enabled — served every pipeline, including the agentic orchestrator's planner and verifier.\n\nOnly LLM tokens were counted; GSQL and Python were treated as zero-cost.\n\nI built P4 as a sanity check, not a competitor: strip out the model entirely, parse the question with regexes, and run one deterministic query. I expected it to score maybe 70% and lose badly.\n\n**It scored 100%.**\n\nThat is the most useful result in the whole benchmark, and it changes how you should read every other row.\n\nThe agentic pipeline is not better than the deterministic control. It is *more expensive* — at 552 tokens per question — for identical accuracy.\n\nSo why keep it?\n\nBecause the control is a boundary marker, not a competitor. It tells you exactly how many questions never needed a model.\n\nAnd because the benchmark questions turned out to be templated, that number was **all of them** — which is precisely the caveat I'll come back to at the end.\n\nThe 150 questions come from five templates. Four of them are structured database queries wearing question marks:\n\n| template | P1 · RAG | P2 · GraphRAG | P3 · Agentic | \n|---|---|---|---|\n| lookup | 63% | 100% | 100% | \n| temporal | 27% | 55% | 100% | \n| multi_hop | 32% | 96% | 100% | \n| aggregation | **5%** | 100% | 100% | \n| superlative | **20%** | 100% | 100% | \n\nPlain RAG scores 63% when the answer is in one document and **5%** when the question asks how many events meet a condition.\n\nThat is not a retrieval-quality problem.\n\nA vector retriever cannot count. It returns chunks, the model reads them, and the model is bad at counting over a set — so no amount of chunk quality fixes it.\n\nGraphRAG fixed it by pushing aggregation into the database.\n\nThe GSQL does **`count`** with a predicate and returns a number. The model's job shrinks from \"compute this\" to \"read this\", which is the job models are actually good at.\n\nThis is the real dividing line between RAG and GraphRAG, and it has nothing to do with graph databases being fashionable.\n\n**It's arithmetic.**\n\nThe agentic pipeline costs 67% fewer tokens than GraphRAG while being 11 points more accurate.\n\nAgents are normally assumed to be expensive — more calls, more context, more chances to ramble.\n\nMine did less work.\n\nThree design decisions did that.\n\nThe orchestrator emits **`STEP: count_above`** and nothing else.\n\n**`sport`**, **` year`**, **` season`**, and **` threshold`** come from a regex parser.\n\nAn earlier version let the planner's parameters override the parsed ones. Because questions say \"biathlon\" while the graph says \"Biathlon\", every aggregation silently answered zero.\n\nSlot filling is mechanical.\n\nThe *plan* is the part that varies with the question, and that's the part worth spending tokens on.\n\nWhen a tool returns a definite value and the deterministic verifier agrees, the loop accepts it and stops rather than paying for another round trip.\n\n**72 of 100 public questions ended that way.**\n\nA planner re-issuing the same query is looping, not reasoning.\n\nDetecting the repeat and telling it to move on is what keeps the average at ~2 tool calls instead of burning the step budget on one empty result.\n\nThe reason P3 is also the *fastest* pipeline is the same mechanism.\n\nGraphRAG's 1,694 tokens are spent almost entirely on the model doing sums that the database could have done, and every token is also latency.\n\nIf you build on TigerGraph, some of this will save you days.\n\n`vectorSearch()` does not work in interpreted queries\nIt fails with:\n\n`GSQL-2500 Unsupported Statement | TOPK_VEC_SEARCH_FUNC`\n\nIt has to live in an *installed* query, which means one **`CREATE QUERY`** + **` INSTALL QUERY`**.\n\nThe **`USE GRAPH`** prefix also has to be repeated in every call because it does not persist between them.\n\n**`INTERPRET QUERY (sp=STRING)`** is rejected whether you pass a dict or a query string.\n\nSo values get inlined as escaped literals.\n\nIf you build this, make sure every inlined value originates from your own parser and not from user input.\n\nIt is stricter than you'd expect:\n\n`SELECT e FROM Event:e`` SELECT` needs a vertex-set variable.`Games:g <-IN_GAMES- Event:e`` FOREACH``MaxAccum`` coalesce``IFF`` SCHEMA_CHANGE JOB`\nA bare **`ALTER ... ADD VECTOR ATTRIBUTE`** is not valid GSQL.\n\nAnd if any type in your schema is declared *global*, the job must be global too.\n\nThe full list is **31 footguns** in the repo, each one of which produced a silently wrong answer before it was found.\n\nThis is the part I'd most want other people to take away.\n\nMy results summary reported **`100/100`** and **` errors: 0`** — both true — while **26 of 150 agentic answers were a whole sentence where a bare number was due**.\n\nFor example:\n\n`\"'Men's foil' has nations=29\"`\n\ninstead of:\n\n`\"29\"`\n\nThe cause: my answer extractor's markers were all *spaced* — **`\" is \"`**, **`\" = \"`**, **`\": \"`**.\n\nThe tool emitted an *unspaced* **`key=value`** tail.\n\nAll three markers missed, the function fell through to **`return line.strip()`**, and the clause became the answer.\n\nI found it by grepping the artefacts for the *shape* of an answer, not by reading the summary.\n\nThe summary was accurate about what it measured and completely silent about what it didn't.\n\nMy scorer returned **`True`** for:\n\n`is_correct(\"c\", [\"Chen Ding\"])`\n\nbecause containment had no length floor, so any single character landing inside the gold scored as correct.\n\nA needle now has to be **4+ characters or a standalone token**, which keeps **`\"29\"`** matching inside **`\"29 nations\"`** while rejecting **`\"4\"`** inside **`\"24\"`**.\n\n**Both defects were scored `correct` the whole time.**\n\nContainment hid the first one, and the second had nothing to fire on because the pipelines emit corpus-exact strings.\n\nI re-scored the entire benchmark with the hardened scorer and *nothing moved*.\n\nThat is the point.\n\nThey were latent, not active.\n\n`\"Aug 7 (prelim), Aug 10 (final)\"`\nThe general lesson:\n\n**Aggregate metrics are blind to shape defects.**\n\nA summary can be entirely correct about accuracy and tell you nothing about whether your answers are in the right form to be graded.\n\nI ended up with **74 tests** using Python's standard-library **`unittest`**.\n\nThey need neither the graph nor the model — which matters more than it sounds, because the model endpoint was down a lot during this build.\n\nTwo things, stated plainly because the results invite the wrong conclusion.\n\n**`changed_strategy`** is **0/100**.\n\nEvery one of **211 tool calls** returned data — no question ever made the first move fail, so the re-planning machinery never fired.\n\nThe agent demonstrably runs multi-step plans and gets them right, but I never observed it recover from a bad step.\n\nAn adversarial probe set — misspelled sport, invented venue, impossible Games edition — would exercise that, and I didn't build one.\n\nWorth noting that such a probe set probably wouldn't show the agent winning anyway.\n\nBoth the agentic and deterministic pipelines share the same regex parser and the same entity linker, so corrupting an entity breaks *both* equally.\n\nMy mental model of \"the router fails, the agent recovers\" was wrong about this architecture.\n\nAll three pipelines call a 27B reasoning model over a tunnel, so those **27–32 second** figures are dominated by network round-trips to a model runtime, not by retrieval.\n\nThe graph itself answers in well under a second.\n\nThe zero-token control finishes a full 150-question pass in **0.1 s average**.\n\nIf you see **32 s for a graph database**, ask what fraction of it was the model.\n\nIf I were picking this up again, the interesting work isn't squeezing accuracy — that's saturated, and the deterministic control already proves it.\n\nIt's building the adversarial probe set, because the only untested claim in the whole system is whether the agent can recover when its first move fails.\n\nEverything else is measured.\n\nFull source, traces, and the hidden-50 outputs:\n\n[https://github.com/Kultzuki/TigerGraph](https://github.com/Kultzuki/TigerGraph)\n\nLive results page (self-contained, no dependencies):\n\n[https://kultzuki.github.io/TigerGraph/olympic-graphrag/results/report.html](https://kultzuki.github.io/TigerGraph/olympic-graphrag/results/report.html)\n\nSetup is one dependency:\n\n```\nbash\npip install -r requirements.txt\n```\n\n", "url": "https://wpnews.pro/news/the-agentic-rag-pipeline-that-was-faster-and-cheaper-and-no-more-accurate-than", "canonical_source": "https://dev.to/kultzuki/the-agentic-rag-pipeline-that-was-faster-and-cheaper-and-no-more-accurate-than-no-agent-at-all-1gbg", "published_at": "2026-10-04 10:35:45+00:00", "updated_at": "2026-10-04 10:42:40.298326+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-research", "mlops"], "entities": ["TigerGraph", "GraphRAG", "Qwen3.8-27B", "GSQL"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/the-agentic-rag-pipeline-that-was-faster-and-cheaper-and-no-more-accurate-than", "markdown": "https://wpnews.pro/news/the-agentic-rag-pipeline-that-was-faster-and-cheaper-and-no-more-accurate-than.md", "text": "https://wpnews.pro/news/the-agentic-rag-pipeline-that-was-faster-and-cheaper-and-no-more-accurate-than.txt", "jsonld": "https://wpnews.pro/news/the-agentic-rag-pipeline-that-was-faster-and-cheaper-and-no-more-accurate-than.jsonld"}}