{"slug": "rag-architecture-diagram-the-boxes-and-what-each-arrow-costs", "title": "RAG Architecture Diagram: the Boxes and What Each Arrow Costs", "summary": "A developer outlined a five-box RAG architecture diagram — index, retriever, reranker, context builder, and loop controller — arguing that each connecting arrow carries a payload paid for twice: once when moved and once when read. The writeup frames the diagram as a per-request budget rather than an org chart, noting that candidate vectors and scored candidates grow with data volume even when traffic does not, and that the index's build step fixes the ceiling for everything downstream. It also cites Zep's temporal knowledge graph as an alternative index shape that stores facts with validity windows instead of passages.", "body_md": "**Short answer:** A RAG architecture diagram is five boxes and the arrows that join them - an index you build, a retriever that turns a query into candidates, a reranker that orders them, a context builder that packs what the model may read, and a loop controller that decides whether any of it runs again. Every arrow carries a payload, and every payload is paid for twice: once when it is moved, once when it is read.\n\n**Key takeaways**\n\nDraw the picture before choosing a vendor. A retrieval stack that survives production is five boxes joined by arrows, and each arrow carries a payload that someone pays for. The index holds what the retriever is allowed to find; the retriever turns one query into candidates; the reranker reorders those candidates; the context builder decides what the model finally reads; the loop controller decides whether the whole thing runs again. [the architecture this diagram draws](https://smartgate.network/industry/advanced-rag-architecture-for-ai-agents) walks the same stack from the implementation side, with the code each box runs and the reason it is written that way. This page is the diagram itself: the boxes, the arrows, and the price attached to each arrow.\n\n| Box | What it holds | The arrow out of it | What that arrow costs | \n|---|---|---|---|\n| Index | chunks, their vectors, the metadata a filter reads | candidate vectors | a build job, plus the storage behind it | \n| Retriever | no state - it is a query path | k candidate passages with scores | a query embedding plus k passages moved | \n| Reranker | no state | m ordered passages | one scoring pass per candidate | \n| Context builder | the packed prompt | the prompt the model reads | the tokens the model is billed for | \n| Loop controller | the state of the request | a rewritten query, or a stop | a full pass through every box above | \n| The layer around them | who is calling, and at what limit | a filter on retrieval | nothing, until it is missing | \n\nRead the table as a budget rather than an org chart. The left column is what you build once; the right column is what runs per request. The two arrows that grow with data volume - candidate vectors and scored candidates - are the ones to watch, because they grow even when traffic does not. And the last row is the one most diagrams leave out, which is why it is the one that shows up in incident reports.\n\nThe index is the only box with a build step, and it fixes the ceiling for everything downstream. If a fact was chunked away, split across two rows, or dropped because its payload was incomplete, no reranker and no prompt rewrite will bring it back - the retriever can only return what the index already holds.\n\nThree things travel together in every row: the vector, the text the model will eventually read, and the metadata a filter matches on. Two of them are cheap to change later. Swapping a metadata field means backfilling the existing rows; swapping the embedding model means re-embedding the whole corpus, because vectors from two models do not share a coordinate system. Chunking sits in the expensive column too, and it is the decision people make last: a 512-token window that cuts a table in half produces two vectors matching neither the table nor the question, and the retriever will confidently return both.\n\nA different index shape is legitimate. Zep's [temporal knowledge graph](https://arxiv.org/abs/2501.13956) stores facts with validity windows rather than passages, which answers \"what was true last Tuesday\" - a question a passage index cannot express - at the price of an extraction pipeline that has to run before anything is searchable. The diagram does not change; only what the first box holds changes.\n\nThe retriever is the shortest code in the stack and the most common place to lose an afternoon. Approximate nearest-neighbour search returns the k closest vectors, and that list is a ranking, not a verdict: a similarity of 0.82 says nothing about whether the passage answers the question. Two decisions belong to this box, and both of them are about the arrow it produces.\n\nFirst, filters belong inside the store rather than in application code afterwards. Post-filtering a k of 20 can leave two usable passages, and the effective k then changes with the data rather than with your configuration. Second, k is not a quality dial on its own: every extra candidate is one more passage the reranker has to score and one more chance for an irrelevant passage to reach the prompt. Measure recall at k of 5, 10 and 20 once, on your own questions, and you will know which of the two failure modes you have - starving the reranker, or diluting it.\n\nA reranker is the easiest box to delete and the easiest box to under-budget. A [cross-encoder](https://www.sbert.net/examples/applications/cross-encoder/README.html) reads the query and one passage together and returns a score for that pair, which means its work is linear in k: fifty candidates are fifty forward passes inside the request path. That is the trade in one sentence - precision at the top of the list, paid for in milliseconds that the user waits through.\n\nTwo conditions make the box worth its latency. The retriever must return more candidates than the prompt can hold, so that there is something to reorder; and the order must matter, which is true when the model reads the top few passages closely and ignores the rest. When the retriever already returns three passages that all belong in the answer, the reranker is decoration. When it returns thirty ranked by vector distance alone, the reranker is the difference between an answer and a plausible-sounding miss.\n\nThe context builder is where the diagram stops being free. Everything above it produces candidates; this box decides how many of them the model may read, and it is the only box in the stack that can shrink the bill without losing the answer. The excerpt below is the arithmetic: the target token count is either given directly or derived from a compression rate, and the rate is then handed to the segmentation step, which splits the context into segments that carry their own settings.\n\n```\n# backend/smartgate/modules/context_gate/algorithm.py — source lines 366–373 (PromptCompressor.structured_compress_prompt)\n        Returns:\n            dict: A dictionary containing:\n                - \"compressed_prompt\" (str): The resulting compressed prompt.\n                - \"origin_tokens\" (int): The original number of tokens in the input.\n                - \"compressed_tokens\" (int): The number of tokens in the compressed output.\n                - \"ratio\" (str): The compression ratio achieved, calculated as the original token number divided by the token number after compression.\n                - \"rate\" (str): The compression rate achieved, in a human-readable format.\n                - \"saving\" (str): Estimated savings in GPT-4 token usage.\n# backend/smartgate/modules/context_gate/algorithm.py — source lines 387–405 (PromptCompressor.structured_compress_prompt)\n        if target_token == -1:\n            target_token = (\n                (\n                    instruction_tokens_length\n                    + question_tokens_length\n                    + sum(context_tokens_length)\n                )\n                * rate\n                - instruction_tokens_length\n                - (question_tokens_length if concate_question else 0)\n            )\n        else:\n            rate = target_token / sum(context_tokens_length)\n        (\n            context,\n            context_segs,\n            context_segs_rate,\n            context_segs_compress,\n        ) = self.segment_structured_context(context, rate)\n```\n\nTwo details in that code are worth copying into your own diagram. The instruction and the question are counted before the rate is applied and subtracted afterwards, so the budget covers the context rather than the whole request - a mistake here quietly under-compresses by exactly the size of the instruction. And the return value is not the compressed text alone: it carries the original token count, the compressed token count, the achieved ratio and an estimated saving, which makes the box measurable in the same request that pays for it. The same five boxes also carry a set of choices - which one to swap, what each swap costs in latency, and which failure each choice invites - and that argument is the whole job of [the decisions behind these boxes](https://smartgate.network/industry/what-is-rag-architecture). If you want the wider catalogue of tricks before changing this box, [token optimization techniques](https://smartgate.network/integration/token-optimization-techniques-for-ai-apps) covers the ones that are not compression.\n\nOne compression rate for a whole request is a blunt instrument, and the code in most stacks assumes it. The alternative is visible in the next excerpt: the context arrives as segments, each tagged with its own rate and a compress flag, and the builder normalises, defaults and validates them before anything is compressed.\n\n```\n# backend/smartgate/modules/context_gate/algorithm.py — source lines 2100–2112 (PromptCompressor.segment_structured_context)\n        for text in context:\n            if not text.startswith(\"<llmlingua\"):\n                text = \"<llmlingua>\" + text\n            if not text.endswith(\"</llmlingua>\"):\n                text = text + \"</llmlingua>\"\n\n            # Regular expression to match <llmlingua, rate=x, compress=y>content</llmlingua>, allowing rate and compress in any order\n            pattern = r\"<llmlingua\\s*(?:,\\s*rate\\s*=\\s*([\\d\\.]+))?\\s*(?:,\\s*compress\\s*=\\s*(True|False))?\\s*(?:,\\s*rate\\s*=\\s*([\\d\\.]+))?\\s*(?:,\\s*compress\\s*=\\s*(True|False))?\\s*>([^<]+)</llmlingua>\"\n            matches = re.findall(pattern, text)\n\n            # Extracting segment contents\n            segments = [match[4] for match in matches]\n# backend/smartgate/modules/context_gate/algorithm.py — source lines 2127–2139 (PromptCompressor.segment_structured_context)\n            segs_compress = [\n                compress if compress is not None else True for compress in segs_compress\n            ]\n            segs_rate = [\n                rate if rate else (global_rate if compress else 1.0)\n                for rate, compress in zip(segs_rate, segs_compress)\n            ]\n            assert (\n                len(segments) == len(segs_rate) == len(segs_compress)\n            ), \"The number of segments, rates, and compress flags should be the same.\"\n            assert all(\n                seg_rate <= 1.0 for seg_rate in segs_rate\n            ), \"Error: 'rate' must not exceed 1.0. The value of 'rate' indicates compression rate and must be within the range [0, 1].\"\n```\n\nThe tags are an interface between boxes, and that is the architectural point rather than a detail. Whoever produces the context - the ingest step, the retriever, the loop controller - can mark a passage as untouchable, and the marking travels with the text instead of living in a configuration file that nobody re-reads. A passage holding a table, a worked example or a citation can be preserved whole while the surrounding prose is compressed hard. The defaults are equally deliberate: a segment with no rate inherits the global rate when compression is allowed and a rate of 1.0 when it is not, so the two settings can never contradict each other, and the asserts refuse a rate above 1.0 or a segment whose counts disagree. A stack that compresses uniformly is not wrong, it is simply leaving its easiest saving on the table - [context window management techniques](https://smartgate.network/industry/context-window-management-techniques) covers the surrounding budget questions.\n\nCompression is usually described as a text operation, and the second model to run in your request path is usually left out of the diagram. It is there in the excerpt: the candidate chunks become a dataset, the dataset is batched through a small trained model with gradients switched off, the logits go through a softmax, and the per-token probabilities are merged back up into per-word probabilities.\n\n```\n# backend/smartgate/modules/context_gate/algorithm.py — source lines 2191–2200 (PromptCompressor.__get_context_prob)\n        chunk_probs = []\n        chunk_words = []\n        with torch.no_grad():\n            for batch in dataloader:\n                ids = batch[\"ids\"].to(self.device, dtype=torch.long)\n                mask = batch[\"mask\"].to(self.device, dtype=torch.long) == 1\n\n                outputs = self.model(input_ids=ids, attention_mask=mask)\n                loss, logits = outputs.loss, outputs.logits\n                probs = F.softmax(logits, dim=-1)\n# backend/smartgate/modules/context_gate/algorithm.py — source lines 2215–2228 (PromptCompressor.__get_context_prob)\n                    (\n                        words,\n                        valid_token_probs,\n                        valid_token_probs_no_force,\n                    ) = self.__merge_token_to_word(\n                        tokens,\n                        token_probs,\n                        force_tokens=force_tokens,\n                        token_map=token_map,\n                        force_reserve_digit=force_reserve_digit,\n                    )\n                    word_probs_no_force = self.__token_prob_to_word_prob(\n                        valid_token_probs_no_force, convert_mode=token_to_word\n                    )\n```\n\nRead the cost of that box before adopting it. It is a second model, inside the request path, whose latency grows with the number of tokens you feed it - and whose output is only a ranking of tokens by how predictable they are, which the compressor then uses to decide what to drop. Nothing about it is a protocol feature: it runs in your process against your own weights, so it is invisible to the caller and to any gateway unless you log it. The configuration in the first excerpt is what keeps it honest: the token-level filter can be switched off entirely, and the context-level and sentence-level filters can be used alone, which is the right answer when the second model's latency is larger than the saving it produces.\n\nNothing in the diagram so far knows who is asking. Retrieval interfaces take a query and a filter; they do not take a principal, and no protocol specification turns a session into an authorization decision. So identity has to be carried into the stack by whichever box already knows it, and the excerpt below is the shape it usually has: a session callback that copies the user id out of the token subject and keeps the role alongside it.\n\n```\n# auth.ts — source lines 39–48 (session)\nasync session({ token, session }) {\n      if (session.user) {\n        if (token.sub) session.user.id = token.sub;\n        (session.user as { role?: UserRole }).role = (token.role as UserRole) || UserRole.USER;\n        session.user.name = token.name ?? null;\n        session.user.email = token.email ?? \"\";\n        session.user.image = token.picture ?? null;\n      }\n      return session;\n    }\n```\n\nThat copied id is the only thing standing between a shared index and a per-tenant leak, and it is a filter value rather than a security boundary. Two consequences follow. The id has to reach the retriever, which means the filter is built at the edge of the stack rather than deep inside the vector code, where it would have to be threaded through every call site. And it has to be recorded with the request, because a retrieval that cannot name its principal cannot be audited afterwards. Agent memory has the same shape - a store keyed by user, agent and run - and [agent memory architecture](https://smartgate.network/industry/agent-memory-architecture) works through the write path when the store fills up.\n\nAn architecture diagram usually stops at the model's answer. The arrow that matters operationally is the one leaving the application: what the caller receives when one of the boxes fails. The excerpt is the mapping - a result object becomes a response envelope, and if the envelope says the call failed, the caller gets an HTTP error with the failure detail attached instead of a partial answer.\n\n```\n# backend/smartgate/core/models.py — source lines 38–46 (smartgate_http_response_from_result)\ndef smartgate_http_response_from_result(result, *, status_code: int = 422):\n    \"\"\"Map ToolResult → API envelope; failed tools become HTTP errors (REST clients).\"\"\"\n    from fastapi import HTTPException\n\n    response = smartgate_response_from_result(result)\n    if not response.success:\n        detail = response.error or {\"code\": \"ERROR\", \"message\": \"request failed\"}\n        raise HTTPException(status_code=status_code, detail=detail)\n    return response\n```\n\nOne status per failure family beats one status for everything, because the transport layer is where retries, budgets and alerts are decided. A caller that can tell a validation failure from a capacity failure from an upstream timeout retries the right things and pages the right person; a caller that receives a 200 with an empty body retries everything and learns nothing. The same envelope discipline is what lets a gateway sit in front of the stack at all, which is the argument in [what a gateway adds](https://smartgate.network/industry/mcp-gateway): limits, keys and audit rows only make sense if a failed call is visible as a failure.\n\nEvery arrow in the diagram is a contract, and the cheapest way to break one is to return nothing. The excerpt is the smallest version of that bug and its fix, from a documentation source config: a node with no children is given a single space, so the renderer below it never has to handle an empty list.\n\n```\n# source.config.ts — source lines 49–53 (onVisitLine)\nonVisitLine(node: { children: unknown[] }) {\n            if (node.children.length === 0) {\n              node.children = [{ type: \"text\", value: \" \" }];\n            }\n          }\n```\n\nThat is a guard against an absent shape, not against empty content. An empty result is fine; a missing field is not. In a retrieval stack the same rule applies three times over: a retriever that finds nothing must return an empty list with the same fields rather than a null; a context builder that drops every passage must still emit a prompt, even a short one that says so; and a loop controller that stops early must still return the shape the caller expects. Teams that skip this spend their first week of production debugging missing keys instead of missing answers, and an agentic stack multiplies the surface, because every extra box and every extra iteration is another chance to hand the next box a hole - [what agentic RAG changes](https://smartgate.network/industry/what-is-agentic-rag) is the version of this diagram where the loop owns more of the boxes.\n\nEverything so far runs once. The loop controller is the box that decides to run it all again, and it is the only component in the diagram whose failure mode is a bill rather than a wrong answer. A second iteration means a rewritten query, a fresh retrieval, another rerank and another packed prompt; the marginal cost is the whole stack, and the marginal benefit is the one passage the first pass missed.\n\nThree limits belong to this box rather than to the caller, because the caller cannot see the loop. A cap on iterations, because the third pass almost never pays. A cap on tokens per request, because the loop's cost is dominated by what each pass carries into the prompt. And a budget check before the second pass rather than after it, because a check that runs at the end of the request reports the overrun instead of preventing it. Measuring this box is a cost question with a known shape - [what a retrieval loop costs](https://smartgate.network/integration/optimize-ai-agent-execution-cost) takes the same stack apart in money - and the loop is also the only place where a retrieval stack becomes an agentic one, which is the line [RAG versus agentic RAG](https://smartgate.network/industry/rag-vs-agentic-rag) draws in detail.\n\nOne confusion costs teams an integration. A tool server exposes capabilities a model can call with arguments, and a directory of such servers is a catalogue of who offers what. A retriever answers a different question: given a query, which passages in my corpus are worth reading. Both end up inside the context window, and only one of them has a reranker, a similarity threshold and a k to tune.\n\nThe practical test is the arrow test. If the answer arrives because the model asked for it with a specific argument, it is a tool call, and [the tool surface a host exposes](https://smartgate.network/industry/mcp-tools-reference) describes how those arguments and results are shaped. If the answer arrives because the system searched for it before the model was invoked, it is retrieval. A server list can be a source of documents - fetch a page with a tool, hand it to the ingest step, index it - but the list itself is not an index, and no reranker will ever see it. The academic version of that boundary, and the design space on both sides of it, is the subject of [the agentic RAG survey](https://smartgate.network/industry/agentic-rag-survey).\n\nWhich of these boxes you own at all — and which belong to a runtime that decides for itself when to\n\nretrieve — is the distinction drawn in [rag vs agentic ai](https://smartgate.network/industry/rag-vs-agentic-ai), which is\n\nthe page to read when the diagram and the deployment disagree about who is in control.\n\nThe interesting comparison is not which box is better but who owns which box. The five-box diagram is the same for everyone; the difference is how much of it you have to build, monitor and pay for.\n\n|  | What it owns | What you still own | \n|---|---|---|\n| A hand-rolled pipeline | nothing until you write it | all five boxes, the rerank, the budget and the logs | \n| A managed vector service | the index and the retriever | chunking, reranking, the context budget, the loop | \n| A model provider's own retrieval | one index inside one vendor | everything outside that vendor | \n| SmartGate | the context builder, deduplication, a hard budget guard, an audit row per call and the loop's limits | your corpus, your chunking and your index if you keep one | \n\nThe reason to hand over the middle of the stack is that the middle is where the money leaks: compression, deduplication and the loop's caps are the boxes that decide the size of the prompt. Free tier covers 2 million tokens a month with all seven tools and no card; Pro starts at 18 dollars a month, Teams at 55, with the same per-key limits on every plan.\n\nStart on the free tier - 2 million tokens a month and all seven tools - from [start free](https://smartgate.network/login?from=/dashboard); the per-plan limits sit on the [pricing page](https://smartgate.network/pricing), the endpoint shape is in the [product docs](https://smartgate.network/docs), and contract traffic begins at the [contact form](https://smartgate.network/contact).\n\n**What are the five boxes of a RAG architecture?**\n\nAn index, a retriever, a reranker, a context builder and a loop controller. The index holds chunks, vectors and filter metadata; the retriever turns a query into candidates; the reranker reorders them; the context builder packs what the model reads; the loop controller decides whether the pass runs again. Identity, limits and logs sit around the boxes rather than inside them.\n\n**Is a reranker always worth the latency?**\n\nNo. It is worth it when the retriever returns more candidates than the prompt can hold and the order of those candidates changes the answer. A cross-encoder scores one query and one passage at a time, so its cost grows with the number of candidates. If the retriever already returns three passages that all belong in the answer, the reranker is decoration.\n\n**Where does a retrieval stack usually break first?**\n\nAt the arrow into the prompt. The index is the ceiling on recall, but the prompt is the ceiling on the bill, and a stack that retrieves well and packs badly pays for tokens the model does not need. The second most common break is a box that returns an empty result with the wrong shape, which surfaces as a missing key rather than a missing answer.\n\n**Can I skip the index and call a search API instead?**\n\nYes, and for a small corpus it is the right first move: a hosted search API becomes your index and retriever in one, and you keep the context builder and the loop. What you give up is control of chunking, filter semantics and score calibration, which are exactly the three knobs you need when recall stops improving.\n\n**How do I know the context builder helped?**\n\nMeasure the token count before and after the same request. A builder worth keeping reports the original token count, the compressed token count and the achieved ratio, so the saving is an observation rather than a belief. Compare the trimmed prompt with the untrimmed one on the same questions: if the answers are identical and the token count halved, the box is doing its job.\n\n**Does the diagram change for an agentic stack?**\n\nThe boxes stay and the loop grows. An agentic system adds iterations, so the loop controller owns more of the cost and needs caps of its own, and every extra box is another shape that has to reach the next box intact. The arrows still carry the same payloads; there are simply more passes over them.\n\nThe code in this article is not transcribed. Each block was cut directly out of the slice body returned by the SmartGate slice API, then re-asserted byte for byte as a substring of that body before publication; the first line inside every fence records the file and the exact source lines. Symbols were pinned with whole-name containment (rule A level 2) and confirmed by the service's slot-proof endpoint - 6 of 8 planned sections pinned, no abstentions, and the two sections that matched no unique symbol are written from published sources with no code, as the abstain rule requires. Some blocks show two line windows rather than one: the gap between them is the part of the function that does not serve this page's diagram, so it is not quoted. The excerpts are one implementation's, chosen because the boxes in this order exist there in code rather than in prose.\n\n| # | SERP keyword | Symbol | File | Source lines | How it was pinned | sha256(12) | \n|---|---|---|---|---|---|---|\n| 1 | model context protocol diagram | `PromptCompressor.structured_compress_prompt` | `backend/smartgate/modules/context_gate/algorithm.py` | 366–373, 387–405 | rule A L2 → slot-proof | `4900c4d45a58` | \n| 2 | model context protocol architecture | `PromptCompressor.segment_structured_context` | `backend/smartgate/modules/context_gate/algorithm.py` | 2100–2112, 2127–2139 | rule A L2 → slot-proof | `298956ac6c1e` | \n| 3 | model context protocol news | `PromptCompressor.__get_context_prob` | `backend/smartgate/modules/context_gate/algorithm.py` | 2191–2200, 2215–2228 | rule A L2 → slot-proof | `dce811214e89` | \n| 4 | mcp specification | `session` | `auth.ts` | 39–48 | rule A L2 → slot-proof | `109468a80ae0` | \n| 5 | streamable http | `smartgate_http_response_from_result` | `backend/smartgate/core/models.py` | 38–46 | rule A L2 → slot-proof | `f715442265ed` | \n| 6 | rag architecture explained | `onVisitLine` | `source.config.ts` | 49–53 | rule A L2 → slot-proof | `b3fc10125ee9` | \n\nEvery fenced block above was cut from the slice body and re-asserted against it byte for byte before\n\npublication. 6 of 8 sections pinned, 0 abstentions, 2 misses.", "url": "https://wpnews.pro/news/rag-architecture-diagram-the-boxes-and-what-each-arrow-costs", "canonical_source": "https://dev.to/smartgate/rag-architecture-diagram-the-boxes-and-what-each-arrow-costs-58ag", "published_at": "2026-09-30 07:06:17+00:00", "updated_at": "2026-09-30 07:16:37.412709+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "mlops", "ai-infrastructure"], "entities": ["Zep"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/rag-architecture-diagram-the-boxes-and-what-each-arrow-costs", "markdown": "https://wpnews.pro/news/rag-architecture-diagram-the-boxes-and-what-each-arrow-costs.md", "text": "https://wpnews.pro/news/rag-architecture-diagram-the-boxes-and-what-each-arrow-costs.txt", "jsonld": "https://wpnews.pro/news/rag-architecture-diagram-the-boxes-and-what-each-arrow-costs.jsonld"}}