RAG Architecture Diagram: the Boxes and What Each Arrow Costs A developer outlined a five-box RAG architecture diagram — index, retriever, reranker, context builder, and loop controller — arguing that each connecting arrow carries a payload paid for twice: once when moved and once when read. The writeup frames the diagram as a per-request budget rather than an org chart, noting that candidate vectors and scored candidates grow with data volume even when traffic does not, and that the index's build step fixes the ceiling for everything downstream. It also cites Zep's temporal knowledge graph as an alternative index shape that stores facts with validity windows instead of passages. Short answer: A RAG architecture diagram is five boxes and the arrows that join them - an index you build, a retriever that turns a query into candidates, a reranker that orders them, a context builder that packs what the model may read, and a loop controller that decides whether any of it runs again. Every arrow carries a payload, and every payload is paid for twice: once when it is moved, once when it is read. Key takeaways Draw the picture before choosing a vendor. A retrieval stack that survives production is five boxes joined by arrows, and each arrow carries a payload that someone pays for. The index holds what the retriever is allowed to find; the retriever turns one query into candidates; the reranker reorders those candidates; the context builder decides what the model finally reads; the loop controller decides whether the whole thing runs again. the architecture this diagram draws https://smartgate.network/industry/advanced-rag-architecture-for-ai-agents walks the same stack from the implementation side, with the code each box runs and the reason it is written that way. This page is the diagram itself: the boxes, the arrows, and the price attached to each arrow. | Box | What it holds | The arrow out of it | What that arrow costs | |---|---|---|---| | Index | chunks, their vectors, the metadata a filter reads | candidate vectors | a build job, plus the storage behind it | | Retriever | no state - it is a query path | k candidate passages with scores | a query embedding plus k passages moved | | Reranker | no state | m ordered passages | one scoring pass per candidate | | Context builder | the packed prompt | the prompt the model reads | the tokens the model is billed for | | Loop controller | the state of the request | a rewritten query, or a stop | a full pass through every box above | | The layer around them | who is calling, and at what limit | a filter on retrieval | nothing, until it is missing | Read the table as a budget rather than an org chart. The left column is what you build once; the right column is what runs per request. The two arrows that grow with data volume - candidate vectors and scored candidates - are the ones to watch, because they grow even when traffic does not. And the last row is the one most diagrams leave out, which is why it is the one that shows up in incident reports. The index is the only box with a build step, and it fixes the ceiling for everything downstream. If a fact was chunked away, split across two rows, or dropped because its payload was incomplete, no reranker and no prompt rewrite will bring it back - the retriever can only return what the index already holds. Three things travel together in every row: the vector, the text the model will eventually read, and the metadata a filter matches on. Two of them are cheap to change later. Swapping a metadata field means backfilling the existing rows; swapping the embedding model means re-embedding the whole corpus, because vectors from two models do not share a coordinate system. Chunking sits in the expensive column too, and it is the decision people make last: a 512-token window that cuts a table in half produces two vectors matching neither the table nor the question, and the retriever will confidently return both. A different index shape is legitimate. Zep's temporal knowledge graph https://arxiv.org/abs/2501.13956 stores facts with validity windows rather than passages, which answers "what was true last Tuesday" - a question a passage index cannot express - at the price of an extraction pipeline that has to run before anything is searchable. The diagram does not change; only what the first box holds changes. The retriever is the shortest code in the stack and the most common place to lose an afternoon. Approximate nearest-neighbour search returns the k closest vectors, and that list is a ranking, not a verdict: a similarity of 0.82 says nothing about whether the passage answers the question. Two decisions belong to this box, and both of them are about the arrow it produces. First, filters belong inside the store rather than in application code afterwards. Post-filtering a k of 20 can leave two usable passages, and the effective k then changes with the data rather than with your configuration. Second, k is not a quality dial on its own: every extra candidate is one more passage the reranker has to score and one more chance for an irrelevant passage to reach the prompt. Measure recall at k of 5, 10 and 20 once, on your own questions, and you will know which of the two failure modes you have - starving the reranker, or diluting it. A reranker is the easiest box to delete and the easiest box to under-budget. A cross-encoder https://www.sbert.net/examples/applications/cross-encoder/README.html reads the query and one passage together and returns a score for that pair, which means its work is linear in k: fifty candidates are fifty forward passes inside the request path. That is the trade in one sentence - precision at the top of the list, paid for in milliseconds that the user waits through. Two conditions make the box worth its latency. The retriever must return more candidates than the prompt can hold, so that there is something to reorder; and the order must matter, which is true when the model reads the top few passages closely and ignores the rest. When the retriever already returns three passages that all belong in the answer, the reranker is decoration. When it returns thirty ranked by vector distance alone, the reranker is the difference between an answer and a plausible-sounding miss. The context builder is where the diagram stops being free. Everything above it produces candidates; this box decides how many of them the model may read, and it is the only box in the stack that can shrink the bill without losing the answer. The excerpt below is the arithmetic: the target token count is either given directly or derived from a compression rate, and the rate is then handed to the segmentation step, which splits the context into segments that carry their own settings. backend/smartgate/modules/context gate/algorithm.py — source lines 366–373 PromptCompressor.structured compress prompt Returns: dict: A dictionary containing: - "compressed prompt" str : The resulting compressed prompt. - "origin tokens" int : The original number of tokens in the input. - "compressed tokens" int : The number of tokens in the compressed output. - "ratio" str : The compression ratio achieved, calculated as the original token number divided by the token number after compression. - "rate" str : The compression rate achieved, in a human-readable format. - "saving" str : Estimated savings in GPT-4 token usage. backend/smartgate/modules/context gate/algorithm.py — source lines 387–405 PromptCompressor.structured compress prompt if target token == -1: target token = instruction tokens length + question tokens length + sum context tokens length rate - instruction tokens length - question tokens length if concate question else 0 else: rate = target token / sum context tokens length context, context segs, context segs rate, context segs compress, = self.segment structured context context, rate Two details in that code are worth copying into your own diagram. The instruction and the question are counted before the rate is applied and subtracted afterwards, so the budget covers the context rather than the whole request - a mistake here quietly under-compresses by exactly the size of the instruction. And the return value is not the compressed text alone: it carries the original token count, the compressed token count, the achieved ratio and an estimated saving, which makes the box measurable in the same request that pays for it. The same five boxes also carry a set of choices - which one to swap, what each swap costs in latency, and which failure each choice invites - and that argument is the whole job of the decisions behind these boxes https://smartgate.network/industry/what-is-rag-architecture . If you want the wider catalogue of tricks before changing this box, token optimization techniques https://smartgate.network/integration/token-optimization-techniques-for-ai-apps covers the ones that are not compression. One compression rate for a whole request is a blunt instrument, and the code in most stacks assumes it. The alternative is visible in the next excerpt: the context arrives as segments, each tagged with its own rate and a compress flag, and the builder normalises, defaults and validates them before anything is compressed. backend/smartgate/modules/context gate/algorithm.py — source lines 2100–2112 PromptCompressor.segment structured context for text in context: if not text.startswith "