Building Knowledge Vault : From RAG Implementation to Retrieval Engineering Jaival Suthar released Knowledge Vault (M1), a local-first PDF retrieval system built on PyMuPDF text extraction, token-aware recursive chunking, BAAI/bge-small-en-v1.5 embeddings stored in Qdrant, and optional BAAI/bge-reranker-base reranking, as the retrieval layer of a local AI stack that follows the Inference Lab (M0) model-serving layer. The project deliberately excludes OCR for scanned PDFs and advanced RAG features, focusing instead on making retrieval behavior inspectable with source-attributed, page-level metadata. Knowledge Vault grounds generation in retrieved evidence but does not guarantee the final answer is faithful to that evidence. Most RAG demos are deceptively simple. Ingest a document. Chunk it. Generate embeddings. Store the vectors. Retrieve a few chunks. Send them to an LLM. The architecture fits neatly on a screen. The difficult part starts when you stop asking, “Does it work?” and start asking, “Why does it work when it does, and why does it fail when it doesn’t?” That’s what I wanted to understand with Knowledge Vault. Knowledge Vault is Milestone 1 M1 of my local AI stack: a local-first PDF retrieval system that turns books and documents into searchable, source-attributed context for an LLM. It builds on Inference Lab M0 , my local model-serving layer which is responsible for running the model. Knowledge Vault is responsible for everything that happens before the model sees the prompt: document ingestion, retrieval, reranking, and context construction. 🔗 Knowledge Vault M1 — GitHub https://github.com/Jaival-Suthar/pdf-rag-system 🔗 Inference Lab M0 — GitHub https://github.com/Jaival-Suthar/inference-lab But I wasn’t trying to build another feature-heavy RAG demo. I wanted to make the retrieval system understandable. So I gave Knowledge Vault a deliberately narrow objective: Build the retrieval layer, understand its behavior, and find out where it breaks. For the project, success meant four things: I deliberately left out the usual temptation to keep adding “advanced RAG” features. First, I wanted to know whether the foundations actually worked. First, I wanted to answer a simpler question: Does the retrieval system I actually built work? And more importantly: Can I prove it? Knowledge Vault has a deliberate boundary between knowledge retrieval and model inference. The retrieval system is responsible for turning documents into useful context. A separate inference service is responsible for executing the language model. The high-level architecture looks like this: The system attempts to ground generation in retrieved evidence, but Knowledge Vault does not guarantee that the final answer is faithful to that evidence. The first stage uses PyMuPDF for page-by-page PDF text extraction. This is intentionally a text-layer PDF pipeline. It preserves source metadata such as document identity and page information so that retrieved chunks can be traced back to where they came from. That traceability matters because a RAG system shouldn’t just return an answer. It should be possible to inspect the evidence behind that answer. It also establishes an important boundary for M1: scanned PDFs requiring OCR are not the problem being solved yet. Extracted text is passed through recursive, token-aware chunking. The goal is not simply to cut the document every N characters. A retrieval chunk should be large enough to preserve useful context but small enough to remain a precise retrieval unit. That creates an unavoidable trade-off. Chunks that are too small can destroy context: Chunk A: "The system uses..."Chunk B: "...a two-stage retrieval process." Chunks that are too large can dilute relevance: Question ↓Large chunk containing five unrelated concepts ↓Weak retrieval precision So chunking is not just preprocessing. It is already a retrieval decision. Each chunk is converted into a vector using: BAAI/bge-small-en-v1.5 The same embedding model is used for the user’s query. Conceptually: Document chunk Question ↓ ↓Embedding model Embedding model ↓ ↓ Vector Query vector The purpose is to move the problem from lexical matching toward semantic similarity. The chunk vectors are stored in Qdrant along with metadata such as: That metadata becomes important later because vector similarity alone doesn’t tell us whether a retrieved result is actually useful evidence. The retrieval pipeline can optionally add: BAAI/bge-reranker-base The architecture is intentionally two-stage: The vector database is responsible for quickly finding candidates. The reranker gets a smaller candidate set and performs a more expensive query-passage relevance calculation. That makes the reranker more than a quality component. It is also a latency-bearing architectural component. This is a classic retrieval architecture, but I didn’t want to assume that adding the second stage automatically made the system better. That had to be measured. The other major architectural choice was separating Knowledge Vault from the model-serving layer. It would have been simpler to put everything into one application: PDF → RAG → Ollama → Answer Instead, the system uses a service boundary: The two systems have different responsibilities. Knowledge Vault is primarily a knowledge-processing and retrieval system. The inference service is a model-execution system. This means I can change the retrieval pipeline without redesigning model serving, and I can change the model runtime without rewriting the document pipeline. There is still coupling. It simply becomes explicit API coupling rather than shared implementation coupling. That distinction matters once a system grows beyond a single experiment. One of the first things Knowledge Vault made clear is that the LLM is actually quite far downstream in the failure chain. Consider the pipeline: Every stage can introduce a different failure. This is why I started thinking about RAG less as a single algorithm and more as a chain of failure boundaries. Understanding those boundaries became one of the central goals of Knowledge Vault. The first version of Knowledge Vault could answer questions over a PDF. That was encouraging. But: “I asked it a question and got a good answer” isn’t an evaluation methodology. It’s a demo. I wanted a controlled experiment. So I selected one real document: The 10X Rule, a 214-page book. Rather than immediately indexing many different documents, I deliberately constrained the experiment. The experiment was run locally using the same document, question set, inference configuration, and final context size across runs. Retrieval configuration was changed only where required by the experiment. Latency was measured per pipeline stage rather than from a single end-to-end timer. The benchmark is intentionally small and controlled; it is intended to compare configurations within this system, not to claim general RAG performance across datasets. This reduces the number of variables and makes it easier to understand why something changed. I created a set of 35 questions designed to exercise different retrieval behaviors. The questions included: The last category is important. A RAG system should not assume every question has an answer in its source. If the document doesn’t contain the required information, the desired behavior is: Question → Retrieve → No sufficient evidence → No grounded answer rather than: Question → LLM → Plausible-sounding answer The deliberately unanswerable questions were included to test how retrieval and generation behaved when the source did not contain sufficient evidence. The most important detail of the experiment is that this was not simply “reranker off versus reranker on.” There were three runs. The comparison was deliberately controlled: The baseline path was: Question → BGE embedding → Qdrant → Top-K candidates → LLM This establishes what the vector retrieval layer can do on its own. The second run added the cross-encoder: Question → BGE embedding → Qdrant → Candidate chunks → BGE reranker → Final Top-K → LLM The initial candidate pool was small. The hypothesis was straightforward: if vector similarity gives us reasonable candidates but imperfect ordering, a cross-encoder should be able to improve the final ranking. The third run changed another important variable. Instead of giving the reranker only the initial small candidate set, I changed: rerank candidate k = 20 while keeping the final context size at the smaller Top-K. The architecture became: This distinction matters. That means Run 3 tests a real retrieval design question: Does giving the reranker a larger candidate pool allow it to recover better evidence? If the correct passage is ranked 8th by vector similarity, a reranker that sees only five candidates can never select it. Increasing the candidate pool to 20 gives it that opportunity. But more candidates also mean more distractors. So the experiment is not simply “5 is bad, 20 is good.” It is a candidate-coverage-versus-ranking trade-off. I realized I was actually evaluating several different properties of the system: These are related, but they are not the same metric. A system can retrieve a relevant passage and still generate the wrong answer. A reranker can change the ranking without improving correctness. And a system can retrieve a highly similar passage that is not actually useful evidence. Before looking at individual failures, I wanted to understand the system at the aggregate level. The reranker changed the Top-1 retrieved result on 27 of 35 queries — that’s 77.1%. Only 8 queries retained the same Top-1 result. At first glance, that sounds like a strong result. A component changed the ranking on more than three quarters of the questions. It would be very easy to write: “The reranker improved retrieval by 77.1%.” But that would be an invalid conclusion. The experiment measured ranking changes, not correctness. A ranking can change from bad → good. It can also change from good → bad. So I needed to inspect what actually moved. Changing rerank candidate k from the smaller candidate pool to 20 produced Top-1 changes on only 6 of 35 queries. This is a useful result precisely because it isn’t spectacular. Increasing the candidate pool gave the reranker more information. But it did not automatically produce better Top-1 retrieval. Only one of these changes was clearly an improvement. Several moved the system toward less useful evidence. And one was an unanswerable query, so changing its retrieved page isn’t inherently a retrieval improvement. The conclusion is therefore: A larger reranking candidate pool increases the search space, but it does not guarantee better evidence selection. More candidates can create more opportunities. They can also create more distractions. The other half of the experiment was performance. I instrumented the pipeline stage by stage rather than measuring only end-to-end latency. The results were revealing. The reranker itself added approximately: This produced a useful architectural insight. I initially expected vector retrieval to be one of the expensive parts of the system. It wasn’t. Qdrant retrieval was around 8–9 ms p50, 12–14 ms p95. The cross-encoder was around ~2.9 seconds per query. The LLM was still the largest stage overall, at roughly 13–14 seconds p50. But the reranker was no longer a negligible preprocessing step. It became a meaningful latency decision. Averages can hide the behavior users actually experience. Some queries were much more expensive in the reranking stage. For example: This changes the engineering question. It is no longer simply: “Does reranking work?” It becomes: “Under what conditions is several additional seconds of reranking worth the retrieval benefit?” That is a much better engineering question because it forces quality and performance into the same decision. Now came the most interesting part of the experiment. Consider a question about the four degrees of action. The book’s Chapter 7 contains the actual explanation. The relevant concepts include: That is the evidence the LLM actually needs. But the book’s Table of Contents also contains: From a semantic matching perspective, the Contents page looks extremely relevant. From an evidence perspective, it is weak. The reranker sometimes preferred exactly that kind of page. For one query, the baseline retrieved substantive Chapter 7 content around page 45, while reranking moved the Top-1 result to the Contents page. For another, the baseline again found Chapter 7 content while the reranker promoted an Index entry around page 196. This is not random noise. It exposes a structural limitation. This became one of the central lessons from the experiment. A reranker evaluates the relationship between: query ↔ passage But a RAG system needs to optimize something closer to: query ↔ passage ↔ answerability ↔ evidence usefulness A Table of Contents can have high semantic relevance because it contains the exact concepts mentioned in the question. An Index can have even more keyword-dense matches. But neither necessarily contains enough information to answer the question. So the system has a subtle failure mode: The problem isn’t necessarily that the reranker doesn’t understand the query. It is that document structure is not the same thing as semantic relevance. A Contents page, Index page, Glossary, Copyright page, and chapter body have very different evidentiary value. The reranker doesn’t automatically know that. The retrieval failures were not the only thing the evaluation exposed. One generated answer was also factually wrong when checked against the source. The question asked: What are the four major mistakes people make when setting goals? The generated answer produced a plausible list involving things such as: The answer sounded reasonable. But the book’s actual four mistakes are: This was a genuine correctness failure. And it exposed another boundary: A fluent answer is not proof that the RAG pipeline retrieved or used the correct evidence. The LLM can produce an answer that sounds consistent with the topic while still contradicting the source. That is why final-answer evaluation alone is insufficient. At this point I had some compelling measurements: But one number was still missing: retrieval accuracy. I could not honestly say: “Reranking improved retrieval accuracy by X%,” because I had not yet established machine-checked passage-level ground truth for every query. And this distinction matters. The experiment proves: It does not yet justify a universal claim that reranking improves retrieval quality. That restraint is part of the result. This experiment changed how I think about RAG evaluation. There are at least three separate questions: 1. Did retrieval find the evidence? Question → Candidate set → Is the correct evidence present? This measures candidate-generation quality. 2. Did reranking put the evidence in a useful position? Candidate set → Reranker → Did the correct evidence move upward? This measures ranking quality. 3. Did the LLM answer from the evidence? Retrieved evidence → LLM → Is the answer supported? This measures grounded answer quality. And there is a fourth question: 4. How did the system behave when sufficient evidence was absent? That matters for the deliberately unanswerable questions. A useful evaluation therefore looks more like: Query ↓Candidate Retrieval ↓Reranker Quality ↓Evidence Quality ↓Answer Grounding ↓Abstention Quality A single “answer accuracy” number hides too much. Assumption 1: Adding a reranker should improve retrieval What happened: The reranker changed the Top-1 result on 27/35 queries. What I learned: A ranking change is not an improvement metric. It has to be evaluated against ground truth. Assumption 2: The vector database would be expensive What happened: Qdrant retrieval was around 8–9 ms p50. What I learned: The vector database was not the bottleneck I expected. Assumption 3: Better semantic relevance means better context What happened: The reranker sometimes promoted Contents and Index pages. What I learned: Semantic relevance and evidentiary usefulness are different objectives. Assumption 4: A correct-looking final answer means retrieval worked What happened: At least one answer sounded plausible but was wrong when checked against the source. What I learned: Retrieval and generation must be evaluated independently. Assumption 5: More candidates should give the reranker more chances to succeed What happened: rerank candidate k = 20 changed only 6/35 Top-1 results, and most of those changes were not improvements. What I learned: More candidates increase opportunity and search space at the same time. The next improvement should not be “use a bigger model.” It should be a more precise evaluation and a targeted retrieval change. First, I want passage-level ground truth for the 35 questions: Then I can calculate meaningful retrieval metrics: For reranking, I want to compare multiple candidate-pool sizes and determine where the correct evidence first appears. Separately, I want to evaluate whether the final answer is actually supported by the retrieved evidence. And for unanswerable questions, I want to measure whether the system avoids producing unsupported answers. Only then can I make a strong claim about whether a particular retrieval configuration is better. The observed Contents and Index failures suggest a concrete hypothesis. Before the reranker sees candidates, Knowledge Vault could incorporate document structure: Qdrant → Candidate chunks → Structural filtering → Reranker → Final context Pages classified as: could be down-ranked or excluded when they are unlikely to provide answer-bearing evidence. But I don’t want to call that a solution yet. It is a hypothesis. The correct engineering process is: Failure → Hypothesis → Implement → Re-run same benchmark → Compare → Keep or reject That keeps the system honest. Knowledge Vault started as a RAG implementation exercise. It became an evaluation exercise. And the most useful thing I learned wasn’t a particular embedding model, vector database, or reranker. It was that a RAG system has several different meanings of “relevance.” A passage can be: Those are not the same property. The experiment made that visible. A Table of Contents can be highly relevant to a question and still be poor evidence. A larger candidate pool can expose the reranker to more relevant chunks and more distractors. A reranker can substantially change rankings without necessarily improving correctness. And an LLM can produce a convincing answer even when the retrieval layer has made a poor choice. That is why I no longer think of RAG as: PDF → Vector DB → LLM I think of it as: Document → Extraction → Chunking → Representation → Candidate Retrieval →Ranking → Evidence Selection → Generation → Verification Every stage has its own failure modes. I started with a working PDF RAG system. Then I instrumented it. I created a controlled 35-question evaluation. I compared retrieval configurations. I measured latency instead of guessing where the bottleneck was. I increased rerank candidate k from the smaller candidate pool to 20 and observed what actually changed. I inspected the retrieved pages against the original source. And I found cases where a component that looked useful at first glance was actually selecting worse evidence. I also found an answer that sounded correct but failed source verification. Those are not failures of the project. Those are the results of the experiment. The system doesn’t need to be perfect for the milestone to be successful. It needs to become understandable. And after this experiment, I understand a lot more about where the retrieval system succeeds, where it fails, what the reranker actually costs, and which assumptions need to be tested next. That is a much more useful result than a benchmark number that simply says: “RAG works.” Because the real engineering question was never whether I could make a RAG system answer a question. It was whether I could build one, measure it , challenge it, find where it breaks, and use those failures to make the next decision. That is what Knowledge Vault was really about. Building Knowledge Vault : From RAG Implementation to Retrieval Engineering https://pub.towardsai.net/building-knowledge-vault-rag-retrieval-engineering-b540161cda6f was originally published in Towards AI https://pub.towardsai.net on Medium, where people are continuing the conversation by highlighting and responding to this story.