{"slug": "building-knowledge-vault-from-rag-implementation-to-retrieval-engineering", "title": "Building Knowledge Vault : From RAG Implementation to Retrieval Engineering", "summary": "Jaival Suthar released Knowledge Vault (M1), a local-first PDF retrieval system built on PyMuPDF text extraction, token-aware recursive chunking, BAAI/bge-small-en-v1.5 embeddings stored in Qdrant, and optional BAAI/bge-reranker-base reranking, as the retrieval layer of a local AI stack that follows the Inference Lab (M0) model-serving layer. The project deliberately excludes OCR for scanned PDFs and advanced RAG features, focusing instead on making retrieval behavior inspectable with source-attributed, page-level metadata. Knowledge Vault grounds generation in retrieved evidence but does not guarantee the final answer is faithful to that evidence.", "body_md": "Most RAG demos are deceptively simple.\n\nIngest a document. Chunk it. Generate embeddings. Store the vectors. Retrieve a few chunks. Send them to an LLM.\n\nThe architecture fits neatly on a screen.\n\nThe difficult part starts when you stop asking, “Does it work?” and start asking, “Why does it work when it does, and why does it fail when it doesn’t?”\n\nThat’s what I wanted to understand with Knowledge Vault.\n\nKnowledge Vault is Milestone 1 (M1) of my local AI stack: a local-first PDF retrieval system that turns books and documents into searchable, source-attributed context for an LLM.\n\nIt builds on Inference Lab (M0), my local model-serving layer which is responsible for running the model. Knowledge Vault is responsible for everything that happens before the model sees the prompt: document ingestion, retrieval, reranking, and context construction.\n\n🔗 [Knowledge Vault (M1) — GitHub](https://github.com/Jaival-Suthar/pdf-rag-system)\n\n🔗 [Inference Lab (M0) — GitHub](https://github.com/Jaival-Suthar/inference-lab)\n\nBut I wasn’t trying to build another feature-heavy RAG demo.\n\nI wanted to make the retrieval system understandable.\n\nSo I gave Knowledge Vault a deliberately narrow objective:\n\nBuild the retrieval layer, understand its behavior, and find out where it breaks.\n\nFor the project, success meant four things:\n\nI deliberately left out the usual temptation to keep adding “advanced RAG” features.\n\n**First, I wanted to know whether the foundations actually worked.**\n\nFirst, I wanted to answer a simpler question:\n\nDoes the retrieval system I actually built work?\n\nAnd more importantly:\n\nCan I prove it?\n\nKnowledge Vault has a deliberate boundary between knowledge retrieval and model inference.\n\nThe retrieval system is responsible for turning documents into useful context. A separate inference service is responsible for executing the language model.\n\nThe high-level architecture looks like this:\n\nThe system attempts to ground generation in retrieved evidence, but Knowledge Vault does not guarantee that the final answer is faithful to that evidence.\n\nThe first stage uses PyMuPDF for page-by-page PDF text extraction.\n\nThis is intentionally a text-layer PDF pipeline. It preserves source metadata such as document identity and page information so that retrieved chunks can be traced back to where they came from.\n\nThat traceability matters because a RAG system shouldn’t just return an answer. It should be possible to inspect the evidence behind that answer.\n\nIt also establishes an important boundary for M1: scanned PDFs requiring OCR are not the problem being solved yet.\n\nExtracted text is passed through recursive, token-aware chunking.\n\nThe goal is not simply to cut the document every N characters.\n\nA retrieval chunk should be large enough to preserve useful context but small enough to remain a precise retrieval unit.\n\nThat creates an unavoidable trade-off.\n\nChunks that are too small can destroy context:\n\n```\nChunk A: \"The system uses...\"Chunk B: \"...a two-stage retrieval process.\"\n```\n\nChunks that are too large can dilute relevance:\n\n```\nQuestion   ↓Large chunk containing five unrelated concepts   ↓Weak retrieval precision\n```\n\nSo chunking is not just preprocessing. It is already a retrieval decision.\n\nEach chunk is converted into a vector using:\n\nBAAI/bge-small-en-v1.5\n\nThe same embedding model is used for the user’s query.\n\nConceptually:\n\n```\nDocument chunk          Question     ↓                      ↓Embedding model        Embedding model     ↓                      ↓   Vector               Query vector\n```\n\nThe purpose is to move the problem from lexical matching toward semantic similarity.\n\nThe chunk vectors are stored in Qdrant along with metadata such as:\n\nThat metadata becomes important later because vector similarity alone doesn’t tell us whether a retrieved result is actually useful evidence.\n\nThe retrieval pipeline can optionally add:\n\nBAAI/bge-reranker-base\n\nThe architecture is intentionally two-stage:\n\nThe vector database is responsible for quickly finding candidates.\n\nThe reranker gets a smaller candidate set and performs a more expensive query-passage relevance calculation.\n\nThat makes the reranker more than a quality component. It is also a latency-bearing architectural component.\n\nThis is a classic retrieval architecture, but I didn’t want to assume that adding the second stage automatically made the system better.\n\nThat had to be measured.\n\nThe other major architectural choice was separating Knowledge Vault from the model-serving layer.\n\nIt would have been simpler to put everything into one application:\n\n```\nPDF → RAG → Ollama → Answer\n```\n\nInstead, the system uses a service boundary:\n\nThe two systems have different responsibilities.\n\nKnowledge Vault is primarily a knowledge-processing and retrieval system.\n\nThe inference service is a model-execution system.\n\nThis means I can change the retrieval pipeline without redesigning model serving, and I can change the model runtime without rewriting the document pipeline.\n\nThere is still coupling. It simply becomes explicit API coupling rather than shared implementation coupling.\n\nThat distinction matters once a system grows beyond a single experiment.\n\nOne of the first things Knowledge Vault made clear is that the LLM is actually quite far downstream in the failure chain.\n\nConsider the pipeline:\n\nEvery stage can introduce a different failure.\n\nThis is why I started thinking about RAG less as a single algorithm and more as a chain of failure boundaries.\n\nUnderstanding those boundaries became one of the central goals of Knowledge Vault.\n\nThe first version of Knowledge Vault could answer questions over a PDF.\n\nThat was encouraging.\n\nBut: *“I asked it a question and got a good answer”* isn’t an evaluation methodology. It’s a demo.\n\nI wanted a controlled experiment.\n\nSo I selected one real document: The 10X Rule, a 214-page book.\n\nRather than immediately indexing many different documents, I deliberately constrained the experiment.\n\nThe experiment was run locally using the same document, question set, inference configuration, and final context size across runs. Retrieval configuration was changed only where required by the experiment. Latency was measured per pipeline stage rather than from a single end-to-end timer.\n\nThe benchmark is intentionally small and controlled; it is intended to compare configurations within this system, not to claim general RAG performance across datasets.\n\nThis reduces the number of variables and makes it easier to understand why something changed.\n\nI created a set of 35 questions designed to exercise different retrieval behaviors.\n\nThe questions included:\n\nThe last category is important.\n\nA RAG system should not assume every question has an answer in its source.\n\nIf the document doesn’t contain the required information, the desired behavior is:\n\n```\nQuestion → Retrieve → No sufficient evidence → No grounded answer\n```\n\nrather than:\n\n```\nQuestion → LLM → Plausible-sounding answer\n```\n\nThe deliberately unanswerable questions were included to test how retrieval and generation behaved when the source did not contain sufficient evidence.\n\nThe most important detail of the experiment is that this was not simply “reranker off versus reranker on.”\n\nThere were three runs.\n\nThe comparison was deliberately controlled:\n\nThe baseline path was:\n\n```\nQuestion → BGE embedding → Qdrant → Top-K candidates → LLM\n```\n\nThis establishes what the vector retrieval layer can do on its own.\n\nThe second run added the cross-encoder:\n\n```\nQuestion → BGE embedding → Qdrant → Candidate chunks → BGE reranker → Final Top-K → LLM\n```\n\nThe initial candidate pool was small.\n\nThe hypothesis was straightforward: if vector similarity gives us reasonable candidates but imperfect ordering, a cross-encoder should be able to improve the final ranking.\n\nThe third run changed another important variable.\n\nInstead of giving the reranker only the initial small candidate set, I changed:\n\nrerank_candidate_k = 20\n\nwhile keeping the final context size at the smaller Top-K.\n\nThe architecture became:\n\nThis distinction matters.\n\nThat means Run 3 tests a real retrieval design question:\n\nDoes giving the reranker a larger candidate pool allow it to recover better evidence?\n\nIf the correct passage is ranked 8th by vector similarity, a reranker that sees only five candidates can never select it. Increasing the candidate pool to 20 gives it that opportunity.\n\nBut more candidates also mean more distractors.\n\nSo the experiment is not simply “5 is bad, 20 is good.” It is a candidate-coverage-versus-ranking trade-off.\n\nI realized I was actually evaluating several different properties of the system:\n\nThese are related, but they are not the same metric.\n\nA system can retrieve a relevant passage and still generate the wrong answer. A reranker can change the ranking without improving correctness. And a system can retrieve a highly similar passage that is not actually useful evidence.\n\nBefore looking at individual failures, I wanted to understand the system at the aggregate level.\n\nThe reranker changed the Top-1 retrieved result on 27 of 35 queries — that’s 77.1%.\n\nOnly 8 queries retained the same Top-1 result.\n\nAt first glance, that sounds like a strong result. A component changed the ranking on more than three quarters of the questions.\n\nIt would be very easy to write: *“The reranker improved retrieval by 77.1%.”*\n\nBut that would be an invalid conclusion.\n\nThe experiment measured ranking changes, not correctness.\n\nA ranking can change from bad → good. It can also change from good → bad.\n\nSo I needed to inspect what actually moved.\n\nChanging rerank_candidate_k from the smaller candidate pool to 20 produced Top-1 changes on only 6 of 35 queries.\n\nThis is a useful result precisely because it isn’t spectacular.\n\nIncreasing the candidate pool gave the reranker more information. But it did not automatically produce better Top-1 retrieval.\n\nOnly one of these changes was clearly an improvement. Several moved the system toward less useful evidence. And one was an unanswerable query, so changing its retrieved page isn’t inherently a retrieval improvement.\n\nThe conclusion is therefore:\n\nA larger reranking candidate pool increases the search space, but it does not guarantee better evidence selection.\n\nMore candidates can create more opportunities. They can also create more distractions.\n\nThe other half of the experiment was performance.\n\nI instrumented the pipeline stage by stage rather than measuring only end-to-end latency.\n\nThe results were revealing.\n\nThe reranker itself added approximately:\n\nThis produced a useful architectural insight.\n\nI initially expected vector retrieval to be one of the expensive parts of the system. It wasn’t.\n\nQdrant retrieval was around 8–9 ms p50, 12–14 ms p95.\n\nThe cross-encoder was around ~2.9 seconds per query.\n\nThe LLM was still the largest stage overall, at roughly 13–14 seconds p50.\n\nBut the reranker was no longer a negligible preprocessing step. It became a meaningful latency decision.\n\nAverages can hide the behavior users actually experience.\n\nSome queries were much more expensive in the reranking stage. For example:\n\nThis changes the engineering question.\n\nIt is no longer simply: *“Does reranking work?”*\n\nIt becomes: *“Under what conditions is several additional seconds of reranking worth the retrieval benefit?”*\n\nThat is a much better engineering question because it forces quality and performance into the same decision.\n\nNow came the most interesting part of the experiment.\n\nConsider a question about the four degrees of action.\n\nThe book’s Chapter 7 contains the actual explanation. The relevant concepts include:\n\nThat is the evidence the LLM actually needs.\n\nBut the book’s Table of Contents also contains:\n\nFrom a semantic matching perspective, the Contents page looks extremely relevant.\n\nFrom an evidence perspective, it is weak.\n\nThe reranker sometimes preferred exactly that kind of page.\n\nFor one query, the baseline retrieved substantive Chapter 7 content around page 45, while reranking moved the Top-1 result to the Contents page.\n\nFor another, the baseline again found Chapter 7 content while the reranker promoted an Index entry around page 196.\n\nThis is not random noise. It exposes a structural limitation.\n\nThis became one of the central lessons from the experiment.\n\nA reranker evaluates the relationship between:\n\n```\nquery ↔ passage\n```\n\nBut a RAG system needs to optimize something closer to:\n\n```\nquery ↔ passage ↔ answerability ↔ evidence usefulness\n```\n\nA Table of Contents can have high semantic relevance because it contains the exact concepts mentioned in the question.\n\nAn Index can have even more keyword-dense matches.\n\nBut neither necessarily contains enough information to answer the question.\n\nSo the system has a subtle failure mode:\n\nThe problem isn’t necessarily that the reranker doesn’t understand the query.\n\nIt is that document structure is not the same thing as semantic relevance.\n\nA Contents page, Index page, Glossary, Copyright page, and chapter body have very different evidentiary value.\n\nThe reranker doesn’t automatically know that.\n\nThe retrieval failures were not the only thing the evaluation exposed.\n\nOne generated answer was also factually wrong when checked against the source.\n\nThe question asked:\n\nWhat are the four major mistakes people make when setting goals?\n\nThe generated answer produced a plausible list involving things such as:\n\nThe answer sounded reasonable.\n\nBut the book’s actual four mistakes are:\n\nThis was a genuine correctness failure.\n\nAnd it exposed another boundary:\n\nA fluent answer is not proof that the RAG pipeline retrieved or used the correct evidence.\n\nThe LLM can produce an answer that sounds consistent with the topic while still contradicting the source.\n\nThat is why final-answer evaluation alone is insufficient.\n\nAt this point I had some compelling measurements:\n\nBut one number was still missing: retrieval accuracy.\n\nI could not honestly say: *“Reranking improved retrieval accuracy by X%,”* because I had not yet established machine-checked passage-level ground truth for every query.\n\nAnd this distinction matters. The experiment proves:\n\nIt does not yet justify a universal claim that reranking improves retrieval quality.\n\nThat restraint is part of the result.\n\nThis experiment changed how I think about RAG evaluation.\n\nThere are at least three separate questions:\n\n1. Did retrieval find the evidence?\n\n```\nQuestion → Candidate set → Is the correct evidence present?\n```\n\nThis measures candidate-generation quality.\n\n2. Did reranking put the evidence in a useful position?\n\n```\nCandidate set → Reranker → Did the correct evidence move upward?\n```\n\nThis measures ranking quality.\n\n3. Did the LLM answer from the evidence?\n\n```\nRetrieved evidence → LLM → Is the answer supported?\n```\n\nThis measures grounded answer quality.\n\nAnd there is a fourth question:\n\n4. How did the system behave when sufficient evidence was absent?\n\nThat matters for the deliberately unanswerable questions.\n\nA useful evaluation therefore looks more like:\n\n```\nQuery  ↓Candidate Retrieval  ↓Reranker Quality  ↓Evidence Quality  ↓Answer Grounding  ↓Abstention Quality\n```\n\nA single “answer accuracy” number hides too much.\n\nAssumption 1: Adding a reranker should improve retrieval What happened: The reranker changed the Top-1 result on 27/35 queries. What I learned: A ranking change is not an improvement metric. It has to be evaluated against ground truth.\n\nAssumption 2: The vector database would be expensive What happened: Qdrant retrieval was around 8–9 ms p50. What I learned: The vector database was not the bottleneck I expected.\n\nAssumption 3: Better semantic relevance means better context What happened: The reranker sometimes promoted Contents and Index pages. What I learned: Semantic relevance and evidentiary usefulness are different objectives.\n\nAssumption 4: A correct-looking final answer means retrieval worked What happened: At least one answer sounded plausible but was wrong when checked against the source. What I learned: Retrieval and generation must be evaluated independently.\n\nAssumption 5: More candidates should give the reranker more chances to succeed What happened: rerank_candidate_k = 20 changed only 6/35 Top-1 results, and most of those changes were not improvements. What I learned: More candidates increase opportunity and search space at the same time.\n\nThe next improvement should not be “use a bigger model.”\n\nIt should be a more precise evaluation and a targeted retrieval change.\n\nFirst, I want passage-level ground truth for the 35 questions:\n\nThen I can calculate meaningful retrieval metrics:\n\nFor reranking, I want to compare multiple candidate-pool sizes and determine where the correct evidence first appears.\n\nSeparately, I want to evaluate whether the final answer is actually supported by the retrieved evidence.\n\nAnd for unanswerable questions, I want to measure whether the system avoids producing unsupported answers.\n\nOnly then can I make a strong claim about whether a particular retrieval configuration is better.\n\nThe observed Contents and Index failures suggest a concrete hypothesis.\n\nBefore the reranker sees candidates, Knowledge Vault could incorporate document structure:\n\n```\nQdrant → Candidate chunks → Structural filtering → Reranker → Final context\n```\n\nPages classified as:\n\ncould be down-ranked or excluded when they are unlikely to provide answer-bearing evidence.\n\nBut I don’t want to call that a solution yet. It is a hypothesis.\n\nThe correct engineering process is:\n\n```\nFailure → Hypothesis → Implement → Re-run same benchmark → Compare → Keep or reject\n```\n\nThat keeps the system honest.\n\nKnowledge Vault started as a RAG implementation exercise.\n\nIt became an evaluation exercise.\n\nAnd the most useful thing I learned wasn’t a particular embedding model, vector database, or reranker. It was that a RAG system has several different meanings of “relevance.”\n\nA passage can be:\n\nThose are not the same property.\n\nThe experiment made that visible.\n\nA Table of Contents can be highly relevant to a question and still be poor evidence.\n\nA larger candidate pool can expose the reranker to more relevant chunks and more distractors.\n\nA reranker can substantially change rankings without necessarily improving correctness.\n\nAnd an LLM can produce a convincing answer even when the retrieval layer has made a poor choice.\n\nThat is why I no longer think of RAG as:\n\n```\nPDF → Vector DB → LLM\n```\n\nI think of it as:\n\n```\nDocument → Extraction → Chunking → Representation → Candidate Retrieval →Ranking → Evidence Selection → Generation → Verification\n```\n\nEvery stage has its own failure modes.\n\nI started with a working PDF RAG system.\n\nThen I instrumented it.\n\nI created a controlled 35-question evaluation.\n\nI compared retrieval configurations.\n\nI measured latency instead of guessing where the bottleneck was.\n\nI increased rerank_candidate_k from the smaller candidate pool to 20 and observed what actually changed.\n\nI inspected the retrieved pages against the original source.\n\nAnd I found cases where a component that looked useful at first glance was actually selecting worse evidence.\n\nI also found an answer that sounded correct but failed source verification.\n\n**Those are not failures** of the project. Those are the results of the experiment.\n\nThe system doesn’t need to be perfect for the milestone to be successful. **It needs to become understandable.**\n\nAnd after this experiment, I understand a lot more about where the retrieval system succeeds, where it fails, what the reranker actually costs, and which assumptions need to be tested next.\n\n**That is a much more useful result** than a benchmark number that simply says: *“RAG works.”*\n\nBecause the real engineering question was never whether I could make a RAG system answer a question.\n\nIt was whether I could **build one, measure it**, challenge it, find where it breaks, and use those failures to make the next decision.\n\n**That is what Knowledge Vault was really about.**\n\n[Building Knowledge Vault : From RAG Implementation to Retrieval Engineering](https://pub.towardsai.net/building-knowledge-vault-rag-retrieval-engineering-b540161cda6f) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.", "url": "https://wpnews.pro/news/building-knowledge-vault-from-rag-implementation-to-retrieval-engineering", "canonical_source": "https://pub.towardsai.net/building-knowledge-vault-rag-retrieval-engineering-b540161cda6f?source=rss----98111c9905da---4", "published_at": "2026-10-01 05:59:43+00:00", "updated_at": "2026-10-01 06:16:57.481132+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-tools", "ai-infrastructure", "mlops"], "entities": ["Jaival Suthar", "Knowledge Vault", "Inference Lab", "PyMuPDF", "BAAI/bge-small-en-v1.5", "BAAI/bge-reranker-base", "Qdrant", "GitHub"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/building-knowledge-vault-from-rag-implementation-to-retrieval-engineering", "markdown": "https://wpnews.pro/news/building-knowledge-vault-from-rag-implementation-to-retrieval-engineering.md", "text": "https://wpnews.pro/news/building-knowledge-vault-from-rag-implementation-to-retrieval-engineering.txt", "jsonld": "https://wpnews.pro/news/building-knowledge-vault-from-rag-implementation-to-retrieval-engineering.jsonld"}}