Building Better RAG A developer outlined a four-part engineering approach to making retrieval-augmented generation (RAG) systems production-ready, arguing that most RAG failures stem from retrieval rather than generation. The writeup covers pre- and post-retrieval techniques such as query rewriting, HyDE, reranking and relevance grading, chunking strategies, embedding and vector-store tuning, and evaluation as the loop that decides which lever to fix next. The developer advises adding each step only in response to measured failures, since every one adds latency and cost. Most RAG systems that disappoint in production fail at retrieval, not generation: the model answers well from the wrong context, or from none at all. Getting a demo to work takes an afternoon; getting it to work on real queries, real documents, and real traffic takes deliberate engineering at every stage of the pipeline. This post walks through four levers that decide whether a RAG system holds up: what you do before and after retrieval, how you chunk your documents, how you tune your embeddings and vector store, and how you measure whether any of it is working. Treat them as a loop. Evaluation tells you which of the other three to fix next. 1. Pre-retrieval and checking strategies User queries are short, ambiguous, and phrased differently from the documents that answer them. Pre-retrieval work closes that gap before you ever touch the index, and checking work catches bad context before it reaches the model. Before retrieval - Query rewriting. Use an LLM to turn a vague or conversational question into a standalone, search-friendly one. In chat settings this also means resolving references like "what about the second one?" using conversation history. - Query expansion and multi-query. Generate several paraphrases of the question, retrieve for each, and merge the results. This improves recall when users and documents use different vocabulary. - HyDE hypothetical document embeddings . Ask the model to write a plausible answer first, then embed that answer instead of the question. Answers tend to sit closer to documents in embedding space than questions do. - Query decomposition. Split a multi-part question "compare plan A and plan B on pricing and limits" into sub-queries, retrieve for each, then synthesize. - Routing. Classify the query and send it to the right place: a specific index, a SQL database, a web search, or no retrieval at all for questions the model can answer directly. Skipping retrieval when it isn't needed saves latency and avoids distracting the model. - Metadata filtering. Extract constraints from the query date range, product, language, document type and apply them as filters so the search space is small and relevant. After retrieval - Reranking. Retrieve a generous candidate set say 30 to 50 , then rescore it with a cross-encoder or reranking model that reads the query and each passage together. This is often the single biggest quality gain per unit of effort. - Relevance checks. Have a lightweight grader decide whether the retrieved passages actually relate to the question. If they don't, rewrite the query and retry, or fall back to another source. This is the idea behind corrective approaches such as CRAG. - Context compression. Trim retrieved passages down to the sentences that matter, so the model isn't diluted by filler and you spend fewer tokens. - Grounding and answer checks. After generation, verify that each claim is supported by the retrieved text, and require citations. If the support isn't there, return "I don't know" or ask a clarifying question instead of guessing. Every one of these steps adds latency and cost, so add them in response to measured failures rather than by default. 2. Managing chunking strategies Chunking decides what a "unit of knowledge" looks like to your retriever. Chunks that are too large bury the relevant sentence in noise and dilute the embedding. Chunks that are too small lose the context that makes them understandable. There is no universal best size, so the goal is to match chunks to how your content is actually structured and queried. | Strategy | How it works | Best for | Watch out for | | Fixed-size | Split every N tokens, usually with overlap | Quick baselines, uniform text | Cuts mid-sentence and mid-idea | | Recursive | Split on paragraphs, then sentences, then words, until chunks fit | General prose; a solid default | Ignores document structure beyond separators | | Structure-aware | Split on headings, sections, code blocks, table boundaries | Docs, wikis, markdown, HTML, legal text | Needs a parser per format | | Semantic | Start a new chunk where embedding similarity between adjacent sentences drops | Long, topic-shifting text | Costs extra embedding calls; uneven chunk sizes | | Parent-child | Embed small chunks, return the larger parent section | Precise matching with full context | More complex storage and indexing | | Sentence-window | Embed single sentences, return the neighbors around each hit | Dense factual text | Window size needs tuning | Practical guidance - Start simple. Recursive splitting at roughly 300 to 800 tokens with 10 to 15 percent overlap is a reasonable baseline. Tune from there using your evaluation set, not intuition. - Respect structure. Never split a table, a code block, or a numbered procedure in half. Keep headings attached to the text beneath them. - Add context to chunks. Prepend the document title and section path to each chunk before embedding, or generate a short LLM summary of where the chunk sits in the document often called contextual retrieval . A chunk that says "it increased 12%" is useless without knowing what "it" is. - Store rich metadata. Source, section, page, date, author, and access permissions let you filter at query time and cite precisely. - Handle non-text content deliberately. Tables, images, and scanned PDFs need their own extraction path. Convert tables to a text format that preserves rows and columns, and caption or OCR images rather than dropping them. - Version your chunking. When you change chunk size or logic, you must re-embed. Keep the chunking config alongside the index so you can compare versions fairly. 3. Optimizing the vector store and embeddings Even with perfect chunks, retrieval quality is capped by how well your embeddings capture meaning in your domain and how well your index finds nearest neighbors at your scale. Embedding choices - Pick a model for your data, not the leaderboard. Public benchmarks are a starting point. Test two or three candidates on your own queries and documents, since domain vocabulary legal, medical, code, multilingual changes the ranking. - Weigh dimensions against cost. Higher-dimensional vectors usually capture more nuance but increase storage, memory, and search latency. Models that support truncated Matryoshka-style embeddings let you shrink vectors with a modest quality loss. - Use the right input format. Some models expect different prefixes or instructions for queries versus documents. Skipping this quietly hurts quality. - Consider fine-tuning. If you have labeled pairs of queries and relevant passages, fine-tuning an embedding model on them can lift retrieval noticeably for specialized domains. Synthetic query generation can bootstrap training data when you have none. - Normalize consistently. Match your distance metric cosine, dot product, or Euclidean to how the model was trained, and use it everywhere. Vector store and index tuning - Choose an index for your scale. Exact flat search is fine for small collections. For larger ones, approximate nearest neighbor indexes such as HNSW or IVF trade a little recall for large speedups. Tune parameters like HNSW's ef search and M against a recall target measured on your own data. - Quantize when memory is the bottleneck. Scalar or product quantization shrinks the index substantially, usually with small recall loss. Measure before and after. - Go hybrid. Combine dense vector search with keyword search BM25 . Dense retrieval handles paraphrase and meaning; keyword search handles exact terms, product codes, names, and rare jargon. Fuse the results with reciprocal rank fusion or a weighted score, then rerank. - Filter efficiently. Index the metadata fields you filter on, and understand whether your store filters before or after the vector search, because post-filtering can return too few results. - Plan for freshness. Decide how updates and deletes propagate, how you detect changed documents, and how you re-embed when you swap models. Keep an alias or versioned index so you can roll out a new embedding model without downtime. - Watch the operational side. Track query latency percentiles, index build time, memory footprint, and cost per query alongside quality, since a more accurate setup that is too slow or expensive won't ship. 4. Evaluating RAG performance You can't improve what you can't measure, and with RAG you need to measure two things separately: did retrieval find the right material, and did generation use it well? Lumping them together makes failures impossible to diagnose. | Stage | Question | Common metrics | | Retrieval | Did we fetch the passages that contain the answer? | Recall@k, precision@k, MRR, nDCG, hit rate | | Generation | Is the answer correct and based on the context? | Faithfulness groundedness , answer relevance, answer correctness | | Context use | Did we retrieve only what was needed? | Context precision, context recall | | System | Is it usable in production? | Latency p50, p95 , cost per query, refusal rate | Build a golden dataset Start with 50 to 200 realistic questions, each paired with the source passages that answer it and, ideally, a reference answer. Pull questions from real user logs where you can. Cover easy lookups, multi-hop questions, ambiguous queries, and questions that have no answer in your corpus, since a good system should say so. Synthetic question generation from your documents is a fast way to scale the set, but have a human review a sample. Score generation with care - LLM-as-judge is the practical way to grade faithfulness and relevance at scale. Give the judge a clear rubric, the question, the context, and the answer, and ask for a score with a short rationale. - Calibrate the judge. Compare its scores against human labels on a sample, watch for bias toward longer or more confident answers, and re-check whenever you change the judge model. - Frameworks help. Tools such as RAGAS, TruLens, DeepEval, and LangSmith implement these metrics so you don't build them from scratch, but understand what each metric actually measures before trusting a number. Run it as a loop 1. Establish a baseline on your golden set before changing anything. 2. Change one variable at a time chunk size, embedding model, reranker and re-run. 3. Break results down by query type, since an average can hide a category that is failing badly. 4. Read the failures by hand. Classify each as a retrieval miss, a ranking problem, a context problem, or a generation problem, then fix the stage that is actually responsible. 5. Gate releases on the eval set in CI so regressions get caught before users see them. 6. Monitor in production: log queries, retrieved chunks, answers, and user feedback thumbs, corrections, follow-up rephrasing , and feed the hard cases back into your golden set. Putting it together A sensible order of operations: build the evaluation set first, get a simple baseline with recursive chunking and a good embedding model, and add hybrid search and reranking next since they tend to pay off early. Then tune chunking and add query-side techniques only where your failure analysis points to them. Every change should earn its place with a measured improvement. RAG quality comes from the whole pipeline working together, and the teams that do it well are the ones who can tell, query by query, which stage let the user down.