Enterprise RAG Engineering — Part 2 of 3 Part 1 argued that most enterprise RAG systems fail at the ingestion layer, before a single query is ever issued — and that the fix is layout-aware parsing rather than a better model. We built that stage. Docling with HeadingHierarchyOptions enabled to recove
Enterprise RAG Engineering — Part 2 of 3 Part 1 argued that most enterprise RAG systems fail at the ingestion layer, before a single query is ever issued — and that the fix is layout-aware parsing rather than a better model. We built that stage. Docling with HeadingHierarchyOptions enabled to recover section structure, generate_parsed_pages on so the font-style signal survives, HybridChunker to split at structural boundaries rather than fixed token counts, and metadata carried through to every chunk: page number, filename, and the heading breadcrumb the chunk sits under. The result is an index where structure survived. Tables did not get flattened into a single text stream. Headings stayed attached to the content they govern. Chunks mean something on their own. That was the write path. This article is about the read path — what happens when the first query arrives. Because a corpus can be indexed perfectly and still fail to return the right chunk. The default retrieval pattern in almost every RAG tutorial is three lines long. Embed the question, run a nearest-neighbour search, return the top k. It works well enough in a proof of concept that most teams ship it. Then a user searches for WH-1000XM5. The system returns wireless headphones. All of them. Not that one. This is not a tuning problem, and no amount of prompt engineering fixes it. It is a property of how embeddings work. An embedding model compresses a span of text into a fixed-width vector that encodes what the text is about. A model number is not about anything. It is an opaque string with no distributional meaning, and whatever position it occupies in vector space is determined almost entirely by the words around it. Two chunks that both discuss wireless headphones sit close together regardless of which specific model each one names. The same failure appears with any exact identifier: a part number, an error code, a control reference, a contract clause, an internal project name. Enterprise corpora are full of them, and users search by them constantly. Lexical search has the exact opposite profile. It matches the token and ranks by term statistics, so it finds WH-1000XM5 on the first hop. And it fails completely when a user asks "how long does the battery last" against a page that only ever says "playback time" — because there is no shared token to match. Neither retriever is better. They fail on different things. That is the entire argument for running both. Running two retrievers is easy. Merging their results is where hybrid search is actually decided, and the obvious approach fails in a way worth understanding precisely. Cosine distance is bounded. For normalised vectors it falls in a known range, and a given value means roughly the same thing across queries and across corpora. Lexical ranking functions are not bounded. Postgres ts_rank_cd and BM25 both produce unbounded positive numbers whose magnitude depends on term frequency, match proximity, document length, and the statistics of the corpus being searched. Index another ten thousand documents and the same document scores differently against the same query. So one retriever returns 0.83 and the other returns 24.7, and there is no constant that reconciles them — because the relationship between the two scales is not fixed. The usual workaround is min-max normalisation: scale each list to the range 0 to 1, then blend with a weight. This introduces a subtler failure. The normalisation is computed over the candidates that were returned, not over the corpus. If the lexical arm returns 50 documents that are all weak matches, the best of those 50 normalises to 1.0 — the same score a perfect match would receive. The normalisation destroys exactly the signal you were trying to preserve, and it does so silently, in precisely the case where the retriever had nothing good to offer. Reciprocal Rank Fusion, introduced by Cormack, Clarke and Büttcher at SIGIR 2009, takes the position that the scores are unsalvageable and discards them entirely. Each document is scored on its position in each list: RRF(d) = Σ 1 / (k + rank_i(d)) i rank_i(d) is the document's 1-indexed position in list i, and k is a smoothing constant. The original paper fixed k = 60 during a pilot investigation and found it near-optimal across TREC collections without being especially sensitive to the exact value. It is now the default in Elasticsearch and OpenSearch. Because it consumes only rank positions, RRF can fuse any number of ranked lists from any retrievers, whatever they score and however they score it. Rank Contribution Share of rank-1 1 1 / 61 = 0.01639 100% 2 1 / 62 = 0.01613 98.4% 20 1 / 80 = 0.01250 76.3% 100 1 / 160 = 0.00625 38.1% The curve is deliberately flat at the top. Moving from rank 1 to rank 2 costs 1.6% of the score. Moving from rank 1 to rank 20 costs about 24%. The consequence is that agreement outranks confidence. A chunk ranked second by both retrievers scores roughly 0.0322. A chunk ranked first by one retriever and missing from the other scores 0.0164. The chunk both retrievers found wins by a factor of two, despite neither ranking it first. With k = 0 the formula collapses to pure reciprocal rank, where rank 1 scores 1.0 and rank 2 scores 0.5, and a single retriever's top hit dominates everything. The 60 damps that until consensus between independent signals becomes the deciding factor. Both retrieval arms, the fusion, and the payload fetch run in one database round trip. The alternative is two queries, both result sets pulled into the application, and fusion in Python. That costs two round trips plus a merge loop, and it moves ranking logic away from the layer that holds the data. The database can express RRF directly. ROW_NUMBER() produces the ranks. FULL OUTER JOIN produces the union. The arithmetic is two divisions and an addition per row. The rows are already there — making the database do the merge is strictly cheaper than shipping them out to merge elsewhere. # Builds an OR query from the question's words, entirely in the database. # Same stemming and stopword handling as the stored index itself. ANY_TSQUERY = r"""CAST(array_to_string(ARRAY( SELECT '''' || replace(replace(lexeme, E'\', E'\\'), '''', '''''') || '''' FROM unnest(tsvector_to_array(to_tsvector('english', :question))) AS lexeme ), ' | ') AS tsquery)""" HYBRID_SQL = f""" WITH keyword AS ( SELECT id, ROW_NUMBER() OVER (ORDER BY ts_rank_cd(fts, {ANY_TSQUERY}) DESC) AS rank FROM chunks WHERE fts @@ {ANY_TSQUERY} ORDER BY ts_rank_cd(fts, {ANY_TSQUERY}) DESC LIMIT :pool ), semantic AS ( SELECT id, ROW_NUMBER() OVER (ORDER BY embedding <=> CAST(:qv AS vector)) AS rank FROM chunks ORDER BY embedding <=> CAST(:qv AS vector) LIMIT :pool ) SELECT c.id, c.content, c.page, c.heading, COALESCE(1.0/(60+k.rank), 0) + COALESCE(1.0/(60+s.rank), 0) AS score, k.rank AS keyword_rank, s.rank AS vector_rank FROM keyword k FULL OUTER JOIN semantic s USING (id) JOIN chunks c USING (id) ORDER BY score DESC LIMIT :pool; """ Building the lexical query in the database, not in Python. The obvious approach is to split the question in Python and join the terms with an OR operator. That introduces a second tokenizer — and Python does not apply the same stemming or stopword list that produced the stored index. The result is a query whose terms do not match what was indexed. ANY_TSQUERY avoids this by running the question through to_tsvector('english', ...), the same function with the same configuration that produced the stored vectors, then unnesting the lexemes, quoting each, and joining them. Query-side and index-side analysis are identical because they are the same call. The replace calls escape backslashes and quotes so a lexeme containing either cannot break out of the quoted literal. OR, not AND. Terms are joined with the OR operator, making this recall-oriented: a chunk matching any term is a candidate. In a fusion pipeline that is correct. The lexical arm's job is to surface anything plausible; precision is handled downstream. FULL OUTER JOIN, not INNER JOIN. This is the structural decision that makes the fusion behave correctly. An INNER JOIN keeps only chunks that both retrievers found — which discards exactly the cases hybrid search exists to catch. The SKU that only the lexical arm found. The paraphrase that only the vector arm found. A LEFT JOIN privileges one arm arbitrarily. With FULL OUTER JOIN, a chunk present in one list and absent from the other survives with a NULL rank on the missing side, and COALESCE(..., 0) contributes zero for that side rather than propagating NULL through the arithmetic. A single-arm hit scores about 0.0164. A two-arm hit scores up to 0.0328. Both are retrievable; the one both retrievers agreed on ranks higher. Deferring the payload fetch. The CTEs carry only ids and ranks. Both arms rank over the full table, and carrying content through them would mean materialising large text values for rows about to be discarded. Joining for the payload after the fusion keeps the intermediate result sets narrow. Two different depths. pool controls how deep each arm searches and how many fused candidates survive. The request's k controls how many documents come back at the end. They are not the same parameter, and binding the request's k to the SQL's final LIMIT would cap the candidate pool at the number of results the user asked for — starving the reranker of exactly the candidates it exists to re-order. The gains are capability-level, and they are specific. Identifier queries work. A search for a model number, a part code or a clause reference now returns the chunk containing that string, because the lexical arm finds it on exact match. Before, those queries returned topically adjacent chunks and nothing else. This is the single largest practical difference for an enterprise corpus. Paraphrase queries still work. The vector arm is unchanged, so questions phrased conceptually against documents using different vocabulary behave exactly as they did. Hybrid search is additive — nothing that worked before stopped working. Single-arm hits are not lost. Because of the FULL OUTER JOIN, a chunk only one retriever found still reaches the candidate set. It ranks below chunks both arms agreed on, which is the correct priority, but it is retrievable. An INNER JOIN implementation would silently drop exactly the results that justify running two retrievers. The candidate pool is wider without being noisier. Two independent retrievers surface a more diverse set of candidates than either alone, and RRF orders them by cross-retriever agreement rather than by whichever scoring function happened to produce a larger number. No scale calibration to maintain. Because RRF consumes only rank positions, there is no normalisation constant that drifts as the corpus grows, and no weight to re-tune after an ingestion run. The fusion behaves the same on 1,000 chunks as on 100,000. One round trip, not two. The entire fusion happens inside the database. There is no second query, no application-side merge loop, and no window where the two result sets could come from different states of the data. A note on evidence: these are mechanical gains, demonstrable by construction. Quantified relevance improvement requires running a labelled evaluation set against both configurations — which is the right next step, and which the candidate-depth and threshold parameters should be tuned against rather than by intuition. Hybrid search produces a better candidate set. It does not produce a better answer, and the reason is worth stating plainly. The maximum possible RRF score in a two-retriever system is 2/61 — about 0.0328 — achieved by a chunk ranked first by both arms. That ceiling is fixed by the formula, not by match quality. A chunk that answers the question perfectly and a chunk that is merely well-positioned in both lists receive identical scores if they occupy identical ranks. The number tells you where a chunk sat in two ordered lists. It tells you nothing about whether the chunk answers anything. This has a direct operational consequence: any threshold applied to a fused score is a threshold on rank position, which is not the same thing as a threshold on relevance and will behave unpredictably as the corpus changes. Both retrieval arms score the query and the document separately. The vector arm embeds each independently and compares the results. The lexical arm computes term statistics without ever considering the query and document jointly. That independence is what makes them fast and indexable — and it is also their ceiling. A cross-encoder reranker processes the query and the document together in a single forward pass. It can model the interaction between them: whether this specific document actually answers this specific question, rather than whether they occupy nearby regions of a vector space or share tokens. That is far more expensive per document, which is exactly why it runs last, over a pool of around 50 candidates rather than the whole corpus. The architecture is deliberate: cheap indexed retrieval narrows millions of chunks to tens, then expensive joint scoring orders those tens. In this system that stage is Amazon Bedrock's amazon.rerank-v1:0, called through the bedrock-agent-runtime endpoint. The API accepts up to 1,000 inline sources per request, so a 50-candidate pool needs no batching. A real quality signal. The reranker's relevanceScore is the only number in the entire pipeline that measures whether a chunk answers the question. Everything upstream measures position or proximity. A threshold that means something. Because that score reflects relevance rather than rank, a minimum threshold becomes a meaningful gate. Chunks that survived retrieval on position alone but do not actually answer the question get dropped before they reach the prompt. Variable result counts, honestly. A request asking for three results may return fewer — or none. In one evaluation run against a compliance corpus, a question with k=3 returned a single source scoring 0.2067, and the generated answer matched the expected ground truth. The system was not failing. It was reporting that one chunk in the corpus answered that question. That behaviour is only possible because a relevance score exists to threshold against. Without it, three chunks come back regardless, and the two weak ones go into the prompt alongside the good one. Cleaner context for generation. Fewer, better chunks means a shorter prompt, fewer tokens, and less opportunity for the model to anchor on a weakly relevant passage. The quality gate pays for itself at the generation stage as well as the retrieval stage. A usable degraded path. Because RRF ordering already exists before the reranker runs, a rerank failure has a sensible fallback: return the fused results truncated to k. This matters in practice — rerank endpoints throttle, and during testing a real ThrottlingException was absorbed this way without failing the request. try: reranked_results = await rerank_documents(k, question, documents) except Exception: logger.exception("rerank failed, falling back to RRF order") return list(islice(documents, k)) Two things about this fallback are deliberate. It catches bare Exception, which is correct here because any rerank failure should degrade rather than fail, and enumerating every exception a client library might raise is a losing game. And it skips the relevance threshold, because fallback documents carry no score to filter on — the caller gets up to k RRF-ranked chunks that may include weak matches. That is the trade: a complete answer on a weaker ranking, rather than no answer. The absent score should be visible to the caller rather than silently defaulted. Exact-identifier queries fail silently. This is the expensive one. A user searching for a model number, an error code or a clause reference gets plausible-looking results that do not contain what they asked for. There is no error and no empty result — just confidently wrong context handed to the model, which then produces a fluent answer about the wrong thing. The failure is invisible until someone checks the citation. Rare and domain-specific terms degrade. Jargon that is thin in the embedding model's training data has no reliable position in vector space. It embeds near whatever it superficially resembles. Lexical matching does not care how common a term is — it either appears or it does not. Recall depends entirely on one model's view of meaning. If the embedding model's notion of similarity does not fit your domain, there is no second signal to compensate. Every retrieval failure traces to the same single point. Negation and precise phrasing blur. Embeddings capture topic well and precise logical structure poorly. Two passages stating opposite conclusions about the same subject can sit close together in vector space. You have no relevance signal at all. This is the fundamental loss. Retrieval scores — fused or raw — measure position and proximity. Nothing in the pipeline measures whether a chunk answers the question. Any threshold you set is a threshold on something other than what you care about. Every query returns exactly k results, whether or not they are any good. A question your corpus cannot answer returns k chunks anyway, because the retrievers always return their top k by position. Those chunks go into the prompt, and the model does what models do with weakly relevant context — it finds something to say. You cannot distinguish "nothing matched" from "something matched." The system loses the ability to say I don't have that, which in a compliance or support context is often the most valuable answer available. The top of the list is nearly undifferentiated. The flat RRF curve that makes consensus robust also means ranks 1 through 5 are separated by a few percent. Served directly, that ordering is weak — the fusion was designed to select a candidate set, not to produce a final ranking. Prompt context gets noisier. More marginal chunks reach the model, consuming tokens and competing for attention with the one that actually answers the question. Hybrid search without reranking gives you a wider net and no way to judge what you caught. Reranking without hybrid search gives you good judgement applied to a candidate set that may never have contained the right answer. The two solve different halves of the same problem. Retrieval decides what is available to answer with. Reranking decides what is worth answering with. Everything described here runs against a database on a laptop. Part 3 takes it to AWS: S3 as the document store, Amazon OpenSearch Service as the vector and lexical index, and an event-driven ingestion path that starts when a file lands in a bucket rather than when someone runs a script. That move changes where this article's fusion logic lives. The FULL OUTER JOIN and the hand-written 1/(60+rank) arithmetic become a search pipeline declaration — OpenSearch implements RRF as a score-ranker-processor with a configurable rank_constant that defaults to the same 60. Part 3 covers what that buys, what it costs, and what the migration does to retrieval quality measured against the same evaluation set. Cormack, Clarke and Büttcher, Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods, SIGIR 2009 Amazon Bedrock Rerank API reference PostgreSQL text search controls
Key Takeaways #
- •Enterprise RAG Engineering — Part 2 of 3 Part 1 argued that most enterprise RAG systems fail at the ingestion layer, before a single query is ever issued — and that the fix is layout-aware parsing rather than a better model. We built that stage
- •This story was reported by Dev.to , covering developments in thedev space.
- •AI advancements continue to reshape industries — read the full article on Dev.to for complete coverage.
📖 Continue reading the full article: