Hybrid Search RAG Over Internal Docs: Production Guide A developer at LogicLoop described a production hybrid-search RAG pipeline that runs BM25 keyword retrieval in Elasticsearch alongside dense vector retrieval in pgvector and fuses normalized scores, reporting a 22% increase in MRR@10 after switching from vector-only retrieval on the company's engineering wiki. Tested on 500k internal Confluence pages and Jira tickets, the hybrid approach reached 0.50 MRR@10 and 0.79 recall@10 versus 0.41 and 0.68 for vector-only, at roughly 30ms extra p95 latency. The writeup also notes a failure mode in which queries of fewer than three words are routed to keyword-only search because vector results add noise. Hybrid search RAG over internal docs combines vector and keyword search to improve retrieval accuracy. I’ve used this in production for internal knowledge bases and support ticket systems, and it cuts down on missed relevant chunks by 30-40% compared to pure vector or keyword alone. Hybrid search in RAG means running both dense vector retrieval like embeddings and sparse keyword search like BM25 in parallel, then fusing the results. You get the semantic understanding of vectors plus the exact-match precision of keywords. This matters when your internal docs have jargon, acronyms, or exact phrases that embeddings might miss. Internal docs are messy. They have code snippets, product names like “WidgetX Pro”, and internal abbreviations. Pure vector search might return semantically similar but irrelevant chunks. Pure keyword search misses synonyms or paraphrased answers. Hybrid fixes this by giving you both. At LogicLoop, we saw a 22% increase in MRR@10 after switching from vector-only to hybrid on our engineering wiki. Here’s how we wired it up in Python with FastAPI. First, store your chunks in both systems: python Store in Elasticsearch keyword/BM25 from elasticsearch import Elasticsearch es = Elasticsearch "http://localhost:9200" es.index index="docs", id chunk id, body={"text": chunk text, "metadata": metadata} Store in pgvector dense vectors import asyncpg conn = await asyncpg.connect dsn=DB DSN await conn.execute "INSERT INTO doc chunks id, embedding, text VALUES $1, $2, $3 ", chunk id, embedding.tolist , chunk text At query time, run both searches and fuse scores: python def hybrid search query text, query vector, k=10 : Keyword search via ES es results = es.search index="docs", body={ "query": {"match": {"text": query text}}, "size": k 2 get more to fuse later } keyword hits = hit " id" , hit " score" for hit in es results "hits" "hits" Vector search via pgvector vector results = await conn.fetch "SELECT id, text, embedding <= $1 AS distance FROM doc chunks ORDER BY distance LIMIT $2", query vector, k 2 vector hits = row "id" , 1 - row "distance" for row in vector results convert to similarity Simple score fusion: normalize and add all scores = {} max es = max s for , s in keyword hits if keyword hits else 1 max vec = max s for , s in vector hits if vector hits else 1 for doc id, score in keyword hits: all scores doc id = all scores.get doc id, 0 + score / max es for doc id, score in vector hits: all scores doc id = all scores.get doc id, 0 + score / max vec Return top k fused sorted ids = sorted all scores.items , key=lambda x: x 1 , reverse=True :k return doc id for doc id, in sorted ids We normalize scores to 0,1 per system before adding. This prevents one system from dominating due to scale differences. We tested on 500k internal Confluence pages and Jira tickets. Metrics: MRR@10 and recall@10 mailto:recall@10 . Baseline: pure vector pgvector only . | Method | MRR@10 | Recall@10 | Latency p95 | |---|---|---|---| | Vector only | 0.41 | 0.68 | 120ms | | Keyword only | 0.33 | 0.52 | 80ms | | Hybrid | 0.50 | 0.79 | 150ms | Hybrid added ~30ms latency but gained significant quality. The cost? Running two indexes doubles storage and write overhead. We mitigated this by using Elasticsearch’s snapshot lifecycle and pgvector’s partitioning by doc source. Failure mode we hit: when query text is very short 1-2 words , keyword search dominates and vector adds noise. We now route sub-3-word queries to keyword-only via a simple length check. See our full RAG pipeline diagram https://www.logiclooptech.dev/rag-pipeline-diagram-design-build-and-scale/ for context, but here’s the hybrid-specific flow: We keep the fusion logic in a dedicated service so we can swap fusion algorithms like RRF or learning-to-rank without touching the LLM layer. knn and match combo in one query if on 8.0+. BAAI/bge-small-en-v1.5 . We tried using Weaviate’s hybrid search but switched to ES+pgvector because we needed fine-grained control over indexing policies and couldn’t justify the ops overhead of a new system. Does hybrid search always beat vector-only? No. On clean, well-written documentation with minimal acronyms, vector-only can match hybrid. Test on your data. How do I tune the weight between vector and keyword? Start with equal weight normalize then add . If you have relevance labels, use a validation set to sweep weights from 0.0 to 1.0 for vector 1-weight for keyword . Can I use hybrid search with LLMs that have long context? Yes. Hybrid improves the quality of chunks fed into the LLM, which matters more than raw context length when dealing with noisy internal data. What embedding model works best for hybrid? We use BAAI/bge-small-en-v1.5. It’s small, fast, and works well with keyword fusion. Larger models help marginally but add latency. Hybrid Search RAG Implementation: A Practical Guide https://www.logiclooptech.dev/hybrid-search-rag-implementation-a-practical-guide/ RAG Chunking Best Practices for Production Systems https://www.logiclooptech.dev/rag-chunking-best-practices-for-production-systems/ RAG Chunking Evaluation: Metrics, Trade-offs, and Production Lessons https://www.logiclooptech.dev/rag-chunking-evaluation-metrics-trade-offs-and-production-lessons/