{"slug": "hybrid-search-rag-over-internal-docs-production-guide", "title": "Hybrid Search RAG Over Internal Docs: Production Guide", "summary": "A developer at LogicLoop described a production hybrid-search RAG pipeline that runs BM25 keyword retrieval in Elasticsearch alongside dense vector retrieval in pgvector and fuses normalized scores, reporting a 22% increase in MRR@10 after switching from vector-only retrieval on the company's engineering wiki. Tested on 500k internal Confluence pages and Jira tickets, the hybrid approach reached 0.50 MRR@10 and 0.79 recall@10 versus 0.41 and 0.68 for vector-only, at roughly 30ms extra p95 latency. The writeup also notes a failure mode in which queries of fewer than three words are routed to keyword-only search because vector results add noise.", "body_md": "Hybrid search RAG over internal docs combines vector and keyword search to improve retrieval accuracy. I’ve used this in production for internal knowledge bases and support ticket systems, and it cuts down on missed relevant chunks by 30-40% compared to pure vector or keyword alone.\n\nHybrid search in RAG means running both dense vector retrieval (like embeddings) and sparse keyword search (like BM25) in parallel, then fusing the results. You get the semantic understanding of vectors plus the exact-match precision of keywords. This matters when your internal docs have jargon, acronyms, or exact phrases that embeddings might miss.\n\nInternal docs are messy. They have code snippets, product names like “WidgetX Pro”, and internal abbreviations. Pure vector search might return semantically similar but irrelevant chunks. Pure keyword search misses synonyms or paraphrased answers. Hybrid fixes this by giving you both. At LogicLoop, we saw a 22% increase in MRR@10 after switching from vector-only to hybrid on our engineering wiki.\n\nHere’s how we wired it up in Python with FastAPI. First, store your chunks in both systems:\n\n``` python\n# Store in Elasticsearch (keyword/BM25)\nfrom elasticsearch import Elasticsearch\nes = Elasticsearch(\"http://localhost:9200\")\nes.index(index=\"docs\", id chunk_id, body={\"text\": chunk_text, \"metadata\": metadata})\n\n# Store in pgvector (dense vectors)\nimport asyncpg\nconn = await asyncpg.connect(dsn=DB_DSN)\nawait conn.execute(\n    \"INSERT INTO doc_chunks (id, embedding, text) VALUES ($1, $2, $3)\",\n    chunk_id, embedding.tolist(), chunk_text\n)\n```\n\nAt query time, run both searches and fuse scores:\n\n``` python\ndef hybrid_search(query_text, query_vector, k=10):\n    # Keyword search via ES\n    es_results = es.search(\n        index=\"docs\",\n        body={\n            \"query\": {\"match\": {\"text\": query_text}},\n            \"size\": k * 2  # get more to fuse later\n        }\n    )\n    keyword_hits = [(hit[\"_id\"], hit[\"_score\"]) for hit in es_results[\"hits\"][\"hits\"]]\n\n    # Vector search via pgvector\n    vector_results = await conn.fetch(\n        \"SELECT id, text, embedding <=> $1 AS distance FROM doc_chunks ORDER BY distance LIMIT $2\",\n        query_vector, k * 2\n    )\n    vector_hits = [(row[\"id\"], 1 - row[\"distance\"]) for row in vector_results]  # convert to similarity\n\n    # Simple score fusion: normalize and add\n    all_scores = {}\n    max_es = max([s for _, s in keyword_hits]) if keyword_hits else 1\n    max_vec = max([s for _, s in vector_hits]) if vector_hits else 1\n\n    for doc_id, score in keyword_hits:\n        all_scores[doc_id] = all_scores.get(doc_id, 0) + score / max_es\n    for doc_id, score in vector_hits:\n        all_scores[doc_id] = all_scores.get(doc_id, 0) + score / max_vec\n\n    # Return top k fused\n    sorted_ids = sorted(all_scores.items(), key=lambda x: x[1], reverse=True)[:k]\n    return [doc_id for doc_id, _ in sorted_ids]\n```\n\nWe normalize scores to [0,1] per system before adding. This prevents one system from dominating due to scale differences.\n\nWe tested on 500k internal Confluence pages and Jira tickets. Metrics: MRR@10 and [recall@10](mailto:recall@10). Baseline: pure vector (pgvector only).  \n\n| Method | MRR@10 | Recall@10 | Latency (p95) | \n|---|---|---|---|\n| Vector only | 0.41 | 0.68 | 120ms | \n| Keyword only | 0.33 | 0.52 | 80ms | \n| Hybrid | 0.50 | 0.79 | 150ms | \n\nHybrid added ~30ms latency but gained significant quality. The cost? Running two indexes doubles storage and write overhead. We mitigated this by using Elasticsearch’s snapshot lifecycle and pgvector’s partitioning by doc source.\n\nFailure mode we hit: when query text is very short (1-2 words), keyword search dominates and vector adds noise. We now route sub-3-word queries to keyword-only via a simple length check.\n\n[See our full RAG pipeline diagram](https://www.logiclooptech.dev/rag-pipeline-diagram-design-build-and-scale/) for context, but here’s the hybrid-specific flow:  \n\nWe keep the fusion logic in a dedicated service so we can swap fusion algorithms (like RRF or learning-to-rank) without touching the LLM layer.\n\n`knn` and `match` combo in one query if on 8.0+.\n`BAAI/bge-small-en-v1.5`).\nWe tried using Weaviate’s hybrid search but switched to ES+pgvector because we needed fine-grained control over indexing policies and couldn’t justify the ops overhead of a new system.\n\n**Does hybrid search always beat vector-only?**\n\nNo. On clean, well-written documentation with minimal acronyms, vector-only can match hybrid. Test on your data.  \n\n**How do I tune the weight between vector and keyword?**\n\nStart with equal weight (normalize then add). If you have relevance labels, use a validation set to sweep weights from 0.0 to 1.0 for vector (1-weight for keyword).  \n\n**Can I use hybrid search with LLMs that have long context?**\n\nYes. Hybrid improves the quality of chunks fed into the LLM, which matters more than raw context length when dealing with noisy internal data.  \n\n**What embedding model works best for hybrid?**\n\nWe use BAAI/bge-small-en-v1.5. It’s small, fast, and works well with keyword fusion. Larger models help marginally but add latency.  \n\n[Hybrid Search RAG Implementation: A Practical Guide](https://www.logiclooptech.dev/hybrid-search-rag-implementation-a-practical-guide/)\n\n[RAG Chunking Best Practices for Production Systems](https://www.logiclooptech.dev/rag-chunking-best-practices-for-production-systems/)\n\n[RAG Chunking Evaluation: Metrics, Trade-offs, and Production Lessons](https://www.logiclooptech.dev/rag-chunking-evaluation-metrics-trade-offs-and-production-lessons/)", "url": "https://wpnews.pro/news/hybrid-search-rag-over-internal-docs-production-guide", "canonical_source": "https://dev.to/ayush_kumar_085a0f2c54e3f/hybrid-search-rag-over-internal-docs-production-guide-2kg4", "published_at": "2026-09-30 10:14:06+00:00", "updated_at": "2026-09-30 10:17:21.867379+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-infrastructure", "mlops", "developer-tools"], "entities": ["LogicLoop", "Elasticsearch", "pgvector", "FastAPI", "Confluence", "Jira", "BM25"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/hybrid-search-rag-over-internal-docs-production-guide", "markdown": "https://wpnews.pro/news/hybrid-search-rag-over-internal-docs-production-guide.md", "text": "https://wpnews.pro/news/hybrid-search-rag-over-internal-docs-production-guide.txt", "jsonld": "https://wpnews.pro/news/hybrid-search-rag-over-internal-docs-production-guide.jsonld"}}