Hybrid search RAG over internal docs combines vector and keyword search to improve retrieval accuracy. I’ve used this in production for internal knowledge bases and support ticket systems, and it cuts down on missed relevant chunks by 30-40% compared to pure vector or keyword alone.
Hybrid search in RAG means running both dense vector retrieval (like embeddings) and sparse keyword search (like BM25) in parallel, then fusing the results. You get the semantic understanding of vectors plus the exact-match precision of keywords. This matters when your internal docs have jargon, acronyms, or exact phrases that embeddings might miss.
Internal docs are messy. They have code snippets, product names like “WidgetX Pro”, and internal abbreviations. Pure vector search might return semantically similar but irrelevant chunks. Pure keyword search misses synonyms or paraphrased answers. Hybrid fixes this by giving you both. At LogicLoop, we saw a 22% increase in MRR@10 after switching from vector-only to hybrid on our engineering wiki.
Here’s how we wired it up in Python with FastAPI. First, store your chunks in both systems:
from elasticsearch import Elasticsearch
es = Elasticsearch("http://localhost:9200")
es.index(index="docs", id chunk_id, body={"text": chunk_text, "metadata": metadata})
import asyncpg
conn = await asyncpg.connect(dsn=DB_DSN)
await conn.execute(
"INSERT INTO doc_chunks (id, embedding, text) VALUES ($1, $2, $3)",
chunk_id, embedding.tolist(), chunk_text
)
At query time, run both searches and fuse scores:
def hybrid_search(query_text, query_vector, k=10):
es_results = es.search(
index="docs",
body={
"query": {"match": {"text": query_text}},
"size": k * 2 # get more to fuse later
}
)
keyword_hits = [(hit["_id"], hit["_score"]) for hit in es_results["hits"]["hits"]]
vector_results = await conn.fetch(
"SELECT id, text, embedding <=> $1 AS distance FROM doc_chunks ORDER BY distance LIMIT $2",
query_vector, k * 2
)
vector_hits = [(row["id"], 1 - row["distance"]) for row in vector_results] # convert to similarity
all_scores = {}
max_es = max([s for _, s in keyword_hits]) if keyword_hits else 1
max_vec = max([s for _, s in vector_hits]) if vector_hits else 1
for doc_id, score in keyword_hits:
all_scores[doc_id] = all_scores.get(doc_id, 0) + score / max_es
for doc_id, score in vector_hits:
all_scores[doc_id] = all_scores.get(doc_id, 0) + score / max_vec
sorted_ids = sorted(all_scores.items(), key=lambda x: x[1], reverse=True)[:k]
return [doc_id for doc_id, _ in sorted_ids]
We normalize scores to [0,1] per system before adding. This prevents one system from dominating due to scale differences.
We tested on 500k internal Confluence pages and Jira tickets. Metrics: MRR@10 and recall@10. Baseline: pure vector (pgvector only).
| Method | MRR@10 | Recall@10 | Latency (p95) |
|---|---|---|---|
| Vector only | 0.41 | 0.68 | 120ms |
| Keyword only | 0.33 | 0.52 | 80ms |
| Hybrid | 0.50 | 0.79 | 150ms |
Hybrid added ~30ms latency but gained significant quality. The cost? Running two indexes doubles storage and write overhead. We mitigated this by using Elasticsearch’s snapshot lifecycle and pgvector’s partitioning by doc source.
Failure mode we hit: when query text is very short (1-2 words), keyword search dominates and vector adds noise. We now route sub-3-word queries to keyword-only via a simple length check.
See our full RAG pipeline diagram for context, but here’s the hybrid-specific flow:
We keep the fusion logic in a dedicated service so we can swap fusion algorithms (like RRF or learning-to-rank) without touching the LLM layer.
knn and match combo in one query if on 8.0+.
BAAI/bge-small-en-v1.5).
We tried using Weaviate’s hybrid search but switched to ES+pgvector because we needed fine-grained control over indexing policies and couldn’t justify the ops overhead of a new system.
Does hybrid search always beat vector-only?
No. On clean, well-written documentation with minimal acronyms, vector-only can match hybrid. Test on your data.
How do I tune the weight between vector and keyword?
Start with equal weight (normalize then add). If you have relevance labels, use a validation set to sweep weights from 0.0 to 1.0 for vector (1-weight for keyword).
Can I use hybrid search with LLMs that have long context?
Yes. Hybrid improves the quality of chunks fed into the LLM, which matters more than raw context length when dealing with noisy internal data.
What embedding model works best for hybrid?
We use BAAI/bge-small-en-v1.5. It’s small, fast, and works well with keyword fusion. Larger models help marginally but add latency.
Hybrid Search RAG Implementation: A Practical Guide
RAG Chunking Best Practices for Production Systems
RAG Chunking Evaluation: Metrics, Trade-offs, and Production Lessons