Hybrid Semantic Search: Keyword Plus Embeddings and Rerank for Docs Chatbots A developer outlined a hybrid retrieval architecture for a multi-tenant healthtech docs chatbot that fuses keyword and embedding candidate lists and then reranks the shortlist before passing passages to a chat model, keeping chat completions as the final step. The design emphasizes per-tenant cost visibility and isolation, requiring tenant IDs at every retrieval stage and bounding the expensive reranking work to a small candidate set. The author notes hybrid retrieval improves recall and ordering but does not validate proposed CRM actions, which still require application-level authorization and review for high-impact, health-data or contractual decisions. Use hybrid retrieval for a docs chatbot that turns healthtech sales-call summaries into CRM actions: collect keyword and embedding candidates, fuse them, then rerank the shortlist before asking a chat model to write anything. The deciding constraint is per-tenant cost visibility. Retrieval must be attributable to one tenant, and the expensive stages must operate on a small, bounded set. TL;DR: exact matching protects product names, contract IDs, and legal terms; embeddings recover paraphrases; reranking decides which passages deserve context space. Keep chat completions last. This design is small enough to ship in a weekly release and clear enough to meter per tenant. A sales rep might ask, “What CRM follow-up did Northstar agree to?” An embedding search can connect follow-up with next action . It may still underweight Northstar , BAA-1047 , or a precise legal phrase. Those strings carry more business value than their linguistic subtlety suggests. Keyword search has the opposite profile. It catches the identifier and misses “send the security packet” when the transcript says “share our compliance materials.” Combining both candidate lists preserves the two kinds of evidence. Reranking then examines the query against each whole passage instead of trusting either first-pass score. That order matters. Sending the whole call archive to a chat model makes cost attribution muddy and gives irrelevant text a chance to steer the answer. I would record tenant ID, retrieval stage, candidate count, and final passage IDs for every run. No global cache key should omit the tenant ID. Consider a query containing both Northstar and BAA-1047 : lexical search should earn a place for the exact contract reference, semantic search should find a passage about sharing compliance materials, and the reranker should compare both against the complete question. If any stage loses the tenant filter, good ranking is irrelevant because the candidate pool is already unsafe. The boundary is equally important: hybrid retrieval improves recall and ordering, but it does not prove that a proposed CRM action is correct. The selected transcript passages remain the evidence. High-impact actions, especially ones involving health data or contractual commitments, still need application-level authorization and review. The core does not need a search framework. It needs two retrievers and one reranker behind narrow interfaces. The example below is executable TypeScript; its sample adapters stand in for your keyword index, embedding index, and reranking provider, so the fusion and tenant isolation can be tested without inventing a vendor request schema. type Hit = { id: string; tenantId: string; text: string; score: number; }; type Search = tenantId: string, query: string, limit: number = Promise