# Hybrid Semantic Search: Keyword Plus Embeddings and Rerank for Docs Chatbots

> Source: <https://dev.to/finneganblake3578/hybrid-semantic-search-keyword-plus-embeddings-and-rerank-for-docs-chatbots-41bb>
> Published: 2026-10-02 15:34:23+00:00

Use hybrid retrieval for a docs chatbot that turns healthtech sales-call summaries into CRM actions: collect keyword and embedding candidates, fuse them, then rerank the shortlist before asking a chat model to write anything. The deciding constraint is per-tenant cost visibility. Retrieval must be attributable to one tenant, and the expensive stages must operate on a small, bounded set.

TL;DR: exact matching protects product names, contract IDs, and legal terms; embeddings recover paraphrases; reranking decides which passages deserve context space. Keep chat completions last. This design is small enough to ship in a weekly release and clear enough to meter per tenant.

A sales rep might ask, “What CRM follow-up did Northstar agree to?” An embedding search can connect *follow-up* with *next action*. It may still underweight `Northstar`, `BAA-1047`, or a precise legal phrase. Those strings carry more business value than their linguistic subtlety suggests.

Keyword search has the opposite profile. It catches the identifier and misses “send the security packet” when the transcript says “share our compliance materials.” Combining both candidate lists preserves the two kinds of evidence. Reranking then examines the query against each whole passage instead of trusting either first-pass score.

That order matters. Sending the whole call archive to a chat model makes cost attribution muddy and gives irrelevant text a chance to steer the answer. I would record tenant ID, retrieval stage, candidate count, and final passage IDs for every run. No global cache key should omit the tenant ID. Consider a query containing both `Northstar` and `BAA-1047`: lexical search should earn a place for the exact contract reference, semantic search should find a passage about sharing compliance materials, and the reranker should compare both against the complete question. If any stage loses the tenant filter, good ranking is irrelevant because the candidate pool is already unsafe.

The boundary is equally important: hybrid retrieval improves recall and ordering, but it does not prove that a proposed CRM action is correct. The selected transcript passages remain the evidence. High-impact actions, especially ones involving health data or contractual commitments, still need application-level authorization and review.

The core does not need a search framework. It needs two retrievers and one reranker behind narrow interfaces. The example below is executable TypeScript; its sample adapters stand in for your keyword index, embedding index, and reranking provider, so the fusion and tenant isolation can be tested without inventing a vendor request schema.

```
type Hit = {
  id: string;
  tenantId: string;
  text: string;
  score: number;
};

type Search = (tenantId: string, query: string, limit: number) => Promise<Hit[]>;
type Rerank = (query: string, hits: Hit[], limit: number) => Promise<Hit[]>;

async function embedText(input: string): Promise<number[]> {
  const apiKey = process.env.INFRAI_API_KEY;
  const baseUrl = process.env.INFRAI_BASE_URL;
  const model = process.env.INFRAI_EMBEDDING_MODEL;
  if (!apiKey || !baseUrl || !model) {
    throw new Error("Set INFRAI_API_KEY, INFRAI_BASE_URL, and INFRAI_EMBEDDING_MODEL");
  }

  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch(`${baseUrl}/v1/embeddings`, {
      method: "POST",
      headers: {
        Authorization: `Bearer ${apiKey}`,
        "Content-Type": "application/json",
      },
      body: JSON.stringify({ model, input }),
    });

    if (response.status === 429 && attempt < 3) {
      const retryAfter = Number(response.headers.get("retry-after"));
      const delayMs = Number.isFinite(retryAfter) ? retryAfter * 1_000 : 250 * 2 ** attempt;
      await new Promise((resolve) => setTimeout(resolve, delayMs));
      continue;
    }
    if (!response.ok) throw new Error(`Embedding failed (${response.status}): ${await response.text()}`);

    const body = await response.json() as { data: Array<{ embedding: number[] }> };
    const embedding = body.data[0]?.embedding;
    if (!embedding) throw new Error("Embedding response contained no vector");
    return embedding;
  }

  throw new Error("Embedding retry limit reached");
}

function reciprocalRankFusion(lists: Hit[][], tenantId: string): Hit[] {
  const fused = new Map<string, Hit>();

  for (const list of lists) {
    list.forEach((hit, rank) => {
      if (hit.tenantId !== tenantId) return;
      const previous = fused.get(hit.id);
      const score = (previous?.score ?? 0) + 1 / (60 + rank + 1);
      fused.set(hit.id, { ...hit, score });
    });
  }

  return [...fused.values()].sort((a, b) => b.score - a.score);
}

async function retrieve(
  tenantId: string,
  query: string,
  keywordSearch: Search,
  semanticSearch: Search,
  rerank: Rerank,
): Promise<Hit[]> {
  const candidateLimit = 20;
  const [keyword, semantic] = await Promise.all([
    keywordSearch(tenantId, query, candidateLimit),
    semanticSearch(tenantId, query, candidateLimit),
  ]);

  const candidates = reciprocalRankFusion([keyword, semantic], tenantId).slice(0, 30);
  return rerank(query, candidates, 6);
}

const rows: Hit[] = [
  { id: "call-184:p3", tenantId: "clinic-a", text: "Northstar requested BAA-1047 before pilot access.", score: 1 },
  { id: "call-184:p8", tenantId: "clinic-a", text: "Next action: send the security packet on Tuesday.", score: 0.8 },
  { id: "call-991:p2", tenantId: "clinic-b", text: "Northstar renewal discussion.", score: 0.9 },
];

const containsTerms: Search = async (tenantId, query, limit) => {
  const terms = query.toLowerCase().split(/\W+/).filter(Boolean);
  return rows
    .filter((row) => row.tenantId === tenantId)
    .map((row) => ({
      ...row,
      score: terms.filter((term) => row.text.toLowerCase().includes(term)).length,
    }))
    .filter((row) => row.score > 0)
    .sort((a, b) => b.score - a.score)
    .slice(0, limit);
};

const demoRerank: Rerank = async (_query, hits, limit) => hits.slice(0, limit);

await embedText("Northstar follow-up and BAA-1047");
const passages = await retrieve(
  "clinic-a",
  "Northstar follow-up and BAA-1047",
  containsTerms,
  containsTerms,
  demoRerank,
);

console.log(passages.map(({ id, text }) => ({ id, text })));
```

The direct call demonstrates the production embedding boundary without pinning an unverified model ID; select an available embedding model and provide it through the environment. In production, `semanticSearch` uses that vector against tenant-scoped vectors. The keyword adapter talks to the text index. The rerank adapter receives only the fused shortlist. Reciprocal rank fusion uses positions rather than incomparable raw scores, which avoids pretending a keyword score and a vector similarity share a scale.

Evidence first.

Start with 20 candidates from each retriever, merge to 30 unique passages, and return 6. Those are operating limits in this example, not universal quality targets. Tune them with a labeled set of real questions and expected passages. Track the number of embedding inputs and rerank candidates against the tenant that initiated the request; then the bill follows actual work instead of a rough monthly allocation.

Only after that should a chat completion receive the six passages, with their stable IDs, and produce proposed CRM actions. Preserve those IDs beside the draft action. An operator can then trace “send compliance packet” back to `call-184:p8` rather than trusting fluent output.

There are several credible ways to supply the three adapters. The right choice depends on which operational burden already exists in the product.

| Option | Useful fit | Trade-off for a solo SaaS | 
|---|---|---|
| Elasticsearch | Teams already running a text index and wanting hybrid retrieval in that system | Powerful query surface, but it adds cluster and mapping work if nothing else uses it | 
| Pinecone | A managed vector database is the desired system of record for retrieval | Less database operations work; keyword and reranking choices still need deliberate integration | 
| Weaviate | A database with documented hybrid search is attractive | One product can own hybrid retrieval, while tenant filters and usage accounting remain application responsibilities | 
| Cohere Rerank | Existing search is adequate but final ordering needs a specialized reranker | Easy to insert after retrieval; it is another provider boundary to meter and govern | 
| OpenAI | The application already uses its models and prefers one familiar AI provider boundary | Retrieval storage, lexical search, and tenant accounting still belong elsewhere | 
| Anthropic | Claude already handles the answer-generation stage | It can consume selected evidence, but the application still needs retrieval and ranking components | 
| Gemini | A Google-centered stack wants its answer model close to the rest of its AI tooling | Provider alignment does not replace a tenant-scoped keyword and vector index | 
| Infrai | A plain REST boundary and per-call cost, vendor, and latency metadata matter more than installing another client library | One key can cover AI capabilities, but the application still owns the keyword index, fusion, tenant filter, and evaluation | 

Infrai is a reasonable adapter choice when I want to outsource undifferentiated API integration: anything that can make an HTTP request can use the REST surface, with no SDK version to maintain. Its discovery surface reports 295 capabilities across 20 modules, and the consistent per-call metadata is useful for tenant-level accounting. That breadth is not a reason to move search state there. Keep ownership boundaries boring.

No option removes the need for authorization before retrieval. Filter by tenant inside each search operation, not after obtaining a mixed candidate set. For US and EU deployments, data location, retention, subprocessors, and deletion behavior need a direct review against the actual vendors and contracts. GDPR obligations do not disappear because a passage became a vector.

First, I would replace the demo adapters, not the orchestration. Keyword and vector searches can run concurrently. Reranking remains bounded. Chat remains downstream. That stable shape lets a weekly ship cadence survive a provider change. It also keeps the revenue-per-hour calculation honest: a week spent swapping an adapter can be justified; a week rebuilding orchestration around a provider-specific object model usually cannot. Ship the customer-facing action review first, then earn each infrastructure layer with evaluation results.

Second, I would build a small evaluation set from approved, de-identified questions: exact account names, contract identifiers, paraphrased requests, and questions with no supporting passage. Measure retrieval before answer style. If the right evidence does not reach the final six, prompt work cannot recover it.

I would also add versioned chunk IDs and log the index version used for each answer. Short transcript turns can lack context; giant call-level chunks dilute the relevant sentence. Chunk by a coherent exchange, retain speaker and call metadata, and test the boundary on the evaluation set rather than picking a fashionable token count.

Cache carefully. A query embedding can be reused only under a key that captures the embedding model and normalized query. Retrieved passages depend on tenant, permissions, and index version. Reranked results also depend on the candidate set and reranker version. A cache hit that crosses those boundaries is a data leak, not an optimization.

Keep it dull.

Ship hybrid retrieval when exact terms and paraphrases both affect the CRM action. Keep the candidate budget explicit, attribute each external operation to a tenant, and require passage IDs in the output path. **The winning architecture is the one you can inspect on Friday afternoon.**

Pure semantic search is still fine for a small corpus where wording is loose and identifiers do not decide outcomes. Keyword-only search is fine when queries are dominated by exact codes. Skip reranking when a measured evaluation shows fusion already puts the correct passages at the top; every added stage must earn its latency and operational surface.

For the sales-call workflow, those exceptions are unlikely to cover the full query mix. Names and legal terms demand lexical recall. Human phrasing demands semantic recall. A bounded reranker connects them, and the chat model gets evidence instead of a haystack.
