A game catalog changes under your feet. Publishers rename editions, players use abbreviations, and yesterday's description can disappear after a rights update. That operational constraint changes the answer: combine keyword and semantic retrieval, but keep both behind a small Node.js contract. Do not let an embedding provider's response shape become your catalog model.
TL;DR: exact terms should go to lexical search, fuzzy intent should get an embedding signal, and the final rank should combine both with observable rules. Store the source text, embedding model identifier, and content version as one revision. That makes re-indexing deliberate and provider changes survivable.
The before/after mental model is short. Before: description -> one search service -> result. After: versioned description -> lexical index plus vector index -> fused candidates -> enrichment job, with trace context and outcome counters crossing every arrow. Three signals matter: exact-match rank, semantic similarity, and source freshness.
Keyword search is excellent when a string carries identity. A query for NG+, 4K, a platform code, or an exact expansion title should not be softened into a vague neighborhood. Lexical indexes also make token matches inspectable. An operator can explain why a document appeared without reverse-engineering a vector.
Exact strings matter.
Semantic search handles a different failure: two descriptions can express the same play style with few shared words. A player may ask for a cooperative puzzle game while a publisher writes about solving rooms with a friend. Embeddings can bring those descriptions closer, but similarity is not truth. It is one ranking input.
The trade-off is concrete. Lexical-only retrieval misses paraphrases. Vector-only retrieval can bury decisive strings and makes model changes an index migration. A fused candidate set preserves both forms of evidence. Freshness is the third signal because a semantically close, retired edition is still the wrong input for enrichment.
This is where portability starts. The application owns the meaning of a CatalogHit; adapters own the mechanics of producing lexical and vector candidates. No provider-specific score crosses that boundary until it has been normalized and labeled.
Keep the contract boring. It should accept stable domain data, return enough evidence to debug a rank, and expose no SDK classes. The example below is intentionally small, but it contains the fields that are painful to reconstruct later.
export type CatalogDocument = {
id: string;
contentVersion: number;
title: string;
description: string;
active: boolean;
};
export type Candidate = {
documentId: string;
source: "lexical" | "semantic";
normalizedScore: number;
contentVersion: number;
};
export interface LexicalIndex {
search(query: string, limit: number): Promise<Candidate[]>;
}
export interface VectorIndex {
search(embedding: readonly number[], limit: number): Promise<Candidate[]>;
}
export interface Embedder {
readonly modelId: string;
embed(text: string): Promise<readonly number[]>;
}
modelId is data, not decoration. An embedding is meaningful only in the coordinate space that produced it, so a model change requires a new index generation or another explicit compatibility plan. Never silently mix generations.
Now fuse candidates without pretending their native scores are directly comparable. Each adapter must normalize its own results into the documented 0..1 range. The fusion layer applies product rules and records the component scores.
type RankedHit = {
documentId: string;
score: number;
lexicalScore: number;
semanticScore: number;
contentVersion: number;
};
export function fuseCandidates(
lexical: readonly Candidate[],
semantic: readonly Candidate[],
activeVersions: ReadonlyMap<string, number>,
): RankedHit[] {
const hits = new Map<string, RankedHit>();
for (const candidate of [...lexical, ...semantic]) {
if (activeVersions.get(candidate.documentId) !== candidate.contentVersion) continue;
const hit = hits.get(candidate.documentId) ?? {
documentId: candidate.documentId,
score: 0,
lexicalScore: 0,
semanticScore: 0,
contentVersion: candidate.contentVersion,
};
if (candidate.source === "lexical") {
hit.lexicalScore = Math.max(hit.lexicalScore, candidate.normalizedScore);
} else {
hit.semanticScore = Math.max(hit.semanticScore, candidate.normalizedScore);
}
hit.score = 0.6 * hit.lexicalScore + 0.4 * hit.semanticScore;
hits.set(candidate.documentId, hit);
}
return [...hits.values()].sort((a, b) => b.score - a.score);
}
The 0.6/0.4 split is an example policy, not a universal optimum. Start from the catalog's error cost. If confusing two editions is worse than missing a paraphrase, weight exact evidence more heavily. Tune against labeled queries, then version the policy beside the evaluation set.
One subtle pitfall lives in the activeVersions check. Deleting a catalog row is not enough when old vectors remain searchable. Filtering by the current content version prevents stale candidates from reaching enrichment while asynchronous index cleanup catches up. The same rule handles an edited description: version 18 cannot masquerade as version 19.
A search endpoint returning 200 does not prove that retrieval helped. Instrument the decision points. For each request, record the query class, index generation, embedding model identifier, candidate counts, selected document versions, and the final outcome. Avoid logging raw descriptions by default; catalog text can still carry licensed or embargoed material.
Use bounded labels for metrics. model_id, index_generation, result, and query_class have controlled cardinality. A document ID or raw query does not, so keep those in sampled traces or access-controlled logs when policy permits. OpenTelemetry provides vendor-neutral APIs and semantic conventions for carrying this telemetry, while W3C Trace Context defines interoperable trace propagation. A useful diagram in words looks like HTTP request -> query classifier -> lexical adapter and embedding adapter -> fusion -> catalog version check -> enrichment worker -> stored proposal. One trace spans those arrows. Metrics answer how often each branch succeeds, logs explain a particular failure, and alerts watch user-visible outcomes such as an empty eligible candidate set or a sustained enrichment rejection rate rather than paging on every adapter retry. This longer chain is useful during an index migration because a single request can reveal which generation ran, which branch timed out, which version passed the freshness check, and whether the worker stored a proposal.
Keep the result taxonomy compact: enriched, no_candidate, stale_candidate, policy_rejected, and dependency_error are enough to begin. This separation matters. A no-match query is a product signal; a dependency error is an operational signal. Combining them into one generic failure rate sends the team to the wrong dashboard.
Short requests may stream progress to a browser with Server-Sent Events. SSE uses the text/event-stream content type and reconnects by default, which suits one-way status updates. It does not replace the durable job record. The database remains the authority for state after a reconnect or process restart.
Provider portability is not the ability to change an environment variable. It is the ability to change an adapter while preserving behavior you have chosen explicitly. Build a small evaluation set from real catalog language, with sensitive text removed or governed appropriately. Include exact titles, abbreviations, paraphrases, renamed editions, inactive products, and descriptions with almost no useful detail.
For every adapter or model generation, run the same cases and retain the ranked evidence. Compare retrieval outcomes, not native similarity numbers. Score scales and distributions can differ, even when two systems return acceptable neighbors.
Use three gates before promotion:
The third gate intentionally has no borrowed percentage. Your threshold needs a labeled set and an error budget tied to this catalog. A public benchmark cannot tell you whether returning the base game for a query about downloadable content is acceptable.
Measure the miss.
Run old and new adapters side by side on a controlled sample, but let only the active path write enrichment proposals. The shadow path produces comparable traces and ranked outputs. Once its behavior passes the gates, switch the read path through configuration and keep rollback available until the new index generation is complete.
This costs extra compute during migration. It also turns a provider change from a blind cutover into an observable comparison. That is a worthwhile trade when enrichment can alter customer-facing catalog data.
The first objection is complexity: why not begin with keyword search and add vectors later? That can be correct. If queries are mostly exact game names, stock-keeping identifiers, and platform codes, a lexical index is a strong first release. Still use the domain contract and content versions. Those pieces are cheap early and prevent the search implementation from leaking across the service. Add semantic retrieval only when labeled misses show that paraphrase is a real problem.
The second objection is latency. Two retrieval paths do create more work, but they do not have to run serially. Start lexical search and embedding generation concurrently, apply independent timeouts, and define degraded behavior. If semantic retrieval misses its deadline, an exact lexical result may still be useful. If lexical retrieval fails, do not conceal that fact inside a blended score. Return the best policy-allowed result and emit the branch outcome.
No magic here. The architecture earns its keep through explicit failure modes: stale content is filtered, partial retrieval is visible, and model migrations create new generations.
The final decision rule is simple. Use lexical search alone while exact language covers the job. Add embeddings when measured paraphrase misses justify them. Once both exist, fuse evidence behind a stable Node.js interface, version every derived artifact, and judge providers by repeatable catalog outcomes rather than SDK ergonomics or opaque scores.