# Add a Reranker or Rebuild Your AI Search Index First?

> Source: <https://www.digitalapplied.com/blog/reranker-vs-reembedding-search-diagnosis>
> Published: 2026-10-10 00:00:00+00:00

Add a reranker when the right evidence is already in the candidate list but appears too low to be used. Investigate ingestion, filters, chunking and first-stage retrieval when that evidence never arrives. Rebuilding the index makes sense only after you have identified a problem the new representation can plausibly fix. The distinction can save a team from an expensive migration that preserves the original failure.

1. 01Inspect candidates firstFind out whether the evidence exists in the material sent to the ranking stage.
2. 02Change one layerA controlled comparison needs a stable corpus and relevance judgments.
3. 03Preserve access filtersBetter relevance never makes an unauthorized document acceptable.
4. 04Migrate with a fallbackKeep the old index available until the replacement passes task-level checks.

## 01 — Failure locationFollow one failed question to its *source*

Start with a real question that the application answered badly and an approved document containing the expected evidence. Trace the document through ingestion, indexing, candidate selection, ranking and final answer construction. This turns a vague complaint about search quality into a specific observation: the source was missing, its passage was excluded, it ranked too low or the answer ignored it.

Consider a hypothetical equipment-support assistant. The relevant manual contains an exception for a particular serial-number range. The assistant returns the general maintenance rule instead. If the exception paragraph was never extracted from the PDF, neither a new embedding model nor a reranker can recover it from an index that does not contain it. Repairing document extraction is the first useful change.

Keep the question, expected passage, candidate identifiers and final context together. The [RAG business guide](https://www.digitalapplied.com/blog/rag-retrieval-augmented-generation-business-guide) explains the overall architecture; this diagnosis concerns where a particular answer went missing. Do not rely on the final prose to infer which documents were actually available to the model.

Use the expected passage to distinguish an evidence problem from a question-design problem. A reviewer may initially believe the manual answers a request, then discover that it covers a different product version. In that case, the honest answer key is uncertainty, not the nearest paragraph. A retrieval experiment built on an incorrect answer key can reward a system for finding the wrong document very consistently, so review the governing version and applicability before interpreting scores.

##### Repair candidate retrieval

Check ingestion, filters, query handling and the representation used to find candidates.

##### Test a reranker

Measure whether the required passage moves into the context the answer model receives.

## 02 — Candidate boundaryUnderstand what a reranker can *change*

A reranker evaluates the relationship between the query and documents supplied to it, then returns a relevance order or scores. It is usually a second stage after lexical or vector search. The Voyage reranker documentation, checked October 11, 2026, lists Rerank 3 and Rerank 3 Lite and describes this candidate-list workflow. That is a capability description, not an independent result on your corpus.

The boundary is decisive: a reranker cannot promote an item it was never given. It can help when a useful exception paragraph is among the retrieved candidates but broad overview pages outrank it. It cannot help if a region filter excluded the exception, the source connector failed or an overly narrow initial search returned only the overview.

Preserve document identifiers through this stage. A ranking API may return an index into the submitted list, which must map back to the correct source and passage. Reordering a list and then attaching citations using the old order creates a convincing answer with the wrong evidence. That is an application bug, and better relevance scores will not expose it automatically.

Reranking also changes the order in which the answer model encounters material. If the application sends only a few passages, a small rank change can determine whether a crucial exception is included at all. If it sends the entire candidate set, ranking may have a different effect and a larger context cost. Record the actual selection rule after reranking; otherwise, the test evaluates an ordering that the production answer path does not use.

A high relevance score is not a calibrated probability that a final answer is correct. Use judged examples and the application outcome rather than inventing a universal score threshold.

## 03 — Retrieval recallCheck the candidate list before changing *models*

Inspect whether the expected passage appears within the candidate budget available to the next stage. For a question with several required sources, mark every required passage rather than counting a single vaguely relevant document as success. The practical question is whether the answer could have been supported from the retrieved material alone.

Try a small set of questions with exact product codes, paraphrased requests and ambiguous terminology. Lexical search may recover a code that semantic search misses, while semantic search may recover a concept expressed differently in the source. Our [hybrid search reference](https://www.digitalapplied.com/blog/hybrid-search-bm25-vector-reranking-reference-2026) explains that combination; use it to test a specific miss rather than adding another component because it is fashionable.

Inspect filters with the same care. A current-document or tenant filter can be correct in principle and wrong in implementation. Expanding the candidate count does not repair a filter that removes the governing document. Preserve the authorization boundary while diagnosing relevance: temporarily removing it from a production query is not a valid way to prove that retrieval improved.

Do not enlarge every candidate list by default. First inspect a miss at the current limit, then test whether a broader list contains the missing evidence. If it does, measure the additional ranking and context burden. If it does not, the failure is probably elsewhere in the path. This staged experiment distinguishes a budget limitation from an indexing or representation problem and gives the team a reason for any extra latency it accepts.

- Confirm the expected passage was ingested and indexed.
- Check filters and identifiers before changing similarity settings.
- Distinguish exact-code misses from semantic paraphrase misses.

## 04 — Evidence unitsInspect chunks before paying for new *embeddings*

An index stores the material produced by your extraction and chunking process. If a heading is separated from its exception, the returned passage may be hard to interpret even when the model represents it well. If a large chunk blends unrelated topics, its similarity can reflect the general theme while obscuring the detail the customer needs.

For the hypothetical manual, retain the serial-number condition beside the maintenance instruction and preserve a link to the original location. Compare a small set of alternative chunk boundaries against the same questions. Do not rewrite the source or add an explanatory answer to the chunk just to make the test pass; that changes the information available to the system.

The [chunking playbook](https://www.digitalapplied.com/blog/rag-chunking-strategies-2026-retrieval-quality-playbook) covers the broader choices. Here, record whether a boundary change fixes the observed miss before considering a full model migration. An embedding rebuild over better chunks and a rebuild using a new model are different experiments, even if both require writing new vectors.

Tables deserve a separate inspection because extraction can separate a row label from its value or repeat column headings inconsistently. A passage that contains the right number without its unit is not adequate evidence. Keep representative table questions in the evaluation and inspect the serialized chunk the model receives. A better embedding of an ambiguous fragment may rank it more confidently without restoring the missing relationship between the label, number and condition.

Read the exact text submitted for embedding and reranking, not only the original document. Parsing, truncation and chunk boundaries can change what the system sees.

## 05 — Controlled trialRun a fixed-candidate ranking *comparison*

Freeze a representative set of questions and the first-stage candidate lists. Have reviewers identify the passages required for a useful answer, including acceptable alternatives. Run the existing ordering and the proposed reranker over those same candidates. This isolates the effect of ranking from changes in document ingestion or candidate retrieval.

Record whether the required evidence reaches the context budget, how much irrelevant material remains and whether important exceptions survive. Also inspect latency and usage for the actual list lengths. Sending more documents increases the opportunity to recover evidence but can add cost and delay. The right limit is a workload decision, not the maximum number the endpoint accepts.

The examples in this guide are proposed tests, not a report of a completed benchmark. If a trial improves one document class and worsens another, keep the split visible. A single average can conceal that scanned manuals or multilingual questions remain unusable. The deployment decision should identify the supported task classes and the unresolved ones.

Have reviewers judge relevance without seeing the candidate model's identity where practical. Preserve disagreements instead of forcing every ambiguous case into a clean pass or fail. Some questions legitimately admit several useful sources, while others require one governing exception. Those differences should be reflected in the answer key. A simple record of sufficient, partial and irrelevant evidence is often more actionable than an elaborate score whose categories nobody can explain.

| Diagnostic routing suggestions, not measured findings or guarantees of a particular repair. |  |  | 
|---|---|---|
| Observation | Likely next investigation | Keep fixed | 
|---|---|---|
| Passage never retrieved | Ingestion, filters or first-stage search | Authorized corpus | 
| Passage retrieved too low | Reranking or context selection | Candidate lists | 
| Passage loses its condition | Chunking or truncation | Source meaning | 
| Evidence supplied, answer wrong | Generation and citation behavior | Retrieved context | 

## 06 — Migration scopeRebuild only when the representation is the *problem*

A different embedding model may be worth testing when the current representation repeatedly misses relevant material despite sound ingestion, filters and query handling. Language coverage, domain terminology or document modality can motivate that test. The motivation is a hypothesis to investigate, not proof that the new model will perform better.

Build a separate candidate index and preserve the original document identifiers, permissions and version metadata. Unless the provider explicitly documents compatible embedding spaces for the selected models, use the corresponding query representation for each index. Equal vector dimensions do not establish compatibility; they only establish that the arrays have the same length.

Compare the indexes on the same question set before switching traffic. Include the costs of reprocessing documents, storing parallel indexes and maintaining updates during the trial. A cheap embedding rate does not make an incomplete migration cheap. Missing permission metadata or stale updates can produce a search system that appears more relevant while becoming less trustworthy.

Plan how new and changed documents reach both indexes during the trial. A candidate index built once from a snapshot will naturally become stale if the original continues receiving updates. Comparing them a week later without accounting for that difference confounds freshness with model quality. Keep synchronization status visible and repeat the comparison on matched source versions before deciding that a new representation has improved or degraded retrieval.

- Version the embedding model, preprocessing and chunking together.
- Maintain both query paths during a controlled comparison.
- Verify document updates and access rules on the candidate index.

## 07 — End-to-end checkEvaluate the answer after retrieval *improves*

Improved ranking is useful only if the answer or task improves. Supply the selected context to the same answer workflow and inspect factual support, missing exceptions and citations. A longer collection of relevant passages can still produce a confused answer if the prompt cannot resolve conflicting versions or the model overlooks the governing condition.

Use cases where the correct outcome is to ask for clarification or decline to answer. Retrieval can return plausible material for a question the corpus does not actually settle. An agent that treats every result list as sufficient evidence will confidently answer outside its knowledge even after its ranking improves. That failure belongs in the acceptance decision.

Our [AI transformation service](https://www.digitalapplied.com/services/ai-transformation) treats retrieval as part of the operating workflow, including what the user sees when evidence is incomplete. Keep a human-readable explanation of the test result: which questions now work, why they work and what still fails. That is more useful to an operator than a model name and a higher aggregate score.

Check citations at the passage level. The answer may state the right rule while linking to a general overview that does not contain the exception. A user following that citation cannot verify the important claim. Preserve the source location through ranking and context assembly, and verify that the link opens the expected authorized material. Retrieval quality includes giving the reader usable evidence, not merely supplying hidden context that produces a plausible sentence.

Inspect the actual context delivered to the answer model. A reranker can select the correct passage while a later truncation step removes it.

## 08 — Operating decisionKeep the smallest repair that earns its *place*

Sometimes the best result is a filter fix, a better extraction step or a modest change to candidate selection. Sometimes a reranker consistently brings the required evidence into view. Sometimes a new embedding index is justified. The diagnostic sequence prevents those choices from being treated as interchangeable upgrades.

Write a rollout rule based on the observed improvement and its boundary. A reranker might apply to long policy queries while exact identifier lookups continue through a simpler path. A new index might begin with a document class whose failures are well understood. Keep the former route available until updates, latency and permission handling have been exercised under realistic conditions.

Keep the failure examples as regression tests. When a later model or connector update arrives, those cases tell you whether useful behavior survived. Search quality is not a property you purchase once with an embedding model; it is a chain of information-preserving decisions that needs to remain observable as the corpus and workload change.

Write down the cases where the chosen repair does not apply. A reranker that improves policy questions may add unnecessary latency to exact record lookups. A new embedding model may require a separate route for image-heavy documents. These boundaries are not weaknesses to conceal; they are the operating rules that keep a successful narrow experiment from becoming an unsupported universal migration across every search feature in the application.

- Repair the stage that lost the evidence.
- Judge ranking and final answers separately.
- Retain a tested fallback and the difficult examples.

### Trace a miss before rebuilding the index

Pick a failed question with a known supporting passage. Find the stage where that passage disappears or becomes unusable, then change one part of the system.

A reranker is a useful ordering tool. A new embedding index is a representation change. Choose between them with evidence from the retrieval path.
