# bge-reranker Truncates at 512 Tokens: 31% of My Chunks Got Cut

> Source: <https://dev.to/ji_ai/bge-reranker-truncates-at-512-tokens-31-of-my-chunks-got-cut-2gm9>
> Published: 2026-10-10 04:50:26+00:00

My RAG pipeline had a question it kept getting wrong. "How do I fix the replication lag alert on the orders DB?" The answer was in my runbooks. Retrieval found the right chunk at rank 3. Then the reranker moved it to rank 14, my top-5 cutoff threw it away, and the LLM confidently answered from a chunk about a *different* database.

The reranker wasn't dumb. It was blind. My bge-reranker truncates at 512 tokens, and the fix commands lived at the bottom of a chunk it never finished reading. When I counted, 31% of my chunks were longer than what the reranker could see.

This post is about that one mechanism: why a cross-encoder reranker truncates at 512 tokens, which part of your text it throws away, and what to do about it.

`BAAI/bge-reranker-base` read the query and passage as `longest_first` truncation, which trims the `(query, passage)` with the `max_length`.
Here's the setup, because the details are where this bug hides:

`RecursiveCharacterTextSplitter`, `chunk_size=2000`. Characters, not tokens.` text-embedding-3-small`, top 20 by cosine similarity.` CrossEncoder("BAAI/bge-reranker-base")` from sentence-transformers, keep top 5.
I had 60 hand-labeled questions. Retrieval put the correct chunk somewhere in the top 20 for 54 of them. After reranking, the correct chunk made the top 5 for only 41.

So the reranker was *losing* answers that retrieval had already found. That's backwards. The whole point of a reranker is to be the smarter, slower second pass.

I went through the 13 losses by hand. In 9 of them, the sentence that actually answered the question sat in the last third of the chunk. Runbooks are written that way: title, symptoms, context, and then, at the very bottom, the `## Resolution` section with the commands you need.

A cross-encoder reranker truncates at 512 tokens because it is a BERT-style encoder with a fixed number of position embeddings, and it reads the query and passage together as a single input. `bge-reranker-base` is built on XLM-RoBERTa, which was trained with 512 positions. Anything past that has no position to sit in, so the tokenizer cuts it before the model runs.

This is different from an embedding model (a bi-encoder). A bi-encoder embeds the query and the passage separately. A cross-encoder concatenates them:

```
<s> query tokens </s></s> passage tokens </s>
```

That's 4 special tokens for XLM-RoBERTa pairs. So your real passage budget is:

```
512 - 4 - len(query_tokens)
```

A 30-token question leaves 478 tokens for the passage. That joint attention over query and passage is exactly why cross-encoders rank better than cosine similarity. It's also why they have this hard ceiling.

The end of the passage gets cut. sentence-transformers' `CrossEncoder` calls the tokenizer with `truncation=True`, which in Hugging Face means the `longest_first` strategy. It removes tokens one at a time from whichever sequence is currently longer. Your query is 30 tokens and your passage is 700, so the passage loses its tail every single time.

You can watch it happen:

``` python
from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("BAAI/bge-reranker-base")

enc = tok(query, passage, truncation=True, max_length=512)
seen = tok.decode(enc["input_ids"], skip_special_tokens=True)

print(seen[-300:])  # the last thing the reranker actually read
```

When I ran this on my replication-lag chunk, the last thing the reranker read was the middle of a paragraph explaining what replication lag *is*. The `pg_stat_replication` query and the fix were gone. The model scored "a chunk that explains the concept" against "a short chunk about a different DB that mentions the alert name in full," and the short chunk won. That's a reasonable judgment on the text it was given.

The silent part is what makes this nasty. With `truncation=True`, the tokenizer does not warn you. Your scores look like normal floats. Nothing in your logs says "I only read 60% of this."

Retrieval didn't catch it because every stage measured the chunk with a different ruler:

| Stage | Unit | Limit | 
|---|---|---|
| Chunker | characters | 2,000 | 
| `text-embedding-3-small` | OpenAI tokens | 8,191 | 
| `bge-reranker-base` | XLM-RoBERTa tokens, query included | 512 | 

The embedding model saw the whole chunk, including the Resolution section, and ranked it well. The reranker saw a truncated version of the same chunk and ranked it badly. The stage with the smallest window had the final say.

The character limit made it worse. 2,000 characters of English prose fits under 512 tokens comfortably. 2,000 characters of runbook does not. Shell commands, hostnames, UUIDs, stack traces and YAML break into many more subword tokens than ordinary words do. My chunks that overflowed were almost all the operational ones, which are exactly the ones people ask about.

Tokenize every `(query, chunk)` pair with the reranker's tokenizer, without truncation, and count how many exceed the limit. It takes a few minutes on a few thousand chunks.

``` python
from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("BAAI/bge-reranker-base")
MAX_LEN = 512
TYPICAL_QUERY = "how do I fix the replication lag alert on the orders db"

def overflow(passage: str) -> int:
    ids = tok(TYPICAL_QUERY, passage, truncation=False)["input_ids"]
    return len(ids) - MAX_LEN

over = [c for c in chunks if overflow(c) > 0]
print(f"{len(over)}/{len(chunks)} chunks truncated "
      f"({len(over) / len(chunks):.0%})")
```

Mine printed 31%. Two more checks worth adding:

There are three fixes, and I ended up using the first two together.

**1. Chunk to the reranker's budget, measured with the reranker's tokenizer.** Reserve room for the query and the special tokens, then cap passages at what's left.

```
QUERY_RESERVE = 64      # longest query you expect, in reranker tokens
PASSAGE_BUDGET = 512 - 4 - QUERY_RESERVE   # 444

def token_len(text: str) -> int:
    return len(tok(text, add_special_tokens=False)["input_ids"])

splitter = RecursiveCharacterTextSplitter(
    chunk_size=PASSAGE_BUDGET,
    chunk_overlap=60,
    length_function=token_len,
)
```

The one-line change that matters is `length_function=token_len`. Now the chunker and the reranker use the same ruler.

**2. Score overlapping windows and take the max.** Some documents shouldn't be split, or you can't re-index right now. Slide a window over the passage, score each window, and keep the highest. In IR research this is known as MaxP scoring.

``` python
def windowed_score(model, query, passage, window=400, stride=200):
    ids = tok(passage, add_special_tokens=False)["input_ids"]
    if len(ids) <= window:
        return float(model.predict([(query, passage)])[0])

    starts = list(range(0, len(ids) - window + 1, stride))
    if starts[-1] + window < len(ids):
        starts.append(len(ids) - window)   # always cover the tail

    windows = [tok.decode(ids[s:s + window]) for s in starts]
    return float(max(model.predict([(query, w) for w in windows])))
```

Watch the `starts.append` line. Without it, a passage whose length isn't a neat multiple of the stride loses its last few dozen tokens. That would be the same bug you're fixing, just smaller.

**3. Use a reranker trained on longer inputs.** This works, but read the cost. Self-attention scales with the square of sequence length, and you run the reranker once per candidate. Going from 512 to 2,048 tokens on 20 candidates per query is a real latency change, so measure it on your own hardware before shipping.

After re-chunking with the reranker's tokenizer and using windowed scoring for the few long documents I kept whole, the correct chunk made the top 5 for 51 of my 60 questions, up from 41. Same embedding model, same reranker, same LLM. The only difference was that the reranker could now read the whole chunk.

Every model in a RAG pipeline has a context window, and the smallest one controls quality. People check the LLM's window and maybe the embedding model's. Almost nobody checks the reranker's, because it returns a clean float either way.

If your documents put the important part at the end, like runbooks, API docs with examples at the bottom, or support tickets with the resolution last, you're in the worst case for `longest_first` truncation.

Yes. `bge-reranker-base` and similar BERT- and XLM-RoBERTa-based cross-encoders read the query and passage as one sequence capped at 512 tokens, including special tokens. sentence-transformers truncates with `longest_first`, which silently drops the end of the passage. If your chunks are sized in characters, or in another model's tokens, a meaningful share of them may be partly invisible to your reranker, and any answer in that tail can't help the score. Measure overflow with the reranker's own tokenizer, size chunks to its budget, and use overlapping-window max scoring for passages you can't split.

*Written by the developer behind [Preterview](https://preterview.com/en), an interview prep platform.*
