bge-reranker Truncates at 512 Tokens: 31% of My Chunks Got Cut A developer found that the bge-reranker-base cross-encoder silently truncates passages at 512 tokens, cutting the tail of 31% of their RAG chunks and causing the reranker to demote correct answers that retrieval had already found. In a 60-question evaluation, retrieval placed the correct chunk in the top 20 for 54 questions, but after reranking only 41 survived the top-5 cutoff, with 9 of the 13 losses having the answering sentence in the final third of the chunk. The cause is XLM-RoBERTa's fixed 512 position embeddings combined with Hugging Face's longest_first truncation, which leaves roughly 512 minus 4 special tokens minus the query length for the passage. My RAG pipeline had a question it kept getting wrong. "How do I fix the replication lag alert on the orders DB?" The answer was in my runbooks. Retrieval found the right chunk at rank 3. Then the reranker moved it to rank 14, my top-5 cutoff threw it away, and the LLM confidently answered from a chunk about a different database. The reranker wasn't dumb. It was blind. My bge-reranker truncates at 512 tokens, and the fix commands lived at the bottom of a chunk it never finished reading. When I counted, 31% of my chunks were longer than what the reranker could see. This post is about that one mechanism: why a cross-encoder reranker truncates at 512 tokens, which part of your text it throws away, and what to do about it. BAAI/bge-reranker-base read the query and passage as longest first truncation, which trims the query, passage with the max length . Here's the setup, because the details are where this bug hides: RecursiveCharacterTextSplitter , chunk size=2000 . Characters, not tokens. text-embedding-3-small , top 20 by cosine similarity. CrossEncoder "BAAI/bge-reranker-base" from sentence-transformers, keep top 5. I had 60 hand-labeled questions. Retrieval put the correct chunk somewhere in the top 20 for 54 of them. After reranking, the correct chunk made the top 5 for only 41. So the reranker was losing answers that retrieval had already found. That's backwards. The whole point of a reranker is to be the smarter, slower second pass. I went through the 13 losses by hand. In 9 of them, the sentence that actually answered the question sat in the last third of the chunk. Runbooks are written that way: title, symptoms, context, and then, at the very bottom, the Resolution section with the commands you need. A cross-encoder reranker truncates at 512 tokens because it is a BERT-style encoder with a fixed number of position embeddings, and it reads the query and passage together as a single input. bge-reranker-base is built on XLM-RoBERTa, which was trained with 512 positions. Anything past that has no position to sit in, so the tokenizer cuts it before the model runs. This is different from an embedding model a bi-encoder . A bi-encoder embeds the query and the passage separately. A cross-encoder concatenates them: