Lost in the Middle: My RAG Found the Right Chunk and the LLM Ignored It A developer building a RAG-based support bot traced a wrong answer to the "lost in the middle" effect, in which an LLM ignored the correct policy chunk placed at position #7 of a 10-chunk prompt and instead used a higher-scoring warranty paragraph near the start. Running a position sweep over 50 real support questions with 10 chunks each (500 calls), the developer measured 45/50 correct answers when the gold chunk was first, 43/50 when last, and roughly 33/50 in slots 5 through 7, confirming the U-shaped accuracy curve reported in the 2023 Liu et al. paper. The writeup attributes the effect to stacked biases in decoder-only transformers, including primacy from causal attention and attention sinks on early tokens. The retrieval log said I was right. The answer said I was wrong. A user asked my support bot how long refund requests stay open. The retriever pulled 10 chunks, and chunk 7 was the exact paragraph from the policy doc: "Refund requests expire after 14 days." The model answered "30 days," which came from chunk 2, a paragraph about warranty claims. I spent an evening blaming the embeddings. The embeddings were fine. The problem was lost in the middle : LLMs use information at the start and end of a long prompt far better than information buried in the middle. My retriever did its job. Then I put the answer in the one place the model barely reads. The lost in the middle effect is the drop in answer accuracy when the relevant information sits in the middle of a long prompt instead of near the beginning or end. The name comes from the 2023 paper "Lost in the Middle: How Language Models Use Long Contexts" by Liu et al., which ran multi-document question answering with the answer-bearing document moved through every position. The shape they found was a U. High accuracy when the gold document came first, a bit lower when it came last, and a sag in between. In some of their settings, a model with the answer sitting in the middle of its context did worse than the same model given no documents at all. Read that again. Adding the correct document, in the wrong spot, made things worse than adding nothing. The nine distractors around it did more damage than the one right answer did good. Run a position sweep. Take questions where you know which chunk holds the answer, keep the other chunks fixed, and move the gold chunk through every slot. If accuracy by slot looks like a smile, you have the problem. Here is the harness I used. call llm and is correct are whatever you already have. python import random def build prompt chunks, question : docs = "\n\n".join f"