{"slug": "lost-in-the-middle-my-rag-found-the-right-chunk-and-the-llm-ignored-it", "title": "Lost in the Middle: My RAG Found the Right Chunk and the LLM Ignored It", "summary": "A developer building a RAG-based support bot traced a wrong answer to the \"lost in the middle\" effect, in which an LLM ignored the correct policy chunk placed at position #7 of a 10-chunk prompt and instead used a higher-scoring warranty paragraph near the start. Running a position sweep over 50 real support questions with 10 chunks each (500 calls), the developer measured 45/50 correct answers when the gold chunk was first, 43/50 when last, and roughly 33/50 in slots 5 through 7, confirming the U-shaped accuracy curve reported in the 2023 Liu et al. paper. The writeup attributes the effect to stacked biases in decoder-only transformers, including primacy from causal attention and attention sinks on early tokens.", "body_md": "The retrieval log said I was right. The answer said I was wrong.\n\nA user asked my support bot how long refund requests stay open. The retriever pulled 10 chunks, and chunk #7 was the exact paragraph from the policy doc: \"Refund requests expire after 14 days.\" The model answered \"30 days,\" which came from chunk #2, a paragraph about warranty claims.\n\nI spent an evening blaming the embeddings. The embeddings were fine. The problem was **lost in the middle**: LLMs use information at the start and end of a long prompt far better than information buried in the middle. My retriever did its job. Then I put the answer in the one place the model barely reads.\n\nThe lost in the middle effect is the drop in answer accuracy when the relevant information sits in the middle of a long prompt instead of near the beginning or end. The name comes from the 2023 paper \"Lost in the Middle: How Language Models Use Long Contexts\" by Liu et al., which ran multi-document question answering with the answer-bearing document moved through every position.\n\nThe shape they found was a U. High accuracy when the gold document came first, a bit lower when it came last, and a sag in between. In some of their settings, a model with the answer sitting in the middle of its context did worse than the same model given no documents at all.\n\nRead that again. Adding the correct document, in the wrong spot, made things worse than adding nothing. The nine distractors around it did more damage than the one right answer did good.\n\nRun a position sweep. Take questions where you know which chunk holds the answer, keep the other chunks fixed, and move the gold chunk through every slot. If accuracy by slot looks like a smile, you have the problem.\n\nHere is the harness I used. `call_llm` and `is_correct` are whatever you already have.\n\n``` python\nimport random\n\ndef build_prompt(chunks, question):\n    docs = \"\\n\\n\".join(\n        f\"<doc id={i+1}>\\n{c}\\n</doc>\" for i, c in enumerate(chunks)\n    )\n    return f\"{docs}\\n\\nQuestion: {question}\\nAnswer using the documents.\"\n\ndef position_sweep(cases, k=10, seed=0):\n    rng = random.Random(seed)\n    hits = [0] * k\n    for case in cases:\n        distractors = rng.sample(case[\"distractors\"], k - 1)\n        for pos in range(k):\n            chunks = distractors[:pos] + [case[\"gold\"]] + distractors[pos:]\n            answer = call_llm(build_prompt(chunks, case[\"question\"]))\n            hits[pos] += is_correct(answer, case[\"expected\"])\n    return [h / len(cases) for h in hits]\n```\n\nTwo details matter. Use real distractors from your own retriever (the chunks it actually returns for that query), not random text. Random text is easy to ignore; near-misses are what confuse the model. And keep the distractor order fixed per case so the only thing changing is the gold position.\n\nMy run: 50 questions from real support tickets, 10 chunks each, 500 calls. The gold chunk in slot 1 got answered correctly 45 times out of 50. Slot 10 got 43. Slots 5 through 7 hovered around 33. Your model and data will give different numbers. What you're looking for is the smile.\n\nMy bug was sitting right at the bottom of it. The refund paragraph had a middling similarity score, so it landed at #7. The warranty paragraph scored higher, landed at #2, and mentioned a number of days. The model grabbed the loud, early, plausible answer.\n\nLLMs ignore the middle because three biases stack up: early tokens get extra attention, tokens near the question get extra attention, and nothing specifically rewards the middle. None of this is a bug in one model. It falls out of how decoder-only transformers are built and trained.\n\n**Primacy from causal attention.** In a decoder, every token can attend to all earlier tokens. The first chunk is visible to every token that follows, so it gets built into the representations of everything downstream. Researchers have also documented \"attention sinks,\" where models park a large share of attention on the very first tokens. The start of the prompt is structurally privileged.\n\n**Recency from distance.** Rotary position embeddings (RoPE), used by most open models, make attention decay somewhat with relative distance. The chunk right before your question is close to where the answer gets generated. The chunk 3,000 tokens back is not.\n\n**Training data shape.** Instructions usually come first. The thing to respond to usually comes last. Long documents where the one crucial fact sits dead center, with the model graded on finding it, are rare in pretraining. The model learned where important stuff usually lives and it is not the middle.\n\nPut those together and the middle of a long prompt is the part with the weakest pull from both ends. Newer long-context models have flattened the curve a lot, and needle-in-a-haystack scores look great. But a needle test uses one distinctive fact in unrelated filler. RAG gives the model ten chunks that all look relevant, and that's the case where the curve comes back.\n\nFix it by sending fewer chunks and putting the best ones at the edges of the context. Those two changes did more for my bot than any embedding swap I tried.\n\nThe cheapest fix is to have less middle. I added a cross-encoder reranker after retrieval and cut from top 10 to top 4. With 4 chunks there is barely a middle to get lost in, and the distractors that caused my bug never made it into the prompt.\n\nRetrieval recall and generation accuracy pull in opposite directions here. More chunks raise the chance the answer is somewhere in the prompt. They also raise the chance the model reads the wrong one. Retrieve wide, rerank, then send narrow.\n\nIf you still need many chunks, don't put them in plain relevance order. That puts #1 at the top, which is great, and buries #2 and #3 in the early middle. Instead, alternate: best first, second best last, and let the weakest ones fall into the center.\n\n``` php\ndef edge_order(chunks):\n    \"\"\"chunks sorted best-first -> best at both ends, worst in the middle\"\"\"\n    front, back = [], []\n    for i, c in enumerate(chunks):\n        (front if i % 2 == 0 else back).append(c)\n    return front + back[::-1]\n\nedge_order([1, 2, 3, 4, 5])  # [1, 3, 5, 4, 2]\n```\n\nLangChain ships the same idea as `LongContextReorder` in `langchain_community.document_transformers`, if you'd rather not own four lines of code.\n\nDocuments first, question last. That puts the instruction in the high-recency zone, right where generation starts. Anthropic's long-context prompting guidance recommends exactly this layout for long inputs. If your template puts the question at the top and then dumps 8K tokens of context under it, flip it. Repeating the question at the top as well costs a few tokens and doesn't hurt.\n\nAsk for the supporting quote first, then the answer:\n\n```\nFirst, copy the exact sentence from the documents that answers the question\ninside <quote> tags, with its doc id. Then answer using only that quote.\nIf no sentence answers it, say so.\n```\n\nThis forces an explicit search step over the whole context before the model commits. It also gives you a free check: if the quoted doc id isn't one your reranker trusted, log it.\n\nNo. A bigger context window lets you fit more chunks. It does not make the model use the middle of them better. Stuffing 50 chunks into a 200K window just gives you a longer, deeper middle. The context window size is a capacity limit. Lost in the middle is a usage problem, and it appears at a few thousand tokens.\n\nAfter the changes (rerank to top 4, edge ordering, question last) I reran the sweep on the same 50 questions. The refund question now gets \"14 days\" every time. More usefully, the gap between the best and worst slot dropped enough that I stopped caring which slot a chunk landed in, and that's the only real sign the problem is gone.\n\nYour LLM ignored the right chunk because of lost in the middle: decoder-only models attend most to the start of the prompt and the text nearest the question, so a correct chunk sitting in position 5 of 10 gets outweighed by plausible distractors at the edges. Retrieval can be perfect and the answer still wrong. Measure it with a position sweep, then cut the chunk count with a reranker, place the strongest chunks first and last, put the question after the documents, and require a quote before the answer. Check where a chunk sits in the prompt, not just whether it's there.\n\n*Written by the developer behind [Preterview](https://preterview.com/en), an interview prep platform.*", "url": "https://wpnews.pro/news/lost-in-the-middle-my-rag-found-the-right-chunk-and-the-llm-ignored-it", "canonical_source": "https://dev.to/ji_ai/lost-in-the-middle-my-rag-found-the-right-chunk-and-the-llm-ignored-it-44ic", "published_at": "2026-09-30 20:59:49+00:00", "updated_at": "2026-09-30 21:16:44.314579+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "natural-language-processing", "ai-agents"], "entities": ["Liu et al.", "Lost in the Middle"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/lost-in-the-middle-my-rag-found-the-right-chunk-and-the-llm-ignored-it", "markdown": "https://wpnews.pro/news/lost-in-the-middle-my-rag-found-the-right-chunk-and-the-llm-ignored-it.md", "text": "https://wpnews.pro/news/lost-in-the-middle-my-rag-found-the-right-chunk-and-the-llm-ignored-it.txt", "jsonld": "https://wpnews.pro/news/lost-in-the-middle-my-rag-found-the-right-chunk-and-the-llm-ignored-it.jsonld"}}