cd /news/large-language-models/lost-in-the-middle-my-rag-found-the-… · home › topics › large-language-models › article
[ARTICLE · art-142815] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Lost in the Middle: My RAG Found the Right Chunk and the LLM Ignored It

A developer building a RAG-based support bot traced a wrong answer to the "lost in the middle" effect, in which an LLM ignored the correct policy chunk placed at position #7 of a 10-chunk prompt and instead used a higher-scoring warranty paragraph near the start. Running a position sweep over 50 real support questions with 10 chunks each (500 calls), the developer measured 45/50 correct answers when the gold chunk was first, 43/50 when last, and roughly 33/50 in slots 5 through 7, confirming the U-shaped accuracy curve reported in the 2023 Liu et al. paper. The writeup attributes the effect to stacked biases in decoder-only transformers, including primacy from causal attention and attention sinks on early tokens.

by read7 min views2 publishedSep 30, 2026

The retrieval log said I was right. The answer said I was wrong.

A user asked my support bot how long refund requests stay open. The retriever pulled 10 chunks, and chunk #7 was the exact paragraph from the policy doc: "Refund requests expire after 14 days." The model answered "30 days," which came from chunk #2, a paragraph about warranty claims.

I spent an evening blaming the embeddings. The embeddings were fine. The problem was lost in the middle: LLMs use information at the start and end of a long prompt far better than information buried in the middle. My retriever did its job. Then I put the answer in the one place the model barely reads.

The lost in the middle effect is the drop in answer accuracy when the relevant information sits in the middle of a long prompt instead of near the beginning or end. The name comes from the 2023 paper "Lost in the Middle: How Language Models Use Long Contexts" by Liu et al., which ran multi-document question answering with the answer-bearing document moved through every position.

The shape they found was a U. High accuracy when the gold document came first, a bit lower when it came last, and a sag in between. In some of their settings, a model with the answer sitting in the middle of its context did worse than the same model given no documents at all.

Read that again. Adding the correct document, in the wrong spot, made things worse than adding nothing. The nine distractors around it did more damage than the one right answer did good.

Run a position sweep. Take questions where you know which chunk holds the answer, keep the other chunks fixed, and move the gold chunk through every slot. If accuracy by slot looks like a smile, you have the problem.

Here is the harness I used. call_llm and is_correct are whatever you already have.

import random

def build_prompt(chunks, question):
    docs = "\n\n".join(
        f"<doc id={i+1}>\n{c}\n</doc>" for i, c in enumerate(chunks)
    )
    return f"{docs}\n\nQuestion: {question}\nAnswer using the documents."

def position_sweep(cases, k=10, seed=0):
    rng = random.Random(seed)
    hits = [0] * k
    for case in cases:
        distractors = rng.sample(case["distractors"], k - 1)
        for pos in range(k):
            chunks = distractors[:pos] + [case["gold"]] + distractors[pos:]
            answer = call_llm(build_prompt(chunks, case["question"]))
            hits[pos] += is_correct(answer, case["expected"])
    return [h / len(cases) for h in hits]

Two details matter. Use real distractors from your own retriever (the chunks it actually returns for that query), not random text. Random text is easy to ignore; near-misses are what confuse the model. And keep the distractor order fixed per case so the only thing changing is the gold position.

My run: 50 questions from real support tickets, 10 chunks each, 500 calls. The gold chunk in slot 1 got answered correctly 45 times out of 50. Slot 10 got 43. Slots 5 through 7 hovered around 33. Your model and data will give different numbers. What you're looking for is the smile.

My bug was sitting right at the bottom of it. The refund paragraph had a middling similarity score, so it landed at #7. The warranty paragraph scored higher, landed at #2, and mentioned a number of days. The model grabbed the loud, early, plausible answer.

LLMs ignore the middle because three biases stack up: early tokens get extra attention, tokens near the question get extra attention, and nothing specifically rewards the middle. None of this is a bug in one model. It falls out of how decoder-only transformers are built and trained.

Primacy from causal attention. In a decoder, every token can attend to all earlier tokens. The first chunk is visible to every token that follows, so it gets built into the representations of everything downstream. Researchers have also documented "attention sinks," where models park a large share of attention on the very first tokens. The start of the prompt is structurally privileged.

Recency from distance. Rotary position embeddings (RoPE), used by most open models, make attention decay somewhat with relative distance. The chunk right before your question is close to where the answer gets generated. The chunk 3,000 tokens back is not.

Training data shape. Instructions usually come first. The thing to respond to usually comes last. Long documents where the one crucial fact sits dead center, with the model graded on finding it, are rare in pretraining. The model learned where important stuff usually lives and it is not the middle.

Put those together and the middle of a long prompt is the part with the weakest pull from both ends. Newer long-context models have flattened the curve a lot, and needle-in-a-haystack scores look great. But a needle test uses one distinctive fact in unrelated filler. RAG gives the model ten chunks that all look relevant, and that's the case where the curve comes back.

Fix it by sending fewer chunks and putting the best ones at the edges of the context. Those two changes did more for my bot than any embedding swap I tried.

The cheapest fix is to have less middle. I added a cross-encoder reranker after retrieval and cut from top 10 to top 4. With 4 chunks there is barely a middle to get lost in, and the distractors that caused my bug never made it into the prompt.

Retrieval recall and generation accuracy pull in opposite directions here. More chunks raise the chance the answer is somewhere in the prompt. They also raise the chance the model reads the wrong one. Retrieve wide, rerank, then send narrow.

If you still need many chunks, don't put them in plain relevance order. That puts #1 at the top, which is great, and buries #2 and #3 in the early middle. Instead, alternate: best first, second best last, and let the weakest ones fall into the center.

def edge_order(chunks):
    """chunks sorted best-first -> best at both ends, worst in the middle"""
    front, back = [], []
    for i, c in enumerate(chunks):
        (front if i % 2 == 0 else back).append(c)
    return front + back[::-1]

edge_order([1, 2, 3, 4, 5])  # [1, 3, 5, 4, 2]

LangChain ships the same idea as LongContextReorder in langchain_community.document_transformers, if you'd rather not own four lines of code.

Documents first, question last. That puts the instruction in the high-recency zone, right where generation starts. Anthropic's long-context prompting guidance recommends exactly this layout for long inputs. If your template puts the question at the top and then dumps 8K tokens of context under it, flip it. Repeating the question at the top as well costs a few tokens and doesn't hurt.

Ask for the supporting quote first, then the answer:

First, copy the exact sentence from the documents that answers the question
inside <quote> tags, with its doc id. Then answer using only that quote.
If no sentence answers it, say so.

This forces an explicit search step over the whole context before the model commits. It also gives you a free check: if the quoted doc id isn't one your reranker trusted, log it.

No. A bigger context window lets you fit more chunks. It does not make the model use the middle of them better. Stuffing 50 chunks into a 200K window just gives you a longer, deeper middle. The context window size is a capacity limit. Lost in the middle is a usage problem, and it appears at a few thousand tokens.

After the changes (rerank to top 4, edge ordering, question last) I reran the sweep on the same 50 questions. The refund question now gets "14 days" every time. More usefully, the gap between the best and worst slot dropped enough that I stopped caring which slot a chunk landed in, and that's the only real sign the problem is gone.

Your LLM ignored the right chunk because of lost in the middle: decoder-only models attend most to the start of the prompt and the text nearest the question, so a correct chunk sitting in position 5 of 10 gets outweighed by plausible distractors at the edges. Retrieval can be perfect and the answer still wrong. Measure it with a position sweep, then cut the chunk count with a reranker, place the strongest chunks first and last, put the question after the documents, and require a quote before the answer. Check where a chunk sits in the prompt, not just whether it's there.

Written by the developer behind Preterview, an interview prep platform.

── more in #large-language-models 4 stories · sorted by recency
── more on @liu et al. 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/lost-in-the-middle-m…] indexed:0 read:7min 2026-09-30 · —