# How to Answer Questions About 500-Page Documents With a 1M-Token Model (Python Map-Reduce Tutorial)

> Source: <https://dev.to/qubax_ai/how-to-answer-questions-about-500-page-documents-with-a-1m-token-model-python-map-reduce-tutorial-2haa>
> Published: 2026-10-08 21:34:52+00:00

Feeding an entire book, contract, or codebase to an AI model used to mean choosing between expensive long-context models or building a retrieval system. That trade-off is gone: several frontier-class models now accept **1,000,000+ tokens in a single request**, and the cheapest of them cost less than a coffee per billion tokens processed.

In this tutorial you'll build a complete long-document Q&A pipeline in Python using the OpenAI-compatible API — no vector database, no embeddings, no chunking library. Just map-reduce over a single context window, with cost tracking at every step.

All prices below are live API prices **as of October 7, 2026, source: [qubax.ai/price-index](https://qubax.ai/price-index)**.

First, the data. Here are the cheapest models that accept ≥1M-token contexts right now, with their price per million tokens and what the same tokens cost on OpenRouter retail:

| Model | Context | Input $/M | Output $/M | OpenRouter input $/M | 
|---|---|---|---|---|
| MiniMax M2 | 1,000,000 | $0.0038 | $0.0151 | $0.26 | 
| DeepSeek V4.1 Flash (0731) | 1,048,576 | $0.0156 | $0.2169 | $0.18 | 
| GLM 5.2 | 1,048,576 | $0.0197 | $0.0788 | $0.09 | 
| Grok 4.20 Beta | 2,000,000 | $0.0626 | $0.2505 | $1.25 | 
| Claude Opus 4.6 | 1,000,000 | $3.5577 | $14.2308 | $5.00 | 

Two things worth noticing:

For this tutorial we'll use DeepSeek V4.1 Flash: 1M context, near-bottom pricing, and the highest real-world long-context traffic of any model right now.

Long-document Q&A has two standard approaches:

RAG is still right when your document exceeds the context window or when latency matters more than recall. But when the document *fits* (and 1M tokens ≈ 750,000 words ≈ five normal novels), single-context beats RAG on quality — the model sees cross-references between chapter 2 and chapter 47 that any retrieval system would miss.

The map-reduce pattern we'll build:

This is more reliable than a single mega-prompt for two reasons: attention degrades over very long inputs (models recall the beginning and end better than the middle — the ["lost in the middle" problem documented by Stanford researchers](https://arxiv.org/abs/2307.03172)), and per-segment extraction keeps each request's output small and cheap.

``` python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["QUBAX_API_KEY"],
    base_url="https://api.qubax.ai/v1",
)

MODEL = "deepseek-v4-flash-0731"

# Live prices (Oct 7, 2026): $0.0156/M input, $0.2169/M output
PRICE_IN = 0.0156 / 1_000_000
PRICE_OUT = 0.2169 / 1_000_000

def cost_of(usage) -> float:
    return usage.prompt_tokens * PRICE_IN + usage.completion_tokens * PRICE_OUT
```

Any OpenAI SDK works (see the [official OpenAI Python SDK docs](https://github.com/openai/openai-python) — only the `base_url` changes). Token counting uses [tiktoken](https://github.com/openai/tiktoken), OpenAI's `o200k_base` tokenizer, which is a close match for how modern models segment text; exact billing always comes from the API's `usage` field.

Don't split by characters — split by tokens, with overlap so no fact is cut in half:

``` python
import tiktoken

enc = tiktoken.get_encoding("o200k_base")

def segment(text: str, max_tokens: int = 300_000, overlap: int = 2_000) -> list[str]:
    """Split text into overlapping token windows."""
    ids = enc.encode(text)
    segments = []
    step = max_tokens - overlap
    for start in range(0, len(ids), step):
        window = ids[start : start + max_tokens]
        segments.append(enc.decode(window))
        if start + max_tokens >= len(ids):
            break
    return segments
```

A 500-page document is roughly 375K tokens, so two segments of 300K with 2K overlap covers it. A 1M-token corpus becomes four segments.

``` php
def extract_from_segment(question: str, segment: str, index: int) -> str:
    resp = client.chat.completions.create(
        model=MODEL,
        messages=[
            {"role": "system", "content":
                "You are a research assistant. Extract EVERY passage, fact, figure, "
                "and quote from the text that is relevant to the user's question. "
                "If nothing is relevant, reply exactly: NO_RELEVANT_CONTENT. "
                "Quote passages verbatim where possible."},
            {"role": "user", "content":
                f"Question: {question}\n\nDocument segment {index}:\n\n{segment}"},
        ],
        temperature=0,
    )
    return resp.choices[0].message.content
php
def answer_question(question: str, document: str) -> dict:
    segs = segment(document)
    findings = []
    total_cost = 0.0

    for i, seg in enumerate(segs):
        out = extract_from_segment(question, seg, i)
        u = out  # each call returns text; usage captured below
        if out.strip() != "NO_RELEVANT_CONTENT":
            findings.append(f"[Segment {i+1}]\n{out}")

    if not findings:
        return {"answer": "The document does not address this question.", "cost": total_cost}

    combined = "\n\n---\n\n".join(findings)
    reduce_resp = client.chat.completions.create(
        model=MODEL,
        messages=[
            {"role": "system", "content":
                "Synthesize a single, complete answer to the question from the "
                "extracted findings below. Cite segments as [S1], [S2]. "
                "Resolve contradictions explicitly."},
            {"role": "user", "content": f"Question: {question}\n\nFindings:\n{combined}"},
        ],
        temperature=0,
    )

    # track cost across all calls
    total_cost = len(enc.encode(document)) * PRICE_IN * 1.05  # map inputs + 5% overlap
    total_cost += reduce_resp.usage.prompt_tokens * PRICE_IN
    total_cost += reduce_resp.usage.completion_tokens * PRICE_OUT

    return {
        "answer": reduce_resp.choices[0].message.content,
        "segments_searched": len(segs),
        "estimated_cost_usd": round(total_cost, 6),
    }
document = open("quarterly_report.pdf.txt").read()  # 500 pages ≈ 375K tokens

result = answer_question(
    "What are all the contingent liabilities mentioned, and in which sections?",
    document,
)
print(result["answer"])
print(f"Searched {result['segments_searched']} segments · cost ≈ ${result['estimated_cost_usd']}")
```

For a 375K-token document on DeepSeek V4.1 Flash, the map pass costs about **$0.006**. The whole query — map plus reduce — lands under **one cent**. The same query against Claude Opus 4.6's 1M context would run about **$1.40** in input alone, and against MiniMax M2 about **$0.0015** — all three fit the document; the budget tier just makes experimentation free.

`tiktoken` encoding of a large document takes a few seconds. Cache segments by document hash so repeat queries skip straight to the map step.`extract_from_segment` calls concurrently (e.g. with `asyncio` + `AsyncOpenAI`). Four segments in parallel cuts wall-clock time ~4×.` usage` per call; sum it rather than estimating from character counts. Our estimate above is deliberately conservative (105% of raw tokens).
Single-context map-reduce wins when the document fits and answer quality matters. Switch to RAG when:

A hybrid works well too: RAG for the fast path, and fall back to full-context map-reduce when retrieval confidence is low.

Roughly 750,000 words, or about 1,500–2,500 PDF pages depending on density. DeepSeek V4.1 Flash accepts 1,048,576 tokens and Grok 4.20 Beta accepts 2,000,000.

Not if the document fits in the context window. Vector search adds infrastructure and can miss cross-references. Use embeddings only when the corpus exceeds the context limit or you need low-latency repeated queries.

As of October 7, 2026, MiniMax M2 at $0.0038 per million input tokens and $0.0151 per million output tokens — roughly 936× cheaper than Claude Opus 4.6's 1M context on input. DeepSeek V4.1 Flash and GLM 5.2 are close behind with 1,048,576-token windows.

You can — and for documents under ~200K tokens it's often the simplest approach. Map-reduce over segments improves recall on very long inputs (models pay less attention to the middle of huge contexts) and keeps each request's output focused, which reduces cost and hallucination.

Yes. All models in the table above are reachable through the standard OpenAI SDK — you only change the `base_url` and API key. The tutorial's code runs unchanged.

**Check live prices for any model at the [Qubax price index](https://qubax.ai/price-index)** — updated continuously from real API traffic. Or [browse all 399+ models](https://qubax.ai/models) to find the right context window and price for your workload.

*Originally published at [qubax.ai](https://qubax.ai/blog/2026-10-07-build-long-document-qa-map-reduce-tutorial). Qubax gives you every top AI model with one API key, cheaper than OpenRouter.*
