How to Answer Questions About 500-Page Documents With a 1M-Token Model (Python Map-Reduce Tutorial) A developer published a Python tutorial for building a long-document question-answering pipeline that uses a single 1M-token context window instead of a vector database, embeddings, or chunking library. The approach splits documents into overlapping token windows with tiktoken and applies a map-reduce pattern over an OpenAI-compatible API, with per-request cost tracking based on live model pricing. The writeup argues single-context Q&A beats retrieval when a document fits the window, since the model can see cross-references that retrieval would miss. Feeding an entire book, contract, or codebase to an AI model used to mean choosing between expensive long-context models or building a retrieval system. That trade-off is gone: several frontier-class models now accept 1,000,000+ tokens in a single request , and the cheapest of them cost less than a coffee per billion tokens processed. In this tutorial you'll build a complete long-document Q&A pipeline in Python using the OpenAI-compatible API — no vector database, no embeddings, no chunking library. Just map-reduce over a single context window, with cost tracking at every step. All prices below are live API prices as of October 7, 2026, source: qubax.ai/price-index https://qubax.ai/price-index . First, the data. Here are the cheapest models that accept ≥1M-token contexts right now, with their price per million tokens and what the same tokens cost on OpenRouter retail: | Model | Context | Input $/M | Output $/M | OpenRouter input $/M | |---|---|---|---|---| | MiniMax M2 | 1,000,000 | $0.0038 | $0.0151 | $0.26 | | DeepSeek V4.1 Flash 0731 | 1,048,576 | $0.0156 | $0.2169 | $0.18 | | GLM 5.2 | 1,048,576 | $0.0197 | $0.0788 | $0.09 | | Grok 4.20 Beta | 2,000,000 | $0.0626 | $0.2505 | $1.25 | | Claude Opus 4.6 | 1,000,000 | $3.5577 | $14.2308 | $5.00 | Two things worth noticing: For this tutorial we'll use DeepSeek V4.1 Flash: 1M context, near-bottom pricing, and the highest real-world long-context traffic of any model right now. Long-document Q&A has two standard approaches: RAG is still right when your document exceeds the context window or when latency matters more than recall. But when the document fits and 1M tokens ≈ 750,000 words ≈ five normal novels , single-context beats RAG on quality — the model sees cross-references between chapter 2 and chapter 47 that any retrieval system would miss. The map-reduce pattern we'll build: This is more reliable than a single mega-prompt for two reasons: attention degrades over very long inputs models recall the beginning and end better than the middle — the "lost in the middle" problem documented by Stanford researchers https://arxiv.org/abs/2307.03172 , and per-segment extraction keeps each request's output small and cheap. python import os from openai import OpenAI client = OpenAI api key=os.environ "QUBAX API KEY" , base url="https://api.qubax.ai/v1", MODEL = "deepseek-v4-flash-0731" Live prices Oct 7, 2026 : $0.0156/M input, $0.2169/M output PRICE IN = 0.0156 / 1 000 000 PRICE OUT = 0.2169 / 1 000 000 def cost of usage - float: return usage.prompt tokens PRICE IN + usage.completion tokens PRICE OUT Any OpenAI SDK works see the official OpenAI Python SDK docs https://github.com/openai/openai-python — only the base url changes . Token counting uses tiktoken https://github.com/openai/tiktoken , OpenAI's o200k base tokenizer, which is a close match for how modern models segment text; exact billing always comes from the API's usage field. Don't split by characters — split by tokens, with overlap so no fact is cut in half: python import tiktoken enc = tiktoken.get encoding "o200k base" def segment text: str, max tokens: int = 300 000, overlap: int = 2 000 - list str : """Split text into overlapping token windows.""" ids = enc.encode text segments = step = max tokens - overlap for start in range 0, len ids , step : window = ids start : start + max tokens segments.append enc.decode window if start + max tokens = len ids : break return segments A 500-page document is roughly 375K tokens, so two segments of 300K with 2K overlap covers it. A 1M-token corpus becomes four segments. php def extract from segment question: str, segment: str, index: int - str: resp = client.chat.completions.create model=MODEL, messages= {"role": "system", "content": "You are a research assistant. Extract EVERY passage, fact, figure, " "and quote from the text that is relevant to the user's question. " "If nothing is relevant, reply exactly: NO RELEVANT CONTENT. " "Quote passages verbatim where possible."}, {"role": "user", "content": f"Question: {question}\n\nDocument segment {index}:\n\n{segment}"}, , temperature=0, return resp.choices 0 .message.content php def answer question question: str, document: str - dict: segs = segment document findings = total cost = 0.0 for i, seg in enumerate segs : out = extract from segment question, seg, i u = out each call returns text; usage captured below if out.strip = "NO RELEVANT CONTENT": findings.append f" Segment {i+1} \n{out}" if not findings: return {"answer": "The document does not address this question.", "cost": total cost} combined = "\n\n---\n\n".join findings reduce resp = client.chat.completions.create model=MODEL, messages= {"role": "system", "content": "Synthesize a single, complete answer to the question from the " "extracted findings below. Cite segments as S1 , S2 . " "Resolve contradictions explicitly."}, {"role": "user", "content": f"Question: {question}\n\nFindings:\n{combined}"}, , temperature=0, track cost across all calls total cost = len enc.encode document PRICE IN 1.05 map inputs + 5% overlap total cost += reduce resp.usage.prompt tokens PRICE IN total cost += reduce resp.usage.completion tokens PRICE OUT return { "answer": reduce resp.choices 0 .message.content, "segments searched": len segs , "estimated cost usd": round total cost, 6 , } document = open "quarterly report.pdf.txt" .read 500 pages ≈ 375K tokens result = answer question "What are all the contingent liabilities mentioned, and in which sections?", document, print result "answer" print f"Searched {result 'segments searched' } segments · cost ≈ ${result 'estimated cost usd' }" For a 375K-token document on DeepSeek V4.1 Flash, the map pass costs about $0.006 . The whole query — map plus reduce — lands under one cent . The same query against Claude Opus 4.6's 1M context would run about $1.40 in input alone, and against MiniMax M2 about $0.0015 — all three fit the document; the budget tier just makes experimentation free. tiktoken encoding of a large document takes a few seconds. Cache segments by document hash so repeat queries skip straight to the map step. extract from segment calls concurrently e.g. with asyncio + AsyncOpenAI . Four segments in parallel cuts wall-clock time ~4×. usage per call; sum it rather than estimating from character counts. Our estimate above is deliberately conservative 105% of raw tokens . Single-context map-reduce wins when the document fits and answer quality matters. Switch to RAG when: A hybrid works well too: RAG for the fast path, and fall back to full-context map-reduce when retrieval confidence is low. Roughly 750,000 words, or about 1,500–2,500 PDF pages depending on density. DeepSeek V4.1 Flash accepts 1,048,576 tokens and Grok 4.20 Beta accepts 2,000,000. Not if the document fits in the context window. Vector search adds infrastructure and can miss cross-references. Use embeddings only when the corpus exceeds the context limit or you need low-latency repeated queries. As of October 7, 2026, MiniMax M2 at $0.0038 per million input tokens and $0.0151 per million output tokens — roughly 936× cheaper than Claude Opus 4.6's 1M context on input. DeepSeek V4.1 Flash and GLM 5.2 are close behind with 1,048,576-token windows. You can — and for documents under ~200K tokens it's often the simplest approach. Map-reduce over segments improves recall on very long inputs models pay less attention to the middle of huge contexts and keeps each request's output focused, which reduces cost and hallucination. Yes. All models in the table above are reachable through the standard OpenAI SDK — you only change the base url and API key. The tutorial's code runs unchanged. Check live prices for any model at the Qubax price index https://qubax.ai/price-index — updated continuously from real API traffic. Or browse all 399+ models https://qubax.ai/models to find the right context window and price for your workload. Originally published at qubax.ai https://qubax.ai/blog/2026-10-07-build-long-document-qa-map-reduce-tutorial . Qubax gives you every top AI model with one API key, cheaper than OpenRouter.