{"slug": "your-7b-model-fits-on-a-4090-until-you-open-a-128k-context-window", "title": "Your 7B Model Fits on a 4090 — Until You Open a 128K Context Window", "summary": "A developer's analysis shows that a 7B model's KV cache, not its weights, is what forces long-context inference onto larger GPUs: Llama 3.1 8B consumes roughly 128 KiB per token in FP16, so a 128K context adds about 16 GB on top of 16 GB of weights and pushes total VRAM past 32 GB. The writeup compares architectures (Qwen2.5 7B needs about half the cache of Llama 3.1 8B at 128K), notes that batch size multiplies cache size and that vLLM's PagedAttention only removes fragmentation, and recommends FP8 cache quantization, GQA-heavy architectures, or a bigger card in that order.", "body_md": "*Researched October 2026. Architecture specs from public model cards; all numbers below are arithmetic, not my own benchmarks.*\n\nIn my [last post](https://dev.to/qisuancloud/your-7b-model-doesn-t-need-an-h100-a-practical-gpu-sizing-guide-for-llm-inference) I showed that a 7B model's weights need ~16 GB in FP16 — a comfortable fit on a 24 GB RTX 4090. Several readers asked the obvious follow-up: *then why does my 7B inference server OOM the moment I enable long context?*\n\nBecause weights are only half the story. The other half is the **KV cache**, and it scales in a way that surprises almost everyone.\n\nFor each token in context, the model must store a key and a value vector per layer:\n\n```\nKV cache per token = 2 × layers × kv_heads × head_dim × bytes_per_value\n```\n\nTake **Llama 3.1 8B** (32 layers, 8 KV heads, 128 head dim, FP16):\n\n```\n2 × 32 × 8 × 128 × 2 bytes = 131,072 bytes ≈ 128 KiB per token\n```\n\nNow multiply by context length:\n\n| Context | KV cache (Llama 3.1 8B, FP16) | + 16 GB weights → total | \n|---|---|---|\n| 8K | ~1 GB | ~17 GB fits 4090 | \n| 32K | ~4 GB | ~21 GB fits 4090 | \n| 128K | ~16 GB | ~33 GB needs 48 GB+ | \n\nSame model, same weights — the context window alone decides whether you need a $0.35/hr card or a $1.50+/hr one. (Prices from my [provider comparison](https://dev.to/qisuancloud/i-compared-5-gpu-clouds-for-llm-inference-in-2026-heres-what-i-found-4884).)\n\nKV head count varies wildly between architectures:\n\n| Model | Layers | KV heads | KiB/token (FP16) | 128K context cache | \n|---|---|---|---|---|\n| Llama 3.1 8B | 32 | 8 | 128 | ~16 GB | \n| Mistral 7B v0.3 | 32 | 8 | 128 | ~16 GB | \n| Qwen2.5 7B | 28 | 4 | 56 | ~7 GB | \n\nQwen2.5 7B at 128K needs roughly **half** the cache of Llama 3.1 8B. Two \"7B models,\" very different GPU bills.\n\n**Llama 3.1 70B** (80 layers, 8 KV heads): 320 KiB/token → 128K context = **~41 GB of KV cache**. Add 140 GB of FP16 weights and you're at ~181 GB — that's three 80 GB cards. Even in FP8 (70 GB weights), you're at ~111 GB: still two cards minimum. Long context is where 70B deployments quietly become multi-GPU projects.\n\n**1. Batch size is a silent multiplier.** KV cache scales with batch × sequence length. Serving 8 concurrent users at 8K context costs roughly the same cache as 1 user at 64K. Your \"fits on a 4090\" math breaks the moment traffic grows.\n\n**2. PagedAttention doesn't shrink the cache.** vLLM's PagedAttention eliminates fragmentation waste, but the bytes are the bytes — it can't make 33 GB fit in 24 GB.\n\n**3. Quantize the cache too.** KV cache in FP8 halves these numbers with minimal quality loss on most workloads. It's the same lever as weight quantization, applied to the forgotten half of VRAM.\n\n```\ncontext_budget = (VRAM − weights − 2 GB overhead) ÷ KiB_per_token ÷ batch_size\n```\n\nIf your planned context exceeds the budget, you have three moves: quantize the cache, pick a GQA-heavy architecture, or rent the bigger card — in that order of cost-effectiveness.\n\nI built a [free GPU Advisor](https://qisuanai.com/advisor) that walks through these tradeoffs with four questions and recommends the cheapest fitting setup. No signup.\n\nWhat's the longest context you're actually serving in production — and did the KV cache math surprise you the first time? Curious what caught people off guard.", "url": "https://wpnews.pro/news/your-7b-model-fits-on-a-4090-until-you-open-a-128k-context-window", "canonical_source": "https://dev.to/qisuancloud/your-7b-model-fits-on-a-4090-until-you-open-a-128k-context-window-31ag", "published_at": "2026-10-11 02:41:11+00:00", "updated_at": "2026-10-11 02:49:55.555042+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "mlops", "ai-tools"], "entities": ["Llama 3.1 8B", "Llama 3.1 70B", "Mistral 7B v0.3", "Qwen2.5 7B", "RTX 4090", "vLLM", "PagedAttention", "GPU Advisor"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/your-7b-model-fits-on-a-4090-until-you-open-a-128k-context-window", "markdown": "https://wpnews.pro/news/your-7b-model-fits-on-a-4090-until-you-open-a-128k-context-window.md", "text": "https://wpnews.pro/news/your-7b-model-fits-on-a-4090-until-you-open-a-128k-context-window.txt", "jsonld": "https://wpnews.pro/news/your-7b-model-fits-on-a-4090-until-you-open-a-128k-context-window.jsonld"}}