Your 7B Model Fits on a 4090 — Until You Open a 128K Context Window A developer's analysis shows that a 7B model's KV cache, not its weights, is what forces long-context inference onto larger GPUs: Llama 3.1 8B consumes roughly 128 KiB per token in FP16, so a 128K context adds about 16 GB on top of 16 GB of weights and pushes total VRAM past 32 GB. The writeup compares architectures (Qwen2.5 7B needs about half the cache of Llama 3.1 8B at 128K), notes that batch size multiplies cache size and that vLLM's PagedAttention only removes fragmentation, and recommends FP8 cache quantization, GQA-heavy architectures, or a bigger card in that order. Researched October 2026. Architecture specs from public model cards; all numbers below are arithmetic, not my own benchmarks. In my last post https://dev.to/qisuancloud/your-7b-model-doesn-t-need-an-h100-a-practical-gpu-sizing-guide-for-llm-inference I showed that a 7B model's weights need ~16 GB in FP16 — a comfortable fit on a 24 GB RTX 4090. Several readers asked the obvious follow-up: then why does my 7B inference server OOM the moment I enable long context? Because weights are only half the story. The other half is the KV cache , and it scales in a way that surprises almost everyone. For each token in context, the model must store a key and a value vector per layer: KV cache per token = 2 × layers × kv heads × head dim × bytes per value Take Llama 3.1 8B 32 layers, 8 KV heads, 128 head dim, FP16 : 2 × 32 × 8 × 128 × 2 bytes = 131,072 bytes ≈ 128 KiB per token Now multiply by context length: | Context | KV cache Llama 3.1 8B, FP16 | + 16 GB weights → total | |---|---|---| | 8K | ~1 GB | ~17 GB fits 4090 | | 32K | ~4 GB | ~21 GB fits 4090 | | 128K | ~16 GB | ~33 GB needs 48 GB+ | Same model, same weights — the context window alone decides whether you need a $0.35/hr card or a $1.50+/hr one. Prices from my provider comparison https://dev.to/qisuancloud/i-compared-5-gpu-clouds-for-llm-inference-in-2026-heres-what-i-found-4884 . KV head count varies wildly between architectures: | Model | Layers | KV heads | KiB/token FP16 | 128K context cache | |---|---|---|---|---| | Llama 3.1 8B | 32 | 8 | 128 | ~16 GB | | Mistral 7B v0.3 | 32 | 8 | 128 | ~16 GB | | Qwen2.5 7B | 28 | 4 | 56 | ~7 GB | Qwen2.5 7B at 128K needs roughly half the cache of Llama 3.1 8B. Two "7B models," very different GPU bills. Llama 3.1 70B 80 layers, 8 KV heads : 320 KiB/token → 128K context = ~41 GB of KV cache . Add 140 GB of FP16 weights and you're at ~181 GB — that's three 80 GB cards. Even in FP8 70 GB weights , you're at ~111 GB: still two cards minimum. Long context is where 70B deployments quietly become multi-GPU projects. 1. Batch size is a silent multiplier. KV cache scales with batch × sequence length. Serving 8 concurrent users at 8K context costs roughly the same cache as 1 user at 64K. Your "fits on a 4090" math breaks the moment traffic grows. 2. PagedAttention doesn't shrink the cache. vLLM's PagedAttention eliminates fragmentation waste, but the bytes are the bytes — it can't make 33 GB fit in 24 GB. 3. Quantize the cache too. KV cache in FP8 halves these numbers with minimal quality loss on most workloads. It's the same lever as weight quantization, applied to the forgotten half of VRAM. context budget = VRAM − weights − 2 GB overhead ÷ KiB per token ÷ batch size If your planned context exceeds the budget, you have three moves: quantize the cache, pick a GQA-heavy architecture, or rent the bigger card — in that order of cost-effectiveness. I built a free GPU Advisor https://qisuanai.com/advisor that walks through these tradeoffs with four questions and recommends the cheapest fitting setup. No signup. What's the longest context you're actually serving in production — and did the KV cache math surprise you the first time? Curious what caught people off guard.