# Your 7B Model Fits on a 4090 — Until You Open a 128K Context Window

> Source: <https://dev.to/qisuancloud/your-7b-model-fits-on-a-4090-until-you-open-a-128k-context-window-31ag>
> Published: 2026-10-11 02:41:11+00:00

*Researched October 2026. Architecture specs from public model cards; all numbers below are arithmetic, not my own benchmarks.*

In my [last post](https://dev.to/qisuancloud/your-7b-model-doesn-t-need-an-h100-a-practical-gpu-sizing-guide-for-llm-inference) I showed that a 7B model's weights need ~16 GB in FP16 — a comfortable fit on a 24 GB RTX 4090. Several readers asked the obvious follow-up: *then why does my 7B inference server OOM the moment I enable long context?*

Because weights are only half the story. The other half is the **KV cache**, and it scales in a way that surprises almost everyone.

For each token in context, the model must store a key and a value vector per layer:

```
KV cache per token = 2 × layers × kv_heads × head_dim × bytes_per_value
```

Take **Llama 3.1 8B** (32 layers, 8 KV heads, 128 head dim, FP16):

```
2 × 32 × 8 × 128 × 2 bytes = 131,072 bytes ≈ 128 KiB per token
```

Now multiply by context length:

| Context | KV cache (Llama 3.1 8B, FP16) | + 16 GB weights → total | 
|---|---|---|
| 8K | ~1 GB | ~17 GB fits 4090 | 
| 32K | ~4 GB | ~21 GB fits 4090 | 
| 128K | ~16 GB | ~33 GB needs 48 GB+ | 

Same model, same weights — the context window alone decides whether you need a $0.35/hr card or a $1.50+/hr one. (Prices from my [provider comparison](https://dev.to/qisuancloud/i-compared-5-gpu-clouds-for-llm-inference-in-2026-heres-what-i-found-4884).)

KV head count varies wildly between architectures:

| Model | Layers | KV heads | KiB/token (FP16) | 128K context cache | 
|---|---|---|---|---|
| Llama 3.1 8B | 32 | 8 | 128 | ~16 GB | 
| Mistral 7B v0.3 | 32 | 8 | 128 | ~16 GB | 
| Qwen2.5 7B | 28 | 4 | 56 | ~7 GB | 

Qwen2.5 7B at 128K needs roughly **half** the cache of Llama 3.1 8B. Two "7B models," very different GPU bills.

**Llama 3.1 70B** (80 layers, 8 KV heads): 320 KiB/token → 128K context = **~41 GB of KV cache**. Add 140 GB of FP16 weights and you're at ~181 GB — that's three 80 GB cards. Even in FP8 (70 GB weights), you're at ~111 GB: still two cards minimum. Long context is where 70B deployments quietly become multi-GPU projects.

**1. Batch size is a silent multiplier.** KV cache scales with batch × sequence length. Serving 8 concurrent users at 8K context costs roughly the same cache as 1 user at 64K. Your "fits on a 4090" math breaks the moment traffic grows.

**2. PagedAttention doesn't shrink the cache.** vLLM's PagedAttention eliminates fragmentation waste, but the bytes are the bytes — it can't make 33 GB fit in 24 GB.

**3. Quantize the cache too.** KV cache in FP8 halves these numbers with minimal quality loss on most workloads. It's the same lever as weight quantization, applied to the forgotten half of VRAM.

```
context_budget = (VRAM − weights − 2 GB overhead) ÷ KiB_per_token ÷ batch_size
```

If your planned context exceeds the budget, you have three moves: quantize the cache, pick a GQA-heavy architecture, or rent the bigger card — in that order of cost-effectiveness.

I built a [free GPU Advisor](https://qisuanai.com/advisor) that walks through these tradeoffs with four questions and recommends the cheapest fitting setup. No signup.

What's the longest context you're actually serving in production — and did the KV cache math surprise you the first time? Curious what caught people off guard.
