cd /news/large-language-models/your-7b-model-fits-on-a-4090-until-y… · home › topics › large-language-models › article
[ARTICLE · art-148988] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Your 7B Model Fits on a 4090 — Until You Open a 128K Context Window

A developer's analysis shows that a 7B model's KV cache, not its weights, is what forces long-context inference onto larger GPUs: Llama 3.1 8B consumes roughly 128 KiB per token in FP16, so a 128K context adds about 16 GB on top of 16 GB of weights and pushes total VRAM past 32 GB. The writeup compares architectures (Qwen2.5 7B needs about half the cache of Llama 3.1 8B at 128K), notes that batch size multiplies cache size and that vLLM's PagedAttention only removes fragmentation, and recommends FP8 cache quantization, GQA-heavy architectures, or a bigger card in that order.

by read3 min views1 publishedOct 11, 2026

Researched October 2026. Architecture specs from public model cards; all numbers below are arithmetic, not my own benchmarks.

In my last post I showed that a 7B model's weights need ~16 GB in FP16 — a comfortable fit on a 24 GB RTX 4090. Several readers asked the obvious follow-up: then why does my 7B inference server OOM the moment I enable long context?

Because weights are only half the story. The other half is the KV cache, and it scales in a way that surprises almost everyone.

For each token in context, the model must store a key and a value vector per layer:

KV cache per token = 2 × layers × kv_heads × head_dim × bytes_per_value

Take Llama 3.1 8B (32 layers, 8 KV heads, 128 head dim, FP16):

2 × 32 × 8 × 128 × 2 bytes = 131,072 bytes ≈ 128 KiB per token

Now multiply by context length:

Context KV cache (Llama 3.1 8B, FP16) + 16 GB weights → total
8K ~1 GB ~17 GB fits 4090
32K ~4 GB ~21 GB fits 4090
128K ~16 GB ~33 GB needs 48 GB+

Same model, same weights — the context window alone decides whether you need a $0.35/hr card or a $1.50+/hr one. (Prices from my provider comparison.)

KV head count varies wildly between architectures:

Model Layers KV heads KiB/token (FP16) 128K context cache
Llama 3.1 8B 32 8 128 ~16 GB
Mistral 7B v0.3 32 8 128 ~16 GB
Qwen2.5 7B 28 4 56 ~7 GB

Qwen2.5 7B at 128K needs roughly half the cache of Llama 3.1 8B. Two "7B models," very different GPU bills.

Llama 3.1 70B (80 layers, 8 KV heads): 320 KiB/token → 128K context = ~41 GB of KV cache. Add 140 GB of FP16 weights and you're at ~181 GB — that's three 80 GB cards. Even in FP8 (70 GB weights), you're at ~111 GB: still two cards minimum. Long context is where 70B deployments quietly become multi-GPU projects.

1. Batch size is a silent multiplier. KV cache scales with batch × sequence length. Serving 8 concurrent users at 8K context costs roughly the same cache as 1 user at 64K. Your "fits on a 4090" math breaks the moment traffic grows.

2. PagedAttention doesn't shrink the cache. vLLM's PagedAttention eliminates fragmentation waste, but the bytes are the bytes — it can't make 33 GB fit in 24 GB.

3. Quantize the cache too. KV cache in FP8 halves these numbers with minimal quality loss on most workloads. It's the same lever as weight quantization, applied to the forgotten half of VRAM.

context_budget = (VRAM − weights − 2 GB overhead) ÷ KiB_per_token ÷ batch_size

If your planned context exceeds the budget, you have three moves: quantize the cache, pick a GQA-heavy architecture, or rent the bigger card — in that order of cost-effectiveness.

I built a free GPU Advisor that walks through these tradeoffs with four questions and recommends the cheapest fitting setup. No signup.

What's the longest context you're actually serving in production — and did the KV cache math surprise you the first time? Curious what caught people off guard.

── more in #large-language-models 4 stories · sorted by recency
── more on @llama 3.1 8b 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/your-7b-model-fits-o…] indexed:0 read:3min 2026-10-11 · —