cd /news/large-language-models/will-it-fit-on-my-gpu-building-an-ll… · home › topics › large-language-models › article
[ARTICLE · art-148783] src=discuss.huggingface.co ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Will it fit on my GPU ? Building an LLM VRAM calculator that models each architecture's real KV cache

StudioTV released a free LLM VRAM Calculator that models each architecture's real KV cache layer by layer instead of assuming full attention, correcting naive estimates that can be off by an order of magnitude. The tool reports that Gemma 4 31B's real 128K-token cache is 10.8 GiB versus 120 GiB for a full-attention formula, Qwen3.6 35B-A3B's is 2.5 GiB, and DeepSeek R1's latent attention cache is 8.6 GiB against a 610 GiB multi-head prediction. It covers 60 GPUs, weight formats from BF16 to GGUF Q3_K_M, vLLM, SGLang, TensorRT-LLM and llama.cpp, and hourly-refreshed rental prices from RunPod, Vast.ai, Verda and Azure.

read4 min views4 publishedOct 10, 2026

Every time a new open model comes out, the same question shows up in every thread: will it run on my card? With 24 GB? With two of them? In 4-bit? At what context length?

Most answers come from a back-of-the-envelope formula: parameters × bytes per weight, plus a KV cache computed as if every layer were a classic full-attention layer. For a Llama-style dense model that is fine. For most models released in the last year, it is wrong, sometimes by an order of magnitude.

I built StudioTV’s LLM VRAM Calculator to answer that question properly, for every recent architecture, and to go one step further: what it costs to rent the hardware, and the exact command to launch it. It is free, has no sign-up, and works for any model on Hugging Face.

The KV cache, the memory that grows with context length, depends on how each layer attends, and modern models mix several kinds of layers:

Sliding-window attention. Gemma 4 31B keeps a 1,024-token window on 50 of its 60 layers. Only

10 global layers cache the whole context. At 128K tokens the real cache is 10.8 GiB; treating

every layer as global gives 120 GiB, 11× too much.

Linear attention. Qwen3.6 35B-A3B caches keys and values on only 10 of its 40 layers; the

other 30 keep a small fixed-size state. Real cache at 128K: 2.5 GiB, not 10.

Latent attention (MLA). DeepSeek R1 stores a 576-wide latent per layer instead of full keys and

values for 128 heads. A multi-head formula would predict about 610 GiB at 128K; the real

cache is 8.6 GiB.

Compressed and sparse attention. DeepSeek V4 Flash compresses the sequence and keeps a

128-token window: about 7× less than the naive estimate.

The calculator reads each model’s architecture layer by layer and computes the cache the way vLLM or llama.cpp actually allocate it, including FP8 cache formats and the window padding llama.cpp

adds.

Pick a model (or paste any Hugging Face link), a context length, a number of concurrent requests, and a GPU. You get:

The answer first. How many GPUs you need (e.g. “1 × NVIDIA H100 · 80 GB”), whether it fits,

is tight, or does not fit, and a stacked bar of what fills each card: weights, KV cache,

recurrent state, runtime overhead.

Weight formats. BF16, FP8, INT8, NVFP4, MXFP4, GPTQ/AWQ, and GGUF Q8_0 down to Q3_K_M, with a

“does it fit?” matrix of every format against every context length. GGUF sizes follow

llama.cpp’s real quantization layout (embeddings and output kept at higher precision, tensors

whose width is not a multiple of 256 falling back to wider formats). Checked against the actual

GGUF files on Hugging Face, the median error is under 1% for every common quant.

60 GPUs. Data-center (H100, H200, B200, MI300X…), workstation and consumer cards, laptop GPUs,

and unified-memory machines like Macs, with the memory each one really exposes.

Multi-GPU and engines. Tensor and pipeline parallelism, for vLLM, SGLang, TensorRT-LLM and

llama.cpp / Ollama / LM Studio, each with its own memory reservation and overhead.

Not enough VRAM? An offload plan: how many MoE experts or layers to move to system RAM

(llama.cpp --n-cpu-moe / -ngl), and what speed to expect from your RAM type.

Fine-tuning. Full fine-tune, LoRA and QLoRA: gradients, optimizer states, activations.

Speed and cost. Rough tokens per second, time to first token, and your cost per million

tokens compared with API prices.

Rent and deploy. On-demand prices from RunPod, Vast.ai, Verda and Azure, refreshed every hour

with price history, the cheapest GPUs that fit your setup, and a ready-to-paste vllm serve /

llama-server / SGLang / TensorRT-LLM command.

Check it against reality. Paste your vLLM or llama.cpp startup log and the calculator compares

what the engine really allocated with its estimate.

Each model has its own page (example: gpt-oss 120B): VRAM by precision and context length, how many GPUs of each type it takes, the cheapest way to run it today, its KV cache curve, and the real sizes of the published GGUF files.

New models appear on their own. A watcher checks Hugging Face every hour, reads each new model’s config.json and safetensors headers (without down the weights), and publishes a page when the architecture is one the calculator handles. The “New & trending” tab follows Hugging Face trends, Ollama’s most pulled models and the latest releases.

MCP server. Ask Claude, ChatGPT or any MCP client “does Qwen3.6 35B fit on my RTX 4090?” and it

calls the calculator. Three tools: estimate_vram, models_that_fit, gpu_prices. No key.

claude mcp add --transport http studiotv-vram https://studiotvai.com/api/mcp

JSON API. /api/models.json, one file per model, and

/api/gpus.json.

VRAM badge for a model card or README, and an RSS feed of new models.

These are estimates. Real usage varies with engine versions and settings, so keep some headroom and confirm with nvidia-smi. Speed figures are rough (±30-50%) and do not model speculative decoding. Some “Rent” links are affiliate links; that is disclosed on the site and does not change the ranking, which is by price.

If you find a model where the estimate is off, paste your engine log into the calculator or tell me: that is exactly the feedback I am looking for !

── more in #large-language-models 4 stories · sorted by recency
── more on @studiotv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/will-it-fit-on-my-gp…] indexed:0 read:4min 2026-10-10 · —