# Will it fit on my GPU ? Building an LLM VRAM calculator that models each architecture's real KV cache

> Source: <https://discuss.huggingface.co/t/will-it-fit-on-my-gpu-building-an-llm-vram-calculator-that-models-each-architectures-real-kv-cache/190291#post_1>
> Published: 2026-10-10 13:48:31+00:00

Every time a new open model comes out, the same question shows up in every thread: will it run on my card? With 24 GB? With two of them? In 4-bit? At what context length?

Most answers come from a back-of-the-envelope formula: parameters × bytes per weight, plus a KV cache computed as if every layer were a classic full-attention layer. For a Llama-style dense model that is fine. For most models released in the last year, it is wrong, sometimes by an order of magnitude.

I built [StudioTV’s LLM VRAM Calculator](https://studiotvai.com/llm-vram-calculator) to answer that question properly, for every recent architecture, and to go one step further: what it costs to rent the hardware, and the exact command to launch it. It is free, has no sign-up, and works for any model on Hugging Face.

The KV cache, the memory that grows with context length, depends on how each layer attends, and modern models mix several kinds of layers:

**Sliding-window attention.** Gemma 4 31B keeps a 1,024-token window on 50 of its 60 layers. Only

10 global layers cache the whole context. At 128K tokens the real cache is **10.8 GiB**; treating

every layer as global gives **120 GiB**, 11× too much.

**Linear attention.** Qwen3.6 35B-A3B caches keys and values on only 10 of its 40 layers; the

other 30 keep a small fixed-size state. Real cache at 128K: **2.5 GiB**, not 10.

**Latent attention (MLA).** DeepSeek R1 stores a 576-wide latent per layer instead of full keys and

values for 128 heads. A multi-head formula would predict about **610 GiB** at 128K; the real

cache is **8.6 GiB**.

**Compressed and sparse attention.** DeepSeek V4 Flash compresses the sequence and keeps a

128-token window: about 7× less than the naive estimate.

The calculator reads each model’s architecture layer by layer and computes the cache the way vLLM or llama.cpp actually allocate it, including FP8 cache formats and the window padding llama.cpp

adds.

Pick a model (or paste any Hugging Face link), a context length, a number of concurrent requests, and a GPU. You get:

**The answer first.** How many GPUs you need (e.g. “1 × NVIDIA H100 · 80 GB”), whether it fits,

is tight, or does not fit, and a stacked bar of what fills each card: weights, KV cache,

recurrent state, runtime overhead.

**Weight formats.** BF16, FP8, INT8, NVFP4, MXFP4, GPTQ/AWQ, and GGUF Q8_0 down to Q3_K_M, with a

“does it fit?” matrix of every format against every context length. GGUF sizes follow

llama.cpp’s real quantization layout (embeddings and output kept at higher precision, tensors

whose width is not a multiple of 256 falling back to wider formats). Checked against the actual

GGUF files on Hugging Face, the median error is under 1% for every common quant.

**60 GPUs.** Data-center (H100, H200, B200, MI300X…), workstation and consumer cards, laptop GPUs,

and unified-memory machines like Macs, with the memory each one really exposes.

**Multi-GPU and engines.** Tensor and pipeline parallelism, for vLLM, SGLang, TensorRT-LLM and

llama.cpp / Ollama / LM Studio, each with its own memory reservation and overhead.

**Not enough VRAM?** An offload plan: how many MoE experts or layers to move to system RAM

(llama.cpp `--n-cpu-moe` / `-ngl`), and what speed to expect from your RAM type.

**Fine-tuning.** Full fine-tune, LoRA and QLoRA: gradients, optimizer states, activations.

**Speed and cost.** Rough tokens per second, time to first token, and your cost per million

tokens compared with API prices.

**Rent and deploy.** On-demand prices from RunPod, Vast.ai, Verda and Azure, refreshed every hour

with price history, the cheapest GPUs that fit your setup, and a ready-to-paste `vllm serve` /

`llama-server` / SGLang / TensorRT-LLM command.

**Check it against reality.** Paste your vLLM or llama.cpp startup log and the calculator compares

what the engine really allocated with its estimate.

Each model has its own page ([example: gpt-oss 120B](https://studiotvai.com/vram-requirements/gpt-oss-120b)): VRAM by precision and context length, how many GPUs of each type it takes, the cheapest way to run it today, its KV cache curve, and the real sizes of the published GGUF files.

New models appear on their own. A watcher checks Hugging Face every hour, reads each new model’s `config.json` and safetensors headers (without downloading the weights), and publishes a page when the architecture is one the calculator handles. The “New & trending” tab follows Hugging Face trends, Ollama’s most pulled models and the latest releases.

**MCP server.** Ask Claude, ChatGPT or any MCP client “does Qwen3.6 35B fit on my RTX 4090?” and it

calls the calculator. Three tools: `estimate_vram`, `models_that_fit`, `gpu_prices`. No key.

```
claude mcp add --transport http studiotv-vram https://studiotvai.com/api/mcp
```

**JSON API.** /api/models.json, one file per model, and

/api/gpus.json.

**VRAM badge** for a model card or README, and an RSS feed of new models.

These are estimates. Real usage varies with engine versions and settings, so keep some headroom and confirm with `nvidia-smi`. Speed figures are rough (±30-50%) and do not model speculative decoding. Some “Rent” links are affiliate links; that is disclosed on the site and does not change the ranking, which is by price.

If you find a model where the estimate is off, paste your engine log into the calculator or tell me: that is exactly the feedback I am looking for !
