Every time a new open model comes out, the same question shows up in every thread: will it run on my card? With 24 GB? With two of them? In 4-bit? At what context length?
Most answers come from a back-of-the-envelope formula: parameters × bytes per weight, plus a KV cache computed as if every layer were a classic full-attention layer. For a Llama-style dense model that is fine. For most models released in the last year, it is wrong, sometimes by an order of magnitude.
I built StudioTV’s LLM VRAM Calculator to answer that question properly, for every recent architecture, and to go one step further: what it costs to rent the hardware, and the exact command to launch it. It is free, has no sign-up, and works for any model on Hugging Face.
The KV cache, the memory that grows with context length, depends on how each layer attends, and modern models mix several kinds of layers:
Sliding-window attention. Gemma 4 31B keeps a 1,024-token window on 50 of its 60 layers. Only
10 global layers cache the whole context. At 128K tokens the real cache is 10.8 GiB; treating
every layer as global gives 120 GiB, 11× too much.
Linear attention. Qwen3.6 35B-A3B caches keys and values on only 10 of its 40 layers; the
other 30 keep a small fixed-size state. Real cache at 128K: 2.5 GiB, not 10.
Latent attention (MLA). DeepSeek R1 stores a 576-wide latent per layer instead of full keys and
values for 128 heads. A multi-head formula would predict about 610 GiB at 128K; the real
cache is 8.6 GiB.
Compressed and sparse attention. DeepSeek V4 Flash compresses the sequence and keeps a
128-token window: about 7× less than the naive estimate.
The calculator reads each model’s architecture layer by layer and computes the cache the way vLLM or llama.cpp actually allocate it, including FP8 cache formats and the window padding llama.cpp
adds.
Pick a model (or paste any Hugging Face link), a context length, a number of concurrent requests, and a GPU. You get:
The answer first. How many GPUs you need (e.g. “1 × NVIDIA H100 · 80 GB”), whether it fits,
is tight, or does not fit, and a stacked bar of what fills each card: weights, KV cache,
recurrent state, runtime overhead.
Weight formats. BF16, FP8, INT8, NVFP4, MXFP4, GPTQ/AWQ, and GGUF Q8_0 down to Q3_K_M, with a
“does it fit?” matrix of every format against every context length. GGUF sizes follow
llama.cpp’s real quantization layout (embeddings and output kept at higher precision, tensors
whose width is not a multiple of 256 falling back to wider formats). Checked against the actual
GGUF files on Hugging Face, the median error is under 1% for every common quant.
60 GPUs. Data-center (H100, H200, B200, MI300X…), workstation and consumer cards, laptop GPUs,
and unified-memory machines like Macs, with the memory each one really exposes.
Multi-GPU and engines. Tensor and pipeline parallelism, for vLLM, SGLang, TensorRT-LLM and
llama.cpp / Ollama / LM Studio, each with its own memory reservation and overhead.
Not enough VRAM? An offload plan: how many MoE experts or layers to move to system RAM
(llama.cpp --n-cpu-moe / -ngl), and what speed to expect from your RAM type.
Fine-tuning. Full fine-tune, LoRA and QLoRA: gradients, optimizer states, activations.
Speed and cost. Rough tokens per second, time to first token, and your cost per million
tokens compared with API prices.
Rent and deploy. On-demand prices from RunPod, Vast.ai, Verda and Azure, refreshed every hour
with price history, the cheapest GPUs that fit your setup, and a ready-to-paste vllm serve /
llama-server / SGLang / TensorRT-LLM command.
Check it against reality. Paste your vLLM or llama.cpp startup log and the calculator compares
what the engine really allocated with its estimate.
Each model has its own page (example: gpt-oss 120B): VRAM by precision and context length, how many GPUs of each type it takes, the cheapest way to run it today, its KV cache curve, and the real sizes of the published GGUF files.
New models appear on their own. A watcher checks Hugging Face every hour, reads each new model’s config.json and safetensors headers (without down the weights), and publishes a page when the architecture is one the calculator handles. The “New & trending” tab follows Hugging Face trends, Ollama’s most pulled models and the latest releases.
MCP server. Ask Claude, ChatGPT or any MCP client “does Qwen3.6 35B fit on my RTX 4090?” and it
calls the calculator. Three tools: estimate_vram, models_that_fit, gpu_prices. No key.
claude mcp add --transport http studiotv-vram https://studiotvai.com/api/mcp
JSON API. /api/models.json, one file per model, and
/api/gpus.json.
VRAM badge for a model card or README, and an RSS feed of new models.
These are estimates. Real usage varies with engine versions and settings, so keep some headroom and confirm with nvidia-smi. Speed figures are rough (±30-50%) and do not model speculative decoding. Some “Rent” links are affiliate links; that is disclosed on the site and does not change the ranking, which is by price.
If you find a model where the estimate is off, paste your engine log into the calculator or tell me: that is exactly the feedback I am looking for !