Will it fit on my GPU ? Building an LLM VRAM calculator that models each architecture's real KV cache StudioTV released a free LLM VRAM Calculator that models each architecture's real KV cache layer by layer instead of assuming full attention, correcting naive estimates that can be off by an order of magnitude. The tool reports that Gemma 4 31B's real 128K-token cache is 10.8 GiB versus 120 GiB for a full-attention formula, Qwen3.6 35B-A3B's is 2.5 GiB, and DeepSeek R1's latent attention cache is 8.6 GiB against a 610 GiB multi-head prediction. It covers 60 GPUs, weight formats from BF16 to GGUF Q3_K_M, vLLM, SGLang, TensorRT-LLM and llama.cpp, and hourly-refreshed rental prices from RunPod, Vast.ai, Verda and Azure. Every time a new open model comes out, the same question shows up in every thread: will it run on my card? With 24 GB? With two of them? In 4-bit? At what context length? Most answers come from a back-of-the-envelope formula: parameters × bytes per weight, plus a KV cache computed as if every layer were a classic full-attention layer. For a Llama-style dense model that is fine. For most models released in the last year, it is wrong, sometimes by an order of magnitude. I built StudioTV’s LLM VRAM Calculator https://studiotvai.com/llm-vram-calculator to answer that question properly, for every recent architecture, and to go one step further: what it costs to rent the hardware, and the exact command to launch it. It is free, has no sign-up, and works for any model on Hugging Face. The KV cache, the memory that grows with context length, depends on how each layer attends, and modern models mix several kinds of layers: Sliding-window attention. Gemma 4 31B keeps a 1,024-token window on 50 of its 60 layers. Only 10 global layers cache the whole context. At 128K tokens the real cache is 10.8 GiB ; treating every layer as global gives 120 GiB , 11× too much. Linear attention. Qwen3.6 35B-A3B caches keys and values on only 10 of its 40 layers; the other 30 keep a small fixed-size state. Real cache at 128K: 2.5 GiB , not 10. Latent attention MLA . DeepSeek R1 stores a 576-wide latent per layer instead of full keys and values for 128 heads. A multi-head formula would predict about 610 GiB at 128K; the real cache is 8.6 GiB . Compressed and sparse attention. DeepSeek V4 Flash compresses the sequence and keeps a 128-token window: about 7× less than the naive estimate. The calculator reads each model’s architecture layer by layer and computes the cache the way vLLM or llama.cpp actually allocate it, including FP8 cache formats and the window padding llama.cpp adds. Pick a model or paste any Hugging Face link , a context length, a number of concurrent requests, and a GPU. You get: The answer first. How many GPUs you need e.g. “1 × NVIDIA H100 · 80 GB” , whether it fits, is tight, or does not fit, and a stacked bar of what fills each card: weights, KV cache, recurrent state, runtime overhead. Weight formats. BF16, FP8, INT8, NVFP4, MXFP4, GPTQ/AWQ, and GGUF Q8 0 down to Q3 K M, with a “does it fit?” matrix of every format against every context length. GGUF sizes follow llama.cpp’s real quantization layout embeddings and output kept at higher precision, tensors whose width is not a multiple of 256 falling back to wider formats . Checked against the actual GGUF files on Hugging Face, the median error is under 1% for every common quant. 60 GPUs. Data-center H100, H200, B200, MI300X… , workstation and consumer cards, laptop GPUs, and unified-memory machines like Macs, with the memory each one really exposes. Multi-GPU and engines. Tensor and pipeline parallelism, for vLLM, SGLang, TensorRT-LLM and llama.cpp / Ollama / LM Studio, each with its own memory reservation and overhead. Not enough VRAM? An offload plan: how many MoE experts or layers to move to system RAM llama.cpp --n-cpu-moe / -ngl , and what speed to expect from your RAM type. Fine-tuning. Full fine-tune, LoRA and QLoRA: gradients, optimizer states, activations. Speed and cost. Rough tokens per second, time to first token, and your cost per million tokens compared with API prices. Rent and deploy. On-demand prices from RunPod, Vast.ai, Verda and Azure, refreshed every hour with price history, the cheapest GPUs that fit your setup, and a ready-to-paste vllm serve / llama-server / SGLang / TensorRT-LLM command. Check it against reality. Paste your vLLM or llama.cpp startup log and the calculator compares what the engine really allocated with its estimate. Each model has its own page example: gpt-oss 120B https://studiotvai.com/vram-requirements/gpt-oss-120b : VRAM by precision and context length, how many GPUs of each type it takes, the cheapest way to run it today, its KV cache curve, and the real sizes of the published GGUF files. New models appear on their own. A watcher checks Hugging Face every hour, reads each new model’s config.json and safetensors headers without downloading the weights , and publishes a page when the architecture is one the calculator handles. The “New & trending” tab follows Hugging Face trends, Ollama’s most pulled models and the latest releases. MCP server. Ask Claude, ChatGPT or any MCP client “does Qwen3.6 35B fit on my RTX 4090?” and it calls the calculator. Three tools: estimate vram , models that fit , gpu prices . No key. claude mcp add --transport http studiotv-vram https://studiotvai.com/api/mcp JSON API. /api/models.json, one file per model, and /api/gpus.json. VRAM badge for a model card or README, and an RSS feed of new models. These are estimates. Real usage varies with engine versions and settings, so keep some headroom and confirm with nvidia-smi . Speed figures are rough ±30-50% and do not model speculative decoding. Some “Rent” links are affiliate links; that is disclosed on the site and does not change the ranking, which is by price. If you find a model where the estimate is off, paste your engine log into the calculator or tell me: that is exactly the feedback I am looking for