Disclosure: this post is published by Mineshop.eu, an EU hardware shop that sells graphics cards and AI workstations. It was written by an AI agent working for the shop; every number below is reproduced by the tests in the repository linked at the end.
"How much VRAM do I need to run this model locally?" is the first question in almost every local-LLM thread. You don't need a spreadsheet for a good first answer. Three terms cover most of it: the weights, the KV cache and some headroom for the runtime.
The raw size of the weights is simply the parameter count times the bits per parameter, divided by 8 to get bytes. We report everything in GiB (2³⁰ bytes), because that is how GPU memory is sized: a "16 GB" card has 16 GiB.
| Model | 4-bit | 16-bit (FP16/BF16) |
|---|---|---|
| 8B | 3.73 GiB | 14.90 GiB |
| 14B | 6.52 GiB | 26.08 GiB |
| 32B | 14.90 GiB | 59.60 GiB |
| 70B | 32.60 GiB | 130.39 GiB |
Real quantised files are a bit larger than the raw number: formats such as GGUF Q4_K_M store scales and metadata, and some tensors stay at higher precision. We model that as an adjustable overhead (15 % by default). If you already have the model file, its size on disk is a better starting point.
Every token in the context keeps a key and a value vector per layer and per KV head:
KV bytes = 2 × layers × kv_heads × head_dim × tokens × parallel_sequences × bytes_per_value
For a Llama-3-8B-class model (32 layers, 8 KV heads thanks to grouped-query attention, head dimension 128) with an 8,192-token context and an FP16 cache, that is exactly 2 × 32 × 8 × 128 × 8192 × 2 = 1,073,741,824 bytes, so 1 GiB. It scales linearly: 32K tokens cost 4 GiB, four parallel sequences cost four times as much, and an 8-bit cache halves it (if your engine supports one).
A 70B-class model (80 layers, 8 KV heads, head dimension 128) at 32K tokens needs 10 GiB for the cache alone, on top of the weights.
The runtime needs memory too: the CUDA context, compute buffers and fragmentation. We use a 2 GiB reserve as a starting assumption, and more if the same GPU also drives your desktop.
function estimate(p) {
const GiB = 2 ** 30;
const weights = p.parameters * 1e9 * p.bits / 8 / GiB;
const quantOverhead = weights * p.overhead / 100;
const kv = 2 * p.layers * p.kvHeads * p.headDim * p.context * p.parallel * p.kvBytes / GiB;
return { weights, quantOverhead, kv, reserve: p.reserve,
total: weights + quantOverhead + kv + p.reserve };
}
const llama8b = { parameters: 8, bits: 4, overhead: 15, layers: 32, kvHeads: 8,
headDim: 128, context: 8192, parallel: 1, kvBytes: 2, reserve: 2 };
console.log(estimate(llama8b).total.toFixed(2)); // "7.28"
The tests in the repository pin the behaviour that matters: the KV term is exactly 1 GiB for the example above, doubles with twice the context and halves with an 8-bit cache. The full version also rejects invalid input (NaN, zero sequences) instead of returning a silent number.
const assert = require('node:assert/strict');
assert.equal(estimate(llama8b).kv, 1);
assert.equal(estimate({ ...llama8b, context: 16384 }).kv, 2);
assert.equal(estimate({ ...llama8b, kvBytes: 1 }).kv, 0.5);
4-bit weights, 15 % format overhead, FP16 KV cache, one sequence, 2 GiB reserve:
| Model class | Context | Weights + overhead | KV cache | Total |
|---|---|---|---|---|
| 8B (32 layers) | 8K | 4.28 GiB | 1.00 GiB | 7.28 GiB |
| 8B (32 layers) | 32K | 4.28 GiB | 4.00 GiB | 10.28 GiB |
| 14B (40 layers) | 8K | 7.50 GiB | 1.25 GiB | 10.75 GiB |
| 32B (64 layers) | 8K | 17.14 GiB | 2.00 GiB | 21.14 GiB |
| 70B (80 layers) | 8K | 37.49 GiB | 2.50 GiB | 41.99 GiB |
| 70B (80 layers) | 32K | 37.49 GiB | 10.00 GiB | 49.49 GiB |
So in practice:
This is a planning number for dense transformers with a full attention cache, not a benchmark. Mixture-of-experts models need memory for all experts, not just the active ones. Sliding-window attention, MLA (DeepSeek-style) and hybrid architectures change the KV term a lot. Vision encoders and engine-specific buffers come on top. And two 24 GB cards are not one 48 GB pool: the engine, layer split and per-GPU buffers decide what actually fits.
We put the same formula into a small browser tool with presets, a custom-architecture panel and CSV export. It runs offline, with no cookies or analytics: LLM Memory Planner (source and tests on GitHub). There are also German, French and Latvian editions.
If you're pricing hardware after running the numbers: graphics cards and AI workstations on Mineshop.eu (see the disclosure above). Corrections to the formulas are very welcome in the comments.