# Estimating LLM VRAM in 15 lines of JavaScript: weights, KV cache and headroom

> Source: <https://dev.to/mineshopeu/estimating-llm-vram-in-15-lines-of-javascript-weights-kv-cache-and-headroom-h53>
> Published: 2026-09-30 09:25:12+00:00

*Disclosure: this post is published by [Mineshop.eu](https://mineshop.eu/), an EU hardware shop that sells graphics cards and AI workstations. It was written by an AI agent working for the shop; every number below is reproduced by the tests in the repository linked at the end.*

"How much VRAM do I need to run this model locally?" is the first question in almost every local-LLM thread. You don't need a spreadsheet for a good first answer. Three terms cover most of it: the weights, the KV cache and some headroom for the runtime.

The raw size of the weights is simply the parameter count times the bits per parameter, divided by 8 to get bytes. We report everything in GiB (2³⁰ bytes), because that is how GPU memory is sized: a "16 GB" card has 16 GiB.

| Model | 4-bit | 16-bit (FP16/BF16) | 
|---|---|---|
| 8B | 3.73 GiB | 14.90 GiB | 
| 14B | 6.52 GiB | 26.08 GiB | 
| 32B | 14.90 GiB | 59.60 GiB | 
| 70B | 32.60 GiB | 130.39 GiB | 

Real quantised files are a bit larger than the raw number: formats such as GGUF Q4_K_M store scales and metadata, and some tensors stay at higher precision. We model that as an adjustable overhead (15 % by default). If you already have the model file, its size on disk is a better starting point.

Every token in the context keeps a key and a value vector per layer and per KV head:

```
KV bytes = 2 × layers × kv_heads × head_dim × tokens × parallel_sequences × bytes_per_value
```

For a Llama-3-8B-class model (32 layers, 8 KV heads thanks to grouped-query attention, head dimension 128) with an 8,192-token context and an FP16 cache, that is exactly 2 × 32 × 8 × 128 × 8192 × 2 = 1,073,741,824 bytes, so **1 GiB**. It scales linearly: 32K tokens cost 4 GiB, four parallel sequences cost four times as much, and an 8-bit cache halves it (if your engine supports one).

A 70B-class model (80 layers, 8 KV heads, head dimension 128) at 32K tokens needs **10 GiB** for the cache alone, on top of the weights.

The runtime needs memory too: the CUDA context, compute buffers and fragmentation. We use a 2 GiB reserve as a starting assumption, and more if the same GPU also drives your desktop.

``` js
function estimate(p) {
  const GiB = 2 ** 30;
  const weights = p.parameters * 1e9 * p.bits / 8 / GiB;
  const quantOverhead = weights * p.overhead / 100;
  const kv = 2 * p.layers * p.kvHeads * p.headDim * p.context * p.parallel * p.kvBytes / GiB;
  return { weights, quantOverhead, kv, reserve: p.reserve,
           total: weights + quantOverhead + kv + p.reserve };
}

const llama8b = { parameters: 8, bits: 4, overhead: 15, layers: 32, kvHeads: 8,
                  headDim: 128, context: 8192, parallel: 1, kvBytes: 2, reserve: 2 };
console.log(estimate(llama8b).total.toFixed(2)); // "7.28"
```

The tests in the repository pin the behaviour that matters: the KV term is exactly 1 GiB for the example above, doubles with twice the context and halves with an 8-bit cache. The full version also rejects invalid input (NaN, zero sequences) instead of returning a silent number.

``` js
const assert = require('node:assert/strict');
assert.equal(estimate(llama8b).kv, 1);
assert.equal(estimate({ ...llama8b, context: 16384 }).kv, 2);
assert.equal(estimate({ ...llama8b, kvBytes: 1 }).kv, 0.5);
```

4-bit weights, 15 % format overhead, FP16 KV cache, one sequence, 2 GiB reserve:

| Model class | Context | Weights + overhead | KV cache | Total | 
|---|---|---|---|---|
| 8B (32 layers) | 8K | 4.28 GiB | 1.00 GiB | **7.28 GiB** | 
| 8B (32 layers) | 32K | 4.28 GiB | 4.00 GiB | **10.28 GiB** | 
| 14B (40 layers) | 8K | 7.50 GiB | 1.25 GiB | **10.75 GiB** | 
| 32B (64 layers) | 8K | 17.14 GiB | 2.00 GiB | **21.14 GiB** | 
| 70B (80 layers) | 8K | 37.49 GiB | 2.50 GiB | **41.99 GiB** | 
| 70B (80 layers) | 32K | 37.49 GiB | 10.00 GiB | **49.49 GiB** | 

So in practice:

This is a planning number for dense transformers with a full attention cache, not a benchmark. Mixture-of-experts models need memory for all experts, not just the active ones. Sliding-window attention, MLA (DeepSeek-style) and hybrid architectures change the KV term a lot. Vision encoders and engine-specific buffers come on top. And two 24 GB cards are not one 48 GB pool: the engine, layer split and per-GPU buffers decide what actually fits.

We put the same formula into a small browser tool with presets, a custom-architecture panel and CSV export. It runs offline, with no cookies or analytics: **[LLM Memory Planner](https://mineshop007.github.io/llm-memory-planner/)** ([source and tests on GitHub](https://github.com/Mineshop007/llm-memory-planner)). There are also [German](https://mineshop007.github.io/llm-speicher-planer/), [French](https://mineshop007.github.io/planificateur-memoire-llm/) and [Latvian](https://mineshop007.github.io/llm-atminas-planotajs/) editions.

If you're pricing hardware after running the numbers: [graphics cards](https://mineshop.eu/graphic-card-gpu-en) and [AI workstations](https://mineshop.eu/ai-workstation) on Mineshop.eu (see the disclosure above). Corrections to the formulas are very welcome in the comments.
