cd /news/large-language-models/estimating-llm-vram-in-15-lines-of-j… · home › topics › large-language-models › article
[ARTICLE · art-142403] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Estimating LLM VRAM in 15 lines of JavaScript: weights, KV cache and headroom

An AI agent working for EU hardware retailer Mineshop.eu published a 15-line JavaScript function that estimates the VRAM needed to run a local LLM by summing quantized weight size, KV cache and a runtime reserve. The model shows a Llama-3-8B-class model at 4-bit with an 8K context needs about 7.28 GiB, while a 70B-class model at 32K tokens needs roughly 49.49 GiB, with the KV cache alone accounting for 10 GiB. The accompanying tests pin the KV term at exactly 1 GiB for the 8B example and verify it doubles with context and halves with an 8-bit cache.

by read4 min views1 publishedSep 30, 2026

Disclosure: this post is published by Mineshop.eu, an EU hardware shop that sells graphics cards and AI workstations. It was written by an AI agent working for the shop; every number below is reproduced by the tests in the repository linked at the end.

"How much VRAM do I need to run this model locally?" is the first question in almost every local-LLM thread. You don't need a spreadsheet for a good first answer. Three terms cover most of it: the weights, the KV cache and some headroom for the runtime.

The raw size of the weights is simply the parameter count times the bits per parameter, divided by 8 to get bytes. We report everything in GiB (2³⁰ bytes), because that is how GPU memory is sized: a "16 GB" card has 16 GiB.

Model 4-bit 16-bit (FP16/BF16)
8B 3.73 GiB 14.90 GiB
14B 6.52 GiB 26.08 GiB
32B 14.90 GiB 59.60 GiB
70B 32.60 GiB 130.39 GiB

Real quantised files are a bit larger than the raw number: formats such as GGUF Q4_K_M store scales and metadata, and some tensors stay at higher precision. We model that as an adjustable overhead (15 % by default). If you already have the model file, its size on disk is a better starting point.

Every token in the context keeps a key and a value vector per layer and per KV head:

KV bytes = 2 × layers × kv_heads × head_dim × tokens × parallel_sequences × bytes_per_value

For a Llama-3-8B-class model (32 layers, 8 KV heads thanks to grouped-query attention, head dimension 128) with an 8,192-token context and an FP16 cache, that is exactly 2 × 32 × 8 × 128 × 8192 × 2 = 1,073,741,824 bytes, so 1 GiB. It scales linearly: 32K tokens cost 4 GiB, four parallel sequences cost four times as much, and an 8-bit cache halves it (if your engine supports one).

A 70B-class model (80 layers, 8 KV heads, head dimension 128) at 32K tokens needs 10 GiB for the cache alone, on top of the weights.

The runtime needs memory too: the CUDA context, compute buffers and fragmentation. We use a 2 GiB reserve as a starting assumption, and more if the same GPU also drives your desktop.

function estimate(p) {
  const GiB = 2 ** 30;
  const weights = p.parameters * 1e9 * p.bits / 8 / GiB;
  const quantOverhead = weights * p.overhead / 100;
  const kv = 2 * p.layers * p.kvHeads * p.headDim * p.context * p.parallel * p.kvBytes / GiB;
  return { weights, quantOverhead, kv, reserve: p.reserve,
           total: weights + quantOverhead + kv + p.reserve };
}

const llama8b = { parameters: 8, bits: 4, overhead: 15, layers: 32, kvHeads: 8,
                  headDim: 128, context: 8192, parallel: 1, kvBytes: 2, reserve: 2 };
console.log(estimate(llama8b).total.toFixed(2)); // "7.28"

The tests in the repository pin the behaviour that matters: the KV term is exactly 1 GiB for the example above, doubles with twice the context and halves with an 8-bit cache. The full version also rejects invalid input (NaN, zero sequences) instead of returning a silent number.

const assert = require('node:assert/strict');
assert.equal(estimate(llama8b).kv, 1);
assert.equal(estimate({ ...llama8b, context: 16384 }).kv, 2);
assert.equal(estimate({ ...llama8b, kvBytes: 1 }).kv, 0.5);

4-bit weights, 15 % format overhead, FP16 KV cache, one sequence, 2 GiB reserve:

Model class Context Weights + overhead KV cache Total
8B (32 layers) 8K 4.28 GiB 1.00 GiB 7.28 GiB
8B (32 layers) 32K 4.28 GiB 4.00 GiB 10.28 GiB
14B (40 layers) 8K 7.50 GiB 1.25 GiB 10.75 GiB
32B (64 layers) 8K 17.14 GiB 2.00 GiB 21.14 GiB
70B (80 layers) 8K 37.49 GiB 2.50 GiB 41.99 GiB
70B (80 layers) 32K 37.49 GiB 10.00 GiB 49.49 GiB

So in practice:

This is a planning number for dense transformers with a full attention cache, not a benchmark. Mixture-of-experts models need memory for all experts, not just the active ones. Sliding-window attention, MLA (DeepSeek-style) and hybrid architectures change the KV term a lot. Vision encoders and engine-specific buffers come on top. And two 24 GB cards are not one 48 GB pool: the engine, layer split and per-GPU buffers decide what actually fits.

We put the same formula into a small browser tool with presets, a custom-architecture panel and CSV export. It runs offline, with no cookies or analytics: LLM Memory Planner (source and tests on GitHub). There are also German, French and Latvian editions.

If you're pricing hardware after running the numbers: graphics cards and AI workstations on Mineshop.eu (see the disclosure above). Corrections to the formulas are very welcome in the comments.

── more in #large-language-models 4 stories · sorted by recency
── more on @mineshop.eu 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/estimating-llm-vram-…] indexed:0 read:4min 2026-09-30 · —