cd /news/large-language-models/llm-vram-requirements-what-fits-on-8… · home › topics › large-language-models › article
[ARTICLE · art-145837] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

LLM VRAM requirements: what fits on 8, 16, 24, 48 and 80GB

A developer published a tier-by-tier breakdown of which large language models fit on 8GB, 16GB, 24GB, 48GB and 80GB of GPU VRAM, assuming standard 4-bit (Q4_K_M) quantization with moderate context headroom. The guide reports that 7B–9B models fit in 8GB, 13B–14B in 16GB, dense 30B–32B and MoE models such as Qwen3 30B-A3B in 24GB, 70B models at roughly 40–42GB in 48GB, and 70B at FP8/INT8 plus large MoE models like Llama 4 Scout (~67GB at Q4) in 80GB. It notes that 48GB is the practical dividing line for 70B-class models, which do not fit on a single 24GB card even at aggressive quantization.

by read6 min views1 publishedOct 6, 2026

Have you ever found yourself stuck in this question: "will this model fit on my GPU?". A lot of us have. The honest answer is always "it depends on quantization and context length," which is true but not actually useful when someone just wants to know if their card can run the model they want to try.

So here is the practical version. A real, tier-by-tier breakdown of what actually fits on the VRAM sizes people are most commonly working with, based on typical 4-bit quantization with reasonable context headroom.

Why does the same model need different amounts of VRAM on different GPUs?

Because model size is not the only variable. Two things change the actual footprint dramatically.

Quantization level: since dropping from full precision to 4-bit roughly quarters the memory a model's weights need

Context length: since every additional token of context adds to the KV cache sitting on top of the base weights

This guide assumes standard 4-bit quantization, commonly written as Q4 or Q4_K_M, with moderate context length, since that is the realistic default for most people running models locally or in production on a budget.

What actually fits on 8GB of VRAM?

Small, genuinely capable models, comfortably.

Llama 3.1 8B, which fits on any 8GB GPU with real room to spare

Qwen 3.5 9B, fitting in around 6.6GB, leaving headroom for context

Most 7B to 9B class models at Q4 quantization, generally landing in the 5 to 6GB range for weights alone

8GB is genuinely usable territory now, mostly thanks to newer, more efficient models in this size class rather than any change in how quantization works.

What actually fits on 16GB of VRAM?

A clear step up, mainly in model class rather than just headroom.

13B to 14B class models, comfortably fitting at Q4 with room left for context

8B class models, but now with substantially more headroom for longer conversations or larger context windows

This is generally considered the point where local LLM use starts feeling genuinely practical rather than tightly constrained

What actually fits on 24GB of VRAM?

This tier opens up a meaningfully larger class of model.

Dense 30B to 32B models, fitting at Q4 quantization, though somewhat tightly with limited context headroom

Mixture-of-experts models like Qwen3 30B-A3B, which carry 30B total parameters but only activate around 3B per token, running noticeably faster than a dense model of the same size while still fitting comfortably

Nvidia’s L4 GPU memory capacity is 24GB. It lands exactly in this range, which is part of why it has become such a common default for cost-efficient serving of 30B-class models in production, particularly MoE architectures that deliver strong performance without demanding the full compute of a much larger dense model.

What actually fits on 48GB of VRAM?

This is genuinely the practical entry point for 70B-class models, though not with much room to spare.

70B models at Q4 quantization, typically landing around 40 to 42GB, fitting on a single 48GB card with modest headroom

The same 70B models do not fit on a single 24GB card even at aggressive quantization, making 48GB the real dividing line for this model class

An alternative path is combining two 24GB GPUs, though multi-GPU setups add interconnect overhead that a single larger card avoids

What actually fits on 80GB of VRAM?

This tier gives you genuine breathing room, and access to model classes that simply do not fit anywhere smaller.

70B models at higher precision, such as FP8 or INT8, with real headroom left over instead of running right at the edge

Large mixture-of-experts models, like a 117B total parameter model with roughly 5B active parameters per token, which runs comfortably on a single 80GB GPU

Models like Llama 4 Scout, which need around 67GB at Q4 quantization since all experts stay resident in memory, fitting 80GB with room but not fitting a single 24GB or even 48GB card at all

Even at 80GB, this is not a limitless tier. Models in the 400 billion parameter range still require multi-GPU setups, commonly four or more 80GB cards working together, regardless of quantization.

How do you double-check these numbers for a specific model before committing to hardware?

The tiers above hold up well as a general reference, but individual models vary a little based on architecture details like the number of attention heads and layers. Before committing budget to a GPU, it is worth checking the specific model card or quantization repository you plan to use, since most popular open-weight models now list exact file sizes for each quantization level directly. That number, plus a reasonable buffer for context and runtime overhead, is a more reliable figure than a general size class alone.

This matters most right at the edge of a tier. A 70B model comfortably fitting 48GB with a short prompt can behave very differently once you push context length up toward 32K tokens or run several requests concurrently, which is exactly the kind of gap a quick model-specific check will catch before it becomes a production problem.

Quick reference: VRAM tier and practical model ceiling

VRAM tier

Practical model ceiling

Example

8GB

7B to 9B dense models

Llama 3.1 8B, Qwen 3.5 9B

16GB

13B to 14B dense models

Comfortable fit with context room

24GB

30B to 32B dense, or 30B-class MoE

Qwen3 30B-A3B

48GB

70B models at Q4

Llama 3.3 70B

80GB

70B at higher precision, or large MoE

117B total MoE, Llama 4 Scout

Where this leaves you

VRAM tiers map fairly predictably to model classes once you account for realistic quantization, which is exactly why this kind of lookup is more useful day to day than working through the full memory formula every time. Match your actual model choice to the tier that comfortably fits it, factor in context length honestly, and you will avoid both the common mistake of undersizing hardware and the equally common one of buying more memory than your actual workload will ever use.

Frequently asked question

Does more context length change these numbers significantly?

Yes, and this is the part people forget most often. These figures assume moderate context. Long context windows, especially past 32K tokens, can add a meaningful amount of memory on top of the base weight footprint, sometimes enough to push a model that "fits" right up against the ceiling.

Is it better to run a smaller model at higher precision or a bigger model quantized down?

Generally, a well-chosen quantized larger model outperforms a smaller model at full precision, especially at Q4 and above. The quality loss from quantization is usually smaller than the capability gap between meaningfully different model sizes.

Can I combine multiple smaller GPUs instead of buying one bigger card?

Yes, but it is not a perfectly clean substitute. Combined VRAM across multiple cards works for many setups, but interconnect speed between GPUs adds overhead that a single larger card with the same total memory does not have, which can noticeably affect throughput on larger models.

Should I always use the most aggressive quantization to fit a bigger model?

Not automatically. Q4 quantization typically costs only 1 to 2 percent in quality compared to full precision, which is usually a fine trade-off. Going more aggressive than that, down to Q2 or Q3, starts to noticeably hurt output quality. Fitting a model technically is not the same as fitting it well.

── more in #large-language-models 4 stories · sorted by recency
── more on @llama 3.1 8b 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/llm-vram-requirement…] indexed:0 read:6min 2026-10-06 · —