cd /news/large-language-models/fine-tuning-a-7b-model-needs-112-gb-… · home › topics › large-language-models › article
[ARTICLE · art-144270] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Fine-tuning a 7B model needs 112 GB. The model is only 14 GB of it.

A developer's breakdown of fine-tuning memory shows that training a 7B-parameter model in mixed precision with Adam requires roughly 112 GB, of which only 14 GB is the fp16 weights — the optimizer state alone accounts for 84 GB, or six times the model. The analysis, drawing on the ZeRO, LoRA and QLoRA papers, explains that LoRA cuts trainable parameters to about 0.06% of the model by freezing the base and training rank-8 adapters, while QLoRA's 4-bit base quantization reduces the frozen weights to about 3.5 GB at the cost of slower dequantization on every pass.

by read3 min views1 publishedOct 3, 2026

Ask how much memory it takes to fine-tune a 7B model and the instinct is "the model's 14 GB in fp16, so a bit more than that". The real figure is about 112 GB, before you've stored a single activation. The model is 14 GB of it.

Once you see where the other 98 GB goes, LoRA and QLoRA stop looking like clever tricks and start looking obvious.

The accounting comes from the ZeRO paper, and it's worth reading in their words:

In total, this results in 2Ψ + 2Ψ + KΨ = 16Ψ bytes of memory requirement. For a model such as GPT-2 with 1.5 Billion parameters, this leads to a memory requirement of at least 24 GB, which is significantly higher than the meager 3 GB of memory required to hold the fp16 parameters alone.

Ψ is the parameter count. Per parameter, mixed-precision training with Adam holds:

Sixteen bytes per parameter. For 7 billion parameters that's 112 GB, and only 14 GB of it is the model you're actually trying to change. The optimizer state alone is 84 GB, six times the weights.

That's the number I wish I'd seen earlier. Fine-tuning memory isn't dominated by the model. It's dominated by the bookkeeping needed to update it.

If most of the cost is gradients and optimizer state, the obvious move is to have far fewer things that need them. That's LoRA. Freeze every pretrained weight, and train a pair of small matrices injected into each layer instead. The frozen base still sits in memory at fp16 — 14 GB — but it carries no gradients and no optimizer state. Only the small matrices do.

How small? Take a Llama-style 7B: 32 layers, hidden size 4096, with adapters on the query and value projections. At rank 8, each adapter is 8 × (4096 + 4096) = 65,536 values. Across 32 layers and two projections that's 4,194,304 trainable parameters.

About 0.06% of the model. You're training four million numbers, not seven billion.

Hu et al. put the GPT-3 result plainly: compared to GPT-3 175B fine-tuned with Adam, LoRA "can reduce the number of trainable parameters by 10,000 times and the GPU memory requirement by 3 times", while it "performs on-par or better than finetuning in model quality" on the models they tested.

And because the low-rank update merges back into the base weight afterwards, there's "no additional inference latency". That's what separated it from earlier adapter methods, which left an extra layer sitting in the forward pass permanently.

LoRA leaves one big cost standing: the frozen base still needs 14 GB at fp16.

QLoRA goes after exactly that. Store the frozen base in 4-bit and it drops to about 3.5 GB — slightly more in practice, because quantization needs constants of its own. Dettmers et al. squeezed those too: their double quantization cuts the overhead "from 32/64 = 0.5 bits, to 8/64 + 32/(64 · 256) = 0.127 bits" per parameter.

Gradients still flow through the 4-bit base into LoRA adapters kept at higher precision. The headline result is the one that made it famous — enough memory saved "to finetune a 65B parameter model on a single 48GB GPU while preserving full 16-bit finetuning task performance".

The cost is time. The 4-bit weights have to be dequantized to do any arithmetic with them, on every pass, so it runs slower than plain LoRA. You're trading speed for memory.

The three methods are really one question: which of the 16 bytes per parameter are you willing to stop paying for?

If someone says their 7B model "needs 14 GB", they're quoting inference. Ask about the other 98. Longer version with the full memory table and a rank-by-rank worked example at diffstudy.com. Sources: ZeRO, arXiv 1910.02054, LoRA, arXiv 2106.09685, QLoRA, arXiv 2305.14314.

── more in #large-language-models 4 stories · sorted by recency
── more on @lora 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/fine-tuning-a-7b-mod…] indexed:0 read:3min 2026-10-03 · —