cd /news/machine-learning/how-lora-actually-works-low-rank-dec… · home › topics › machine-learning › article
[ARTICLE · art-146793] src=dev.to ↗ pub= topic=machine-learning verified=true sentiment=· neutral

How LoRA Actually Works: Low-Rank Decomposition, Weight Merging, and Memory Breakdown Under the Hood

A developer's technical breakdown explains why full-parameter fine-tuning of an 8-billion-parameter model in 16-bit precision exceeds 80 GB of VRAM, since AdamW training consumes roughly 16 bytes per trainable parameter across weights, gradients, master weights, and two optimizer moments. The writeup details how Low-Rank Adaptation freezes base weights and decomposes the update matrix into two small matrices, cutting trainable parameters by about 128x per layer at rank 16 and reducing memory footprint by over 70% while retaining 99% of full fine-tuning performance.

by read7 min views3 publishedOct 7, 2026

If you try to full-parameter fine-tune an 8-billion parameter language model in 16-bit precision, your GPU memory requirement immediately explodes past 80 GB.

The raw model weights only occupy 16 GB of VRAM ($8 \times 10^9 \text{ parameters} \times 2 \text{ bytes}$). Yet standard training crashes on a single 80 GB A100 or H100.

Where does the missing 64+ GB of VRAM go? And why can Low-Rank Adaptation (LoRA) reduce that memory footprint by over 70% while updating fewer than 0.1% of the parameters and retaining 99% of full fine-tuning performance?

Here is what actually happens mathematically, mechanically, and in physical GPU memory during LoRA fine-tuning and inference.

When training a neural network with the standard AdamW optimizer, model weights are only a tiny fraction of your memory consumption.

For every single trainable parameter, the system must track:

$$\text{Total Memory per Trainable Parameter} = 2 + 2 + 4 + 4 + 4 = 16 \text{ bytes}$$

+-------------------------------------------------------------+
| Full Parameter Training (16 bytes per parameter)            |
+-------------------------------------------------------------+
| Weights: 2B | Grads: 2B | Master: 4B | Mom1: 4B | Mom2: 4B  |
+-------------------------------------------------------------+

For an 8B parameter model:

This is why full fine-tuning requires multi-GPU distributed clusters (DeepSpeed ZeRO or FSDP).

LoRA solves this problem at its root: by freezing the base weights, it eliminates optimizer states and gradients for 99.9% of the model.

Why does freezing the model work?

A standard linear layer in a transformer computes:

$$h = W_0 x$$

where $W_0 \in \mathbb{R}^{d \times k}$, $x \in \mathbb{R}^{k}$, and $h \in \mathbb{R}^{d}$. For Llama 3 8B, the hidden dimension $d = 4096$. Each weight matrix has up to $4096 \times 4096 \approx 16.7\text{M}$ parameters with full mathematical rank 4096.

In 2021, Edward Hu et al. published the foundational hypothesis behind LoRA:

When adapting a pre-trained language model to a specific task or instruction format, the weight update matrix $\Delta W$ has an extremely low "intrinsic rank" ($r \ll d$).

The pre-trained model already knows syntax, world facts, and language logic. Fine-tuning is merely steering attention and adjusting feature selection. You do not need to update 16 million degrees of freedom per matrix; you only need to update a tiny subspace of dimension $r$ (typically $r = 8, 16, \text{ or } 32$).

Instead of directly learning a full $d \times k$ matrix $\Delta W$, LoRA decomposes $\Delta W$ into the product of two small, low-rank matrices:

$$\Delta W = B \cdot A$$

where:

   Full Update Matrix ΔW (d × k)            LoRA Decomposition
   +---------------------------+       +-------+
   |                           |       |       |       +---------------------------+
 d |                           |   = d |   B   |   × r |             A             |
   |                           |       | (d×r) |       +---------------------------+
   +---------------------------+       +-------+                     k
                 k

Suppose $d = 4096, k = 4096$, and you choose rank $r = 16$:

$$\text{Reduction Factor} = \frac{16,777,216}{131,072} = 128\times \text{ fewer parameters (99.2% reduction)}$$

Multiply that across 32 transformer layers, and your trainable parameters drop from 8 billion down to ~20 million.

Two critical implementation details make LoRA stable and practical:

How do you prevent random noise from wrecking the pre-trained model at the start of training?

At step 0 of training:

$$\Delta W = B \cdot A = 0 \cdot A = 0$$

The adapter initially produces an exact zero vector. The model starts training with the exact output of the pre-trained base model, ensuring smooth and stable gradient descent.

The modified forward pass with LoRA includes a constant scaling multiplier $\frac{\alpha}{r}$:

$$h = W_0 x + \frac{\alpha}{r} (B \cdot A) x$$

Why does $\frac{\alpha}{r}$ exist?

When you experiment with different ranks (e.g., jumping from $r=8$ to $r=64$), the magnitude of the matrix multiplication $BAx$ naturally scales with $r$. The $\frac{\alpha}{r}$ term normalizes the magnitude of the adapter's update. This allows you to change the rank $r$ without needing to retune your learning rate from scratch.

During training, the computation branches into two parallel paths:

                      Input x
                     /       \
                    /         \
    [Frozen Base Weight]    [Matrix A (r × k)]  (Down-project to r)
          W₀ x                     |
            |                 [Matrix B (d × r)]  (Up-project to d)
            |                      |
            |               Scale by (α / r)
            \                     /
             \                   /
              +------> (+) <----+
                        |
                     Output h

During backpropagation:

One of the biggest advantages of LoRA over earlier adapter architectures (like Houlsby adapters or prompt tuning) is that LoRA introduces zero additional inference latency.

In production deployment, you do not keep two separate matrix multiplication branches in memory. Because matrix multiplication is distributive:

$$h = W_0 x + \Delta W x = (W_0 + \Delta W) x$$

Before exporting your model for deployment, you merge the adapter weights directly into the base weights with a single matrix addition:

$$W_{\text{serving}} = W_0 + \frac{\alpha}{r} (B \cdot A)$$

Once added, the adapter matrices $A$ and $B$ are discarded. The final serving model has the exact same architecture, tensor count, and latency as the original base model.

+-----------------------------------------------------------+
| Base Weights W₀           LoRA Weights (B × A)            |
| [ 4096 × 4096 ]     +     [ 4096 × 4096 ]                 |
+-----------------------------------------------------------+
                             ↓
+-----------------------------------------------------------+
| Merged Weights W_serving                                  |
| [ 4096 × 4096 ] → Standard inference speed, zero latency  |
+-----------------------------------------------------------+

If you need to switch tasks dynamically (multi-tenant serving), systems like S-LoRA or vLLM Multi-LoRA keep the base model frozen in VRAM and compute the small $BAx$ branch on the fly for distinct requests using custom batched GEMM kernels (like Punica).

To see how straightforward the mechanics are, here is a complete, working LoRA linear layer implemented in pure PyTorch:

import math
import torch
import torch.nn as nn

class LoRALinear(nn.Module):
    def __init__(
        self,
        in_features: int,
        out_features: int,
        r: int = 16,
        lora_alpha: float = 32.0,
        lora_dropout: float = 0.05
    ):
        super().__init__()
        self.base_layer = nn.Linear(in_features, out_features, bias=False)
        self.base_layer.weight.requires_grad = False

        self.r = r
        self.lora_alpha = lora_alpha
        self.scaling = lora_alpha / r

        if r > 0:
            self.lora_A = nn.Parameter(torch.empty(r, in_features))
            self.lora_B = nn.Parameter(torch.empty(out_features, r))
            self.dropout = nn.Dropout(p=lora_dropout) if lora_dropout > 0 else nn.Identity()

            nn.init.kaiming_uniform_(self.lora_A, a=math.sqrt(5))
            nn.init.zeros_(self.lora_B)

        self.merged = False

    def forward(self, x: torch.Tensor) -> torch.Tensor:
        if self.merged or self.r == 0:
            return self.base_layer(x)

        base_out = self.base_layer(x)
        lora_out = (self.dropout(x) @ self.lora_A.T) @ self.lora_B.T * self.scaling
        return base_out + lora_out

    def merge(self):
        """Fold LoRA weights into base weights for zero-overhead serving."""
        if self.r > 0 and not self.merged:
            delta_w = (self.lora_B @ self.lora_A) * self.scaling
            self.base_layer.weight.data += delta_w
            self.merged = True

    def unmerge(self):
        """Subtract LoRA weights to resume training or swap adapters."""
        if self.r > 0 and self.merged:
            delta_w = (self.lora_B @ self.lora_A) * self.scaling
            self.base_layer.weight.data -= delta_w
            self.merged = False

Here is how the memory footprint compares on an 8-billion parameter model (batch size 2, sequence length 2048):

Component Full Fine-Tuning (BF16) LoRA ($r=16$, All Linears) QLoRA (4-bit Base + LoRA)
Base Model Weights 16.0 GB (BF16) 16.0 GB (BF16) 4.5 GB (NF4)
Adapter Weights N/A 0.08 GB (BF16) 0.08 GB (BF16)
Gradients 16.0 GB 0.08 GB 0.08 GB
Optimizer States (AdamW) 96.0 GB 0.48 GB 0.48 GB
Activations (w/ checkpointing) ~6.0 GB ~6.0 GB ~6.0 GB
Total VRAM Needed ~134 GB ~22.6 GB ~11.1 GB
Hardware Required 2x A100 (80GB) 1x RTX 3090 / 4090 (24GB) 1x RTX 3060 (12GB)

The original 2021 LoRA paper demonstrated proof-of-concept results by attaching adapters only to Query and Value projections. However, multiple empirical studies (such as QLoRA by Dettmers et al.) prove that targeting all linear layers ($q, k, v, o$, as well as MLP layers $gate, up, down_proj$) yields significantly higher accuracy and expressive capacity, even with a smaller rank like $r=8$ or $r=16$.

A common convention is setting $\alpha = 2 \times r$ or $\alpha = r$. If you double your rank $r$ from 16 to 32 but leave $\alpha$ unchanged at 32, you have inadvertently halved the effective learning rate of your adapter ($\frac{\alpha}{r}$ drops from 2.0 to 1.0). Keep your $\frac{\alpha}{r}$ ratio consistent when sweeping rank values.

If you are deploying a single specialized model in production, leaving LoRA as separate branches adds unnecessary kernel launches and memory bandwidth overhead. Always call .merge_and_unload() in Hugging Face PEFT or manually fold weights before saving your deployment artifact.

Setting $r=128$ or $r=256$ rarely improves downstream task accuracy for instruction fine-tuning, but it drastically increases memory consumption and risk of catastrophic forgetting. For standard instruction tuning and classification, $r \in [8, 32]$ is almost always the empirical sweet spot.

── more in #machine-learning 4 stories · sorted by recency
── more on @lora 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-lora-actually-wo…] indexed:0 read:7min 2026-10-07 · —