{"slug": "how-lora-actually-works-low-rank-decomposition-weight-merging-and-memory-under", "title": "How LoRA Actually Works: Low-Rank Decomposition, Weight Merging, and Memory Breakdown Under the Hood", "summary": "A developer's technical breakdown explains why full-parameter fine-tuning of an 8-billion-parameter model in 16-bit precision exceeds 80 GB of VRAM, since AdamW training consumes roughly 16 bytes per trainable parameter across weights, gradients, master weights, and two optimizer moments. The writeup details how Low-Rank Adaptation freezes base weights and decomposes the update matrix into two small matrices, cutting trainable parameters by about 128x per layer at rank 16 and reducing memory footprint by over 70% while retaining 99% of full fine-tuning performance.", "body_md": "If you try to full-parameter fine-tune an 8-billion parameter language model in 16-bit precision, your GPU memory requirement immediately explodes past 80 GB.\n\nThe raw model weights only occupy 16 GB of VRAM ($8 \\times 10^9 \\text{ parameters} \\times 2 \\text{ bytes}$). Yet standard training crashes on a single 80 GB A100 or H100.\n\nWhere does the missing 64+ GB of VRAM go? And why can Low-Rank Adaptation (LoRA) reduce that memory footprint by over 70% while updating fewer than 0.1% of the parameters and retaining 99% of full fine-tuning performance?\n\nHere is what actually happens mathematically, mechanically, and in physical GPU memory during LoRA fine-tuning and inference.\n\nWhen training a neural network with the standard AdamW optimizer, model weights are only a tiny fraction of your memory consumption.\n\nFor every single trainable parameter, the system must track:\n\n$$\\text{Total Memory per Trainable Parameter} = 2 + 2 + 4 + 4 + 4 = 16 \\text{ bytes}$$\n\n```\n+-------------------------------------------------------------+\n| Full Parameter Training (16 bytes per parameter)            |\n+-------------------------------------------------------------+\n| Weights: 2B | Grads: 2B | Master: 4B | Mom1: 4B | Mom2: 4B  |\n+-------------------------------------------------------------+\n```\n\nFor an 8B parameter model:\n\nThis is why full fine-tuning requires multi-GPU distributed clusters (DeepSpeed ZeRO or FSDP).\n\nLoRA solves this problem at its root: by freezing the base weights, **it eliminates optimizer states and gradients for 99.9% of the model**.\n\nWhy does freezing the model work?\n\nA standard linear layer in a transformer computes:\n\n$$h = W_0 x$$\n\nwhere $W_0 \\in \\mathbb{R}^{d \\times k}$, $x \\in \\mathbb{R}^{k}$, and $h \\in \\mathbb{R}^{d}$. For Llama 3 8B, the hidden dimension $d = 4096$. Each weight matrix has up to $4096 \\times 4096 \\approx 16.7\\text{M}$ parameters with full mathematical rank 4096.\n\nIn 2021, Edward Hu et al. published the foundational hypothesis behind LoRA:\n\nWhen adapting a pre-trained language model to a specific task or instruction format, the weight update matrix $\\Delta W$ has an extremely low \"intrinsic rank\" ($r \\ll d$).\n\nThe pre-trained model already knows syntax, world facts, and language logic. Fine-tuning is merely steering attention and adjusting feature selection. You do not need to update 16 million degrees of freedom per matrix; you only need to update a tiny subspace of dimension $r$ (typically $r = 8, 16, \\text{ or } 32$).\n\nInstead of directly learning a full $d \\times k$ matrix $\\Delta W$, LoRA decomposes $\\Delta W$ into the product of two small, low-rank matrices:\n\n$$\\Delta W = B \\cdot A$$\n\nwhere:\n\n```\n   Full Update Matrix ΔW (d × k)            LoRA Decomposition\n   +---------------------------+       +-------+\n   |                           |       |       |       +---------------------------+\n d |                           |   = d |   B   |   × r |             A             |\n   |                           |       | (d×r) |       +---------------------------+\n   +---------------------------+       +-------+                     k\n                 k\n```\n\nSuppose $d = 4096, k = 4096$, and you choose rank $r = 16$:\n\n$$\\text{Reduction Factor} = \\frac{16,777,216}{131,072} = 128\\times \\text{ fewer parameters (99.2% reduction)}$$\n\nMultiply that across 32 transformer layers, and your trainable parameters drop from 8 billion down to ~20 million.\n\nTwo critical implementation details make LoRA stable and practical:\n\nHow do you prevent random noise from wrecking the pre-trained model at the start of training?\n\nAt step 0 of training:\n\n$$\\Delta W = B \\cdot A = 0 \\cdot A = 0$$\n\nThe adapter initially produces an exact zero vector. The model starts training with the exact output of the pre-trained base model, ensuring smooth and stable gradient descent.\n\nThe modified forward pass with LoRA includes a constant scaling multiplier $\\frac{\\alpha}{r}$:\n\n$$h = W_0 x + \\frac{\\alpha}{r} (B \\cdot A) x$$\n\nWhy does $\\frac{\\alpha}{r}$ exist?\n\nWhen you experiment with different ranks (e.g., jumping from $r=8$ to $r=64$), the magnitude of the matrix multiplication $BAx$ naturally scales with $r$. The $\\frac{\\alpha}{r}$ term normalizes the magnitude of the adapter's update. This allows you to change the rank $r$ without needing to retune your learning rate from scratch.\n\nDuring training, the computation branches into two parallel paths:\n\n```\n                      Input x\n                     /       \\\n                    /         \\\n    [Frozen Base Weight]    [Matrix A (r × k)]  (Down-project to r)\n          W₀ x                     |\n            |                 [Matrix B (d × r)]  (Up-project to d)\n            |                      |\n            |               Scale by (α / r)\n            \\                     /\n             \\                   /\n              +------> (+) <----+\n                        |\n                     Output h\n```\n\nDuring backpropagation:\n\nOne of the biggest advantages of LoRA over earlier adapter architectures (like Houlsby adapters or prompt tuning) is that **LoRA introduces zero additional inference latency**.\n\nIn production deployment, you do not keep two separate matrix multiplication branches in memory. Because matrix multiplication is distributive:\n\n$$h = W_0 x + \\Delta W x = (W_0 + \\Delta W) x$$\n\nBefore exporting your model for deployment, you merge the adapter weights directly into the base weights with a single matrix addition:\n\n$$W_{\\text{serving}} = W_0 + \\frac{\\alpha}{r} (B \\cdot A)$$\n\nOnce added, the adapter matrices $A$ and $B$ are discarded. The final serving model has the exact same architecture, tensor count, and latency as the original base model.\n\n```\n+-----------------------------------------------------------+\n| Base Weights W₀           LoRA Weights (B × A)            |\n| [ 4096 × 4096 ]     +     [ 4096 × 4096 ]                 |\n+-----------------------------------------------------------+\n                             ↓\n+-----------------------------------------------------------+\n| Merged Weights W_serving                                  |\n| [ 4096 × 4096 ] → Standard inference speed, zero latency  |\n+-----------------------------------------------------------+\n```\n\nIf you need to switch tasks dynamically (multi-tenant serving), systems like **S-LoRA** or **vLLM Multi-LoRA** keep the base model frozen in VRAM and compute the small $BAx$ branch on the fly for distinct requests using custom batched GEMM kernels (like Punica).\n\nTo see how straightforward the mechanics are, here is a complete, working LoRA linear layer implemented in pure PyTorch:\n\n``` python\nimport math\nimport torch\nimport torch.nn as nn\n\nclass LoRALinear(nn.Module):\n    def __init__(\n        self,\n        in_features: int,\n        out_features: int,\n        r: int = 16,\n        lora_alpha: float = 32.0,\n        lora_dropout: float = 0.05\n    ):\n        super().__init__()\n        # 1. Base pre-trained linear layer (frozen)\n        self.base_layer = nn.Linear(in_features, out_features, bias=False)\n        self.base_layer.weight.requires_grad = False\n\n        self.r = r\n        self.lora_alpha = lora_alpha\n        self.scaling = lora_alpha / r\n\n        if r > 0:\n            # 2. Low-rank matrices A and B\n            self.lora_A = nn.Parameter(torch.empty(r, in_features))\n            self.lora_B = nn.Parameter(torch.empty(out_features, r))\n            self.dropout = nn.Dropout(p=lora_dropout) if lora_dropout > 0 else nn.Identity()\n\n            # 3. Initialization: Kaiming uniform for A, exact zeros for B\n            nn.init.kaiming_uniform_(self.lora_A, a=math.sqrt(5))\n            nn.init.zeros_(self.lora_B)\n\n        self.merged = False\n\n    def forward(self, x: torch.Tensor) -> torch.Tensor:\n        if self.merged or self.r == 0:\n            return self.base_layer(x)\n\n        # Standard LoRA forward pass: W₀x + (α/r) * B(A(dropout(x)))\n        base_out = self.base_layer(x)\n        lora_out = (self.dropout(x) @ self.lora_A.T) @ self.lora_B.T * self.scaling\n        return base_out + lora_out\n\n    def merge(self):\n        \"\"\"Fold LoRA weights into base weights for zero-overhead serving.\"\"\"\n        if self.r > 0 and not self.merged:\n            # Compute ΔW = (α/r) * B @ A\n            delta_w = (self.lora_B @ self.lora_A) * self.scaling\n            self.base_layer.weight.data += delta_w\n            self.merged = True\n\n    def unmerge(self):\n        \"\"\"Subtract LoRA weights to resume training or swap adapters.\"\"\"\n        if self.r > 0 and self.merged:\n            delta_w = (self.lora_B @ self.lora_A) * self.scaling\n            self.base_layer.weight.data -= delta_w\n            self.merged = False\n```\n\nHere is how the memory footprint compares on an 8-billion parameter model (batch size 2, sequence length 2048):\n\n| Component | Full Fine-Tuning (BF16) | LoRA ($r=16$, All Linears) | QLoRA (4-bit Base + LoRA) | \n|---|---|---|---|\n| **Base Model Weights** | 16.0 GB (BF16) | 16.0 GB (BF16) | 4.5 GB (NF4) | \n| **Adapter Weights** | N/A | 0.08 GB (BF16) | 0.08 GB (BF16) | \n| **Gradients** | 16.0 GB | 0.08 GB | 0.08 GB | \n| **Optimizer States (AdamW)** | 96.0 GB | 0.48 GB | 0.48 GB | \n| **Activations (w/ checkpointing)** | ~6.0 GB | ~6.0 GB | ~6.0 GB | \n| **Total VRAM Needed** | **~134 GB** | **~22.6 GB** | **~11.1 GB** | \n| **Hardware Required** | 2x A100 (80GB) | 1x RTX 3090 / 4090 (24GB) | 1x RTX 3060 (12GB) | \n\nThe original 2021 LoRA paper demonstrated proof-of-concept results by attaching adapters only to Query and Value projections. However, multiple empirical studies (such as QLoRA by Dettmers et al.) prove that targeting **all linear layers** ($q, k, v, o$, as well as MLP layers $gate, up, down_proj$) yields significantly higher accuracy and expressive capacity, even with a smaller rank like $r=8$ or $r=16$.\n\nA common convention is setting $\\alpha = 2 \\times r$ or $\\alpha = r$. If you double your rank $r$ from 16 to 32 but leave $\\alpha$ unchanged at 32, you have inadvertently halved the effective learning rate of your adapter ($\\frac{\\alpha}{r}$ drops from 2.0 to 1.0). Keep your $\\frac{\\alpha}{r}$ ratio consistent when sweeping rank values.\n\nIf you are deploying a single specialized model in production, leaving LoRA as separate branches adds unnecessary kernel launches and memory bandwidth overhead. Always call `.merge_and_unload()` in Hugging Face PEFT or manually fold weights before saving your deployment artifact.\n\nSetting $r=128$ or $r=256$ rarely improves downstream task accuracy for instruction fine-tuning, but it drastically increases memory consumption and risk of catastrophic forgetting. For standard instruction tuning and classification, $r \\in [8, 32]$ is almost always the empirical sweet spot.", "url": "https://wpnews.pro/news/how-lora-actually-works-low-rank-decomposition-weight-merging-and-memory-under", "canonical_source": "https://dev.to/syed_anzar/how-lora-actually-works-low-rank-decomposition-weight-merging-and-memory-breakdown-under-the-hood-2b18", "published_at": "2026-10-07 12:35:30+00:00", "updated_at": "2026-10-07 12:47:26.123171+00:00", "lang": "en", "topics": ["machine-learning", "large-language-models", "ai-research", "ai-infrastructure", "mlops"], "entities": ["LoRA", "AdamW", "Llama 3 8B", "Edward Hu", "DeepSpeed ZeRO", "FSDP", "A100", "H100"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-lora-actually-works-low-rank-decomposition-weight-merging-and-memory-under", "markdown": "https://wpnews.pro/news/how-lora-actually-works-low-rank-decomposition-weight-merging-and-memory-under.md", "text": "https://wpnews.pro/news/how-lora-actually-works-low-rank-decomposition-weight-merging-and-memory-under.txt", "jsonld": "https://wpnews.pro/news/how-lora-actually-works-low-rank-decomposition-weight-merging-and-memory-under.jsonld"}}