Unsloth vs Axolotl vs TRL: 87% of Your Fine-Tuning VRAM Goes to a Tensor You Never Wrote
A first-principles analysis of LoRA fine-tuning memory usage reveals that 87.3% of VRAM scaling with sequence length for Llama 3.1 8B is consumed by the cross-entropy loss head tensor, not the model, …