Fine-tuning a 7B model needs 112 GB. The model is only 14 GB of it.
A developer's breakdown of fine-tuning memory shows that training a 7B-parameter model in mixed precision with Adam requires roughly 112 GB, of which only 14 GB is the fp16 weights — the optimizer sta…