For people who run and fine-tune smaller models locally (1B to 7B range), what VRAM have you found to be the practical minimum on a laptop GPU? Most laptop GPUs I see listed have 8GB, and some newer ones have 12GB, and I’m not sure how far that goes with quantization and LoRA. Also, is it usually better to do the training in the cloud and just use the laptop for inference and testing? Would like to hear what setups have actually worked for you.
Trying to improve output speed with MLX models on LM Studio