10:05
2026-08-10
huggingface.co
machine-learning
Making Knowledge Distillation Cheap Enough to Run at Scale
A new paper from Hugging Face, 'Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss,' cuts the VRAM required for knowledge distillation from roughly 250GB to neβ¦