Chunked KL loss, running Knowledge Distillation locally in less <6GB VRAM
A new CUDA benchmark compares three knowledge-distillation loss implementations—Full Dense KL, Forward-Chunked Loss, and Full Chunked KL—showing that the Full Chunked KL method fuses the output projec…