Stabilizing language models under continual learning via condition-anchored distillation
Condition-anchored generative distillation (CAGD) reduced four-task final held-out loss in a 219M masked diffusion language model from 2.927 to 1.114 in one task order and from 2.168 to 0.891 in exact…