{"slug": "stabilizing-language-models-under-continual-learning-via-condition-anchored", "title": "Stabilizing language models under continual learning via condition-anchored distillation", "summary": "Condition-anchored generative distillation (CAGD) reduced four-task final held-out loss in a 219M masked diffusion language model from 2.927 to 1.114 in one task order and from 2.168 to 0.891 in exact reverse, according to an arXiv paper (2610.06940v1). The method retains a small set of old prompts, uses a frozen previous model to reconstruct completions and generation states, and matches its predictive distributions while learning the next task; identical teacher-generated support gave soft targets a 0.055 lower final average loss than hard replay. The direction persisted on fresh facts and natural instructions across SMDM and Qwen3, though on GSM8K exact-match retention for Qwen was seed-mixed at 0.6B and worsened at 1.7B.", "body_md": "arXiv:2610.06940v1 Announce Type: new \nAbstract: Continual adaptation of language models can change their output distribution on prompts learned earlier, while retaining every old prompt-answer pair may be undesirable or impossible. We study condition-anchored generative distillation (CAGD): retain a small set of old prompts, use a frozen previous model to reconstruct completions and generation states, and match its predictive distributions while learning the next task. The formulation separates three roles that ordinary replay conflates: conditions select the behavior to protect, teacher generations locate relevant states, and soft targets specify how predictions may change. For autoregressive language generation, teacher-rollout distillation admits an exact chain-rule decomposition of sequence divergence. For masked-diffusion language modeling, our implementation directly controls local denoising drift on teacher-generated completions. In continual adaptation of a 219M masked diffusion language model, CAGD reduces four-task final held-out loss from 2.927 to 1.114 in one task order and from 2.168 to 0.891 in exact reverse. The same soft targets lower final average loss by 0.055 over hard replay when teacher-generated support is held identical. The direction persists on fresh facts and natural instructions across SMDM and Qwen3. On GSM8K, Qwen adaptation preserves answer-format compliance, but exact-match retention is seed-mixed at 0.6B and worsens at 1.7B. These results support condition-anchored functional preservation as a common design principle across the tested language-generation objectives.", "url": "https://wpnews.pro/news/stabilizing-language-models-under-continual-learning-via-condition-anchored", "canonical_source": "https://arxiv.org/abs/2610.06940", "published_at": "2026-10-07 04:00:00+00:00", "updated_at": "2026-10-07 04:19:30.685946+00:00", "lang": "en", "topics": ["large-language-models", "machine-learning", "ai-research", "natural-language-processing"], "entities": ["CAGD", "SMDM", "Qwen3", "GSM8K", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/stabilizing-language-models-under-continual-learning-via-condition-anchored", "markdown": "https://wpnews.pro/news/stabilizing-language-models-under-continual-learning-via-condition-anchored.md", "text": "https://wpnews.pro/news/stabilizing-language-models-under-continual-learning-via-condition-anchored.txt", "jsonld": "https://wpnews.pro/news/stabilizing-language-models-under-continual-learning-via-condition-anchored.jsonld"}}