cd /news/large-language-models/stabilizing-language-models-under-co… · home › topics › large-language-models › article
[ARTICLE · art-146568] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Stabilizing language models under continual learning via condition-anchored distillation

Condition-anchored generative distillation (CAGD) reduced four-task final held-out loss in a 219M masked diffusion language model from 2.927 to 1.114 in one task order and from 2.168 to 0.891 in exact reverse, according to an arXiv paper (2610.06940v1). The method retains a small set of old prompts, uses a frozen previous model to reconstruct completions and generation states, and matches its predictive distributions while learning the next task; identical teacher-generated support gave soft targets a 0.055 lower final average loss than hard replay. The direction persisted on fresh facts and natural instructions across SMDM and Qwen3, though on GSM8K exact-match retention for Qwen was seed-mixed at 0.6B and worsened at 1.7B.

by read1 min views1 publishedOct 7, 2026

arXiv:2610.06940v1 Announce Type: new Abstract: Continual adaptation of language models can change their output distribution on prompts learned earlier, while retaining every old prompt-answer pair may be undesirable or impossible. We study condition-anchored generative distillation (CAGD): retain a small set of old prompts, use a frozen previous model to reconstruct completions and generation states, and match its predictive distributions while learning the next task. The formulation separates three roles that ordinary replay conflates: conditions select the behavior to protect, teacher generations locate relevant states, and soft targets specify how predictions may change. For autoregressive language generation, teacher-rollout distillation admits an exact chain-rule decomposition of sequence divergence. For masked-diffusion language modeling, our implementation directly controls local denoising drift on teacher-generated completions. In continual adaptation of a 219M masked diffusion language model, CAGD reduces four-task final held-out loss from 2.927 to 1.114 in one task order and from 2.168 to 0.891 in exact reverse. The same soft targets lower final average loss by 0.055 over hard replay when teacher-generated support is held identical. The direction persists on fresh facts and natural instructions across SMDM and Qwen3. On GSM8K, Qwen adaptation preserves answer-format compliance, but exact-match retention is seed-mixed at 0.6B and worsens at 1.7B. These results support condition-anchored functional preservation as a common design principle across the tested language-generation objectives.

── more in #large-language-models 4 stories · sorted by recency
── more on @cagd 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/stabilizing-language…] indexed:0 read:1min 2026-10-07 · —