cd /news/artificial-intelligence/activation-conditioned-self-distilla… · home › topics › artificial-intelligence › article
[ARTICLE · art-143640] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Activation-Conditioned Self-Distillation

A new arXiv paper (2609.38342v1) introduces Activation-Conditioned Self-Distillation (ACSD), a method that extracts a steering vector by contrasting activations of self-generated trajectories reaching verified correct answers within a generation budget against all remaining trajectories, then applies that vector from a frozen copy of the base model at each prediction position. On DeepSeek-R1-0528-Qwen3-8B, ACSD reaches 71.9% mean mathematical accuracy and 70.9% LiveCodeBench v6 pass@12, versus 69.0% and 66.3% for the reference-conditioned OPSD baseline, and it achieves the highest mean accuracy across four mathematical benchmarks on each of five models. ACSD requires neither problem-specific reference text nor teacher parameter updates, and the distilled student is used alone at inference.

by read1 min views1 publishedOct 2, 2026

arXiv:2609.38342v1 Announce Type: new Abstract: On-policy self-distillation uses a model as its own teacher to provide dense supervision for reasoning, often through reference-solution conditioning. Providing privileged information does not by itself ensure effective token-level supervision throughout long responses. We introduce Activation-Conditioned Self-Distillation (ACSD), which extracts a steering vector by contrasting activations of self-generated trajectories that reach verified correct answers within a generation budget with those of all remaining trajectories. A frozen copy of the base model applies this vector at each prediction position, and the student learns from its next-token distributions on student-generated prefixes. Outcome verification is used for direction construction and calibration; distillation requires neither problem-specific reference text nor teacher parameter updates. The distilled student is used alone at inference. On each of five models, ACSD achieves the highest mean accuracy over four mathematical benchmarks among the evaluated methods. On DeepSeek-R1-0528-Qwen3-8B, mean mathematical accuracy reaches 71.9% and LiveCodeBench v6 pass@12 reaches 70.9%, compared with 69.0% and 66.3% for the reference-conditioned OPSD baseline. Contrasts among correct trajectories also support distillation, and extracted directions can be reused across mathematical training datasets. On fixed student trajectories, ACSD maintains more stable late-position logit-update magnitudes than OPSD.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @activation-conditioned self-distillation 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/activation-condition…] indexed:0 read:1min 2026-10-02 · —