cd /news/ai-safety/distillation-for-incrimination-and-d… · home › topics › ai-safety › article
[ARTICLE · art-148045] src=arxiv.org ↗ pub= topic=ai-safety verified=true sentiment=· neutral

Distillation for Incrimination and Distillation for Capabilities

A new arXiv paper (2610.11012v1) introduces two distillation methods for AI safety: Distillation for Incrimination (DFI), which transfers misalignment but not concealment ability, and Distillation for Capabilities (DFC), which transfers capabilities while reducing misalignment transfer. Distilling AuditBench's secret-keeping models into their underlying instruction-tuned models produced students significantly more likely than their teachers to admit hidden behavior when asked, though the confession gains largely disappeared when the student did not share the teacher's pretrained base. Among DFC techniques evaluated, inoculation prompting and training for more epochs on fewer unique examples both preserved standard distillation's capability gains while substantially reducing subliminal transfer of an animal preference used as a misalignment proxy.

by read1 min views1 publishedOct 9, 2026

arXiv:2610.11012v1 Announce Type: new Abstract: Powerful misaligned AI models might recognize alignment evaluations and strategically behave well on them, rendering direct audits uninformative. However, distilling such a model into a weaker benign student places the teacher in a Distillation Double Bind: if misalignment transfers, the student may conceal it less effectively, revealing evidence about the teacher; if it does not, the student may learn useful capabilities while remaining benign. We introduce two distinct distillation approaches, one targeting each outcome. Distillation for Incrimination (DFI) aims to transfer misalignment but not the ability to conceal it. Distilling AuditBench's secret-keeping models into their underlying instruction-tuned model produces students that are significantly more likely than their teachers to admit their hidden behavior when asked, suggesting that knowledge of the behavior transferred more readily than the propensity to conceal it. Confession gains largely disappear when the student does not share the teacher's pretrained base, so DFI should target the teacher's own pre-RL checkpoint, which is weaker than the teacher but shares its base model. Distillation for Capabilities (DFC) aims to transfer capabilities but not misalignment. Among several techniques we evaluate, two are effective: inoculation prompting and training for more epochs on fewer unique examples. Both preserve the capability gains of standard distillation while substantially reducing the subliminal transfer of an animal preference, our proxy for misalignment. Together, these findings demonstrate two ways distillation can be used for AI safety: incriminating misaligned models, and extracting their capabilities without their misalignment.

── more in #ai-safety 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/distillation-for-inc…] indexed:0 read:1min 2026-10-09 · —