cd /news/ai-safety/distillation-defenses-easily-break-a… · home › topics › ai-safety › article
[ARTICLE · art-142845] src=arxiv.org ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Distillation Defenses Easily Break After Reinforcement Learning

A September 28, 2026 arXiv paper (2609.35699) argues that distillation defenses for closed-source large language models are typically evaluated immediately after distillation and break once attackers apply reinforcement learning on top of the stolen traces. The authors show simple attacks can steal reasoning capabilities using data easily obtainable from current APIs, yielding reasoning improvements equivalent to more sophisticated attacks that extract full hidden traces, and conclude that any distillation defense leaking enough information to reconstruct approximate reasoning traces is likely ineffective. They propose batch-level distillation defenses as a potentially more effective alternative.

read2 min views1 publishedSep 30, 2026
Distillation Defenses Easily Break After Reinforcement Learning
Image: source
  [Submitted on 28 Sep 2026]


[View PDF](https://arxiv.org/pdf/2609.35699)

[HTML (experimental)](https://arxiv.org/html/2609.35699v1)

Abstract:Distillation attacks copy the reasoning capabilities of closed-source large language models, allowing bad actors to replicate state-of-the-art performance at low cost. Attackers systematically collect a large volume of frontier model reasoning traces and then train (i.e., "distill") their own models on these traces. Existing defenses against distillation attacks are typically evaluated immediately after distillation, implicitly assuming attackers do not train their models any further. In this paper, we argue that a more realistic threat model includes further training with reinforcement learning after distillation. A misspecified threat model can give a false sense of security -- some defenses that seem effective after distillation can be broken after subsequent reinforcement learning. Practically, reinforcement learning lowers the bar for a distillation attack to be effective. We show that simple attacks can steal reasoning capabilities from existing closed-source language models using data easily obtainable from current APIs, yielding reasoning improvements equivalent to more sophisticated attacks that extract the full hidden traces. Results indicate that any distillation defense that leaks sufficient information to reconstruct approximate reasoning traces is likely ineffective. We conclude by discussing broader implications and batch-level distillation defenses which could be more effective.

Current browse context:

cs.LG

References & Citations

...

Bibliographic Explorer

(What is the Explorer?) Connected Papers

(What is Connected Papers?) Litmaps

(What is Litmaps?) scite Smart Citations

(What are Smart Citations?) alphaXiv

(What is alphaXiv?) CatalyzeX Code Finder for Papers

(What is CatalyzeX?) DagsHub

(What is DagsHub?) Gotit.pub

(What is GotitPub?) Hugging Face

(What is Huggingface?) ScienceCast

(What is ScienceCast?) Influence Flower

(What are Influence Flowers?) CORE Recommender

(What is CORE?) IArxiv Recommender

(What is IArxiv?) arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

── more in #ai-safety 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/distillation-defense…] indexed:0 read:2min 2026-09-30 · —