{"slug": "distillation-defenses-easily-break-after-reinforcement-learning", "title": "Distillation Defenses Easily Break After Reinforcement Learning", "summary": "A September 28, 2026 arXiv paper (2609.35699) argues that distillation defenses for closed-source large language models are typically evaluated immediately after distillation and break once attackers apply reinforcement learning on top of the stolen traces. The authors show simple attacks can steal reasoning capabilities using data easily obtainable from current APIs, yielding reasoning improvements equivalent to more sophisticated attacks that extract full hidden traces, and conclude that any distillation defense leaking enough information to reconstruct approximate reasoning traces is likely ineffective. They propose batch-level distillation defenses as a potentially more effective alternative.", "body_md": "# Computer Science > Machine Learning\n\n  [Submitted on 28 Sep 2026]\n\n# Title:Distillation Defenses Easily Break After Reinforcement Learning\n\n[View PDF](https://arxiv.org/pdf/2609.35699)\n\n[HTML (experimental)](https://arxiv.org/html/2609.35699v1)\n\nAbstract:Distillation attacks copy the reasoning capabilities of closed-source large language models, allowing bad actors to replicate state-of-the-art performance at low cost. Attackers systematically collect a large volume of frontier model reasoning traces and then train (i.e., \"distill\") their own models on these traces. Existing defenses against distillation attacks are typically evaluated immediately after distillation, implicitly assuming attackers do not train their models any further. In this paper, we argue that a more realistic threat model includes further training with reinforcement learning after distillation. A misspecified threat model can give a false sense of security -- some defenses that seem effective after distillation can be broken after subsequent reinforcement learning. Practically, reinforcement learning lowers the bar for a distillation attack to be effective. We show that simple attacks can steal reasoning capabilities from existing closed-source language models using data easily obtainable from current APIs, yielding reasoning improvements equivalent to more sophisticated attacks that extract the full hidden traces. Results indicate that any distillation defense that leaks sufficient information to reconstruct approximate reasoning traces is likely ineffective. We conclude by discussing broader implications and batch-level distillation defenses which could be more effective.\n    \n\n### Current browse context:\n\ncs.LG\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\nIArxiv Recommender\n\n*(*[What is IArxiv?](https://iarxiv.org/about))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/distillation-defenses-easily-break-after-reinforcement-learning", "canonical_source": "https://arxiv.org/abs/2609.35699", "published_at": "2026-09-30 21:58:38+00:00", "updated_at": "2026-09-30 22:18:32.529320+00:00", "lang": "en", "topics": ["ai-safety", "large-language-models", "ai-research", "machine-learning", "artificial-intelligence"], "entities": ["arXiv", "cs.LG"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/distillation-defenses-easily-break-after-reinforcement-learning", "markdown": "https://wpnews.pro/news/distillation-defenses-easily-break-after-reinforcement-learning.md", "text": "https://wpnews.pro/news/distillation-defenses-easily-break-after-reinforcement-learning.txt", "jsonld": "https://wpnews.pro/news/distillation-defenses-easily-break-after-reinforcement-learning.jsonld"}}