{"slug": "emergent-misalignment-recruits-a-pre-existing-persona-subspace", "title": "Emergent Misalignment Recruits a Pre-Existing Persona Subspace", "summary": "A new study from researchers at an undisclosed institution finds that fine-tuning an aligned language model on a narrow stream of bad advice recruits a pre-existing persona subspace, causing broad misalignment on unrelated questions. Using the Qwen2.5-14B-Instruct model, the team extracted per-domain persona subspaces and found that four unrelated domains share a low-rank core at 657 times a random-subspace null, with 82% of that core lying outside a style core. Projecting the subspace out of the residual stream throughout fine-tuning prevented broad misalignment (27.7% to 0.0% of judged generations), while injecting it into the never-fine-tuned model induced misalignment that grew with dose to 45.4%.", "body_md": "# Computer Science > Machine Learning\n\n[Submitted on 23 Jul 2026]\n\n# Title:Emergent Misalignment Recruits a Pre-existing Persona Subspace\n\n[View PDF](/pdf/2607.21356)\n\n[HTML (experimental)](https://arxiv.org/html/2607.21356v1)\n\nAbstract:Fine-tuning an aligned language model on a narrow stream of bad advice can make it broadly misaligned on questions unrelated to the training data, a phenomenon called emergent misalignment. We ask why the narrow lesson generalizes at all, and we find that narrow fine-tuning recruits a persona structure that is present in the model before the fine-tune exists. From a frozen instruction-tuned model (Qwen2.5-14B-Instruct) we extract per-domain persona subspaces by contrastive teacher forcing and find that 4 unrelated domains share one low-rank core at 657x a random-subspace null, with 82% of that core lying outside a style core built at matched diversity. The literal first optimizer step of fine-tuning on insecure code climbs a broad-misalignment margin harder than the same code framed as educational, and forecasts realized margin movement out to 375 steps. Projecting the subspace out of the residual stream throughout fine-tuning prevents broad misalignment (27.7% to 0.0% of judged generations) while a matched-rank random subspace changes nothing; injecting it into the never-fine-tuned model induces misalignment that grows with dose to 45.4%, past the fine-tuned model it is measured against. The same projection applied to the weight gradient is inert, and three post-hoc weight edits leave the disposition in place: the sharpest edit suppresses the behavior rather than removing it, and the ablated structure re-forms inside the subspace the edit cleared. Spreading a fixed budget of bad data across 4 domains produces more broad misalignment than mechanical weight superposition and matched diversity jointly account for. All measurements come from one model at 14B; the extraction is from an aligned instruction-tuned checkpoint, which leaves the structure's provenance open; and the intervention that prevents misalignment also abolishes the narrow trained behavior.\n\n## Submission history\n\nFrom: Mohammed Suhail B Nadaf [[view email](/show-email/7480247d/2607.21356)]\n\n**[v1]** Thu, 23 Jul 2026 14:19:28 UTC (864 KB)\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer\n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers\n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps\n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations\n\n*(*[What are Smart Citations?](https://www.scite.ai/))# Code, Data and Media Associated with this Article\n\nalphaXiv\n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers\n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub\n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub\n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face\n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast\n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower\n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender\n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\nIArxiv Recommender\n\n*(*[What is IArxiv?](https://iarxiv.org/about))# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [ Learn more about arXivLabs](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/emergent-misalignment-recruits-a-pre-existing-persona-subspace", "canonical_source": "https://arxiv.org/abs/2607.21356", "published_at": "2026-07-24 11:07:06+00:00", "updated_at": "2026-07-24 11:22:11.765441+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-safety", "large-language-models"], "entities": ["Qwen2.5-14B-Instruct"], "alternates": {"html": "https://wpnews.pro/news/emergent-misalignment-recruits-a-pre-existing-persona-subspace", "markdown": "https://wpnews.pro/news/emergent-misalignment-recruits-a-pre-existing-persona-subspace.md", "text": "https://wpnews.pro/news/emergent-misalignment-recruits-a-pre-existing-persona-subspace.txt", "jsonld": "https://wpnews.pro/news/emergent-misalignment-recruits-a-pre-existing-persona-subspace.jsonld"}}