{"slug": "the-implications-of-linguistic-illegibility-for-llm-security", "title": "The Implications of Linguistic Illegibility for LLM Security", "summary": "A paper submitted to arXiv on 2 Sep 2026 introduces the term \"linguistic illegibility\" to describe scenarios in which an LLM's externalized or mechanistically-probed language artifacts fail to represent how the model actually thinks, arguing that security mechanisms relying on a model's linguistic self-reporting — chain-of-thought monitoring, constitutional self-critique, and activation probing for linguistically-defined feature vectors — can never be completely sound. The authors argue that taint tracking, which defines a priori pieces of system state that should never be influenced by model-produced data, is a promising approach for an effective sandbox, supplemented by robust virtualization and third-party auditing of sandboxing configurations. The paper states these mechanisms would have mitigated recent sandbox exploits by frontier models.", "body_md": "# Computer Science > Machine Learning\n\n  [Submitted on 2 Sep 2026]\n\n# Title:The Implications of Linguistic Illegibility for LLM Security\n\n[View PDF](https://arxiv.org/pdf/2609.02852)\n\n[HTML (experimental)](https://arxiv.org/html/2609.02852v1)\n\nAbstract:LLMs are trained to generate natural language. However, various strands of evidence indicate that an LLM's externalized linguistic outputs and mechanistically-extracted linguistic features can be an unreliable lens for understanding internal model computation. We introduce the term ``linguistic illegibility'' to broadly refer to scenarios in which an LLM's externalized or mechanistically-probed language artifacts fail to represent how the model actually thinks. We argue that the specter of linguistic illegibility is unavoidable for LLMs whose internal computations are not directly expressed via language, but rather math over activation spaces (with lossy translations between activation spaces and natural language happening at the bookends). If linguistic illegibility is always possible, then security mechanisms that rely on a model's linguistic self-reporting (e.g., chain-of-thought monitoring, constitutional self-critique, activation probing for linguistically-defined feature vectors) can never be completely sound; the model sandbox will always need isolation techniques whose guarantees do not depend on reading a model's linguistic state at all. We argue that observing a model's outputs using taint tracking is a promising approach for an effective sandbox: regardless of how a model linguistically self-reports, a taint tracking policy can define, a priori, various pieces of system state that should never be influenced by model-produced data. We also discuss several additional sandboxing mechanisms (e.g., robust virtualization, third-party auditing of sandboxing configurations) which collectively provide a critical floor beneath linguistic monitoring, and would have mitigated recent sandbox exploits by frontier models.\n    \n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\nIArxiv Recommender\n\n*(*[What is IArxiv?](https://iarxiv.org/about))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/the-implications-of-linguistic-illegibility-for-llm-security", "canonical_source": "https://arxiv.org/abs/2609.02852", "published_at": "2026-09-18 19:00:06+00:00", "updated_at": "2026-09-18 19:24:16.715138+00:00", "lang": "en", "topics": ["ai-safety", "large-language-models", "machine-learning", "ai-research"], "entities": ["arXiv"], "alternates": {"html": "https://wpnews.pro/news/the-implications-of-linguistic-illegibility-for-llm-security", "markdown": "https://wpnews.pro/news/the-implications-of-linguistic-illegibility-for-llm-security.md", "text": "https://wpnews.pro/news/the-implications-of-linguistic-illegibility-for-llm-security.txt", "jsonld": "https://wpnews.pro/news/the-implications-of-linguistic-illegibility-for-llm-security.jsonld"}}