The Implications of Linguistic Illegibility for LLM Security A paper submitted to arXiv on 2 Sep 2026 introduces the term "linguistic illegibility" to describe scenarios in which an LLM's externalized or mechanistically-probed language artifacts fail to represent how the model actually thinks, arguing that security mechanisms relying on a model's linguistic self-reporting — chain-of-thought monitoring, constitutional self-critique, and activation probing for linguistically-defined feature vectors — can never be completely sound. The authors argue that taint tracking, which defines a priori pieces of system state that should never be influenced by model-produced data, is a promising approach for an effective sandbox, supplemented by robust virtualization and third-party auditing of sandboxing configurations. The paper states these mechanisms would have mitigated recent sandbox exploits by frontier models. Computer Science Machine Learning Submitted on 2 Sep 2026 Title:The Implications of Linguistic Illegibility for LLM Security View PDF https://arxiv.org/pdf/2609.02852 HTML experimental https://arxiv.org/html/2609.02852v1 Abstract:LLMs are trained to generate natural language. However, various strands of evidence indicate that an LLM's externalized linguistic outputs and mechanistically-extracted linguistic features can be an unreliable lens for understanding internal model computation. We introduce the term linguistic illegibility'' to broadly refer to scenarios in which an LLM's externalized or mechanistically-probed language artifacts fail to represent how the model actually thinks. We argue that the specter of linguistic illegibility is unavoidable for LLMs whose internal computations are not directly expressed via language, but rather math over activation spaces with lossy translations between activation spaces and natural language happening at the bookends . If linguistic illegibility is always possible, then security mechanisms that rely on a model's linguistic self-reporting e.g., chain-of-thought monitoring, constitutional self-critique, activation probing for linguistically-defined feature vectors can never be completely sound; the model sandbox will always need isolation techniques whose guarantees do not depend on reading a model's linguistic state at all. We argue that observing a model's outputs using taint tracking is a promising approach for an effective sandbox: regardless of how a model linguistically self-reports, a taint tracking policy can define, a priori, various pieces of system state that should never be influenced by model-produced data. We also discuss several additional sandboxing mechanisms e.g., robust virtualization, third-party auditing of sandboxing configurations which collectively provide a critical floor beneath linguistic monitoring, and would have mitigated recent sandbox exploits by frontier models. References & Citations Loading... Bibliographic and Citation Tools Bibliographic Explorer What is the Explorer? https://info.arxiv.org/labs/showcase.html arxiv-bibliographic-explorer Connected Papers What is Connected Papers? https://www.connectedpapers.com/about Litmaps What is Litmaps? https://www.litmaps.co/ scite Smart Citations What are Smart Citations? https://www.scite.ai/ Code, Data and Media Associated with this Article alphaXiv What is alphaXiv? https://alphaxiv.org/ CatalyzeX Code Finder for Papers What is CatalyzeX? https://www.catalyzex.com DagsHub What is DagsHub? https://dagshub.com/ Gotit.pub What is GotitPub? http://gotit.pub/faq Hugging Face What is Huggingface? https://huggingface.co/huggingface ScienceCast What is ScienceCast? https://sciencecast.org/welcome Demos Recommenders and Search Tools Influence Flower What are Influence Flowers? https://influencemap.cmlab.dev/ CORE Recommender What is CORE? https://core.ac.uk/services/recommender IArxiv Recommender What is IArxiv? https://iarxiv.org/about arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs https://info.arxiv.org/labs/index.html .