cd /news/ai-safety/the-implications-of-linguistic-illeg… · home topics ai-safety article
[ARTICLE · art-134060] src=arxiv.org ↗ pub= topic=ai-safety verified=true sentiment=· neutral

The Implications of Linguistic Illegibility for LLM Security

A paper submitted to arXiv on 2 Sep 2026 introduces the term "linguistic illegibility" to describe scenarios in which an LLM's externalized or mechanistically-probed language artifacts fail to represent how the model actually thinks, arguing that security mechanisms relying on a model's linguistic self-reporting — chain-of-thought monitoring, constitutional self-critique, and activation probing for linguistically-defined feature vectors — can never be completely sound. The authors argue that taint tracking, which defines a priori pieces of system state that should never be influenced by model-produced data, is a promising approach for an effective sandbox, supplemented by robust virtualization and third-party auditing of sandboxing configurations. The paper states these mechanisms would have mitigated recent sandbox exploits by frontier models.

read2 min views1 publishedSep 18, 2026
The Implications of Linguistic Illegibility for LLM Security
Image: source
  [Submitted on 2 Sep 2026]


[View PDF](https://arxiv.org/pdf/2609.02852)

[HTML (experimental)](https://arxiv.org/html/2609.02852v1)

Abstract:LLMs are trained to generate natural language. However, various strands of evidence indicate that an LLM's externalized linguistic outputs and mechanistically-extracted linguistic features can be an unreliable lens for understanding internal model computation. We introduce the term ``linguistic illegibility'' to broadly refer to scenarios in which an LLM's externalized or mechanistically-probed language artifacts fail to represent how the model actually thinks. We argue that the specter of linguistic illegibility is unavoidable for LLMs whose internal computations are not directly expressed via language, but rather math over activation spaces (with lossy translations between activation spaces and natural language happening at the bookends). If linguistic illegibility is always possible, then security mechanisms that rely on a model's linguistic self-reporting (e.g., chain-of-thought monitoring, constitutional self-critique, activation probing for linguistically-defined feature vectors) can never be completely sound; the model sandbox will always need isolation techniques whose guarantees do not depend on reading a model's linguistic state at all. We argue that observing a model's outputs using taint tracking is a promising approach for an effective sandbox: regardless of how a model linguistically self-reports, a taint tracking policy can define, a priori, various pieces of system state that should never be influenced by model-produced data. We also discuss several additional sandboxing mechanisms (e.g., robust virtualization, third-party auditing of sandboxing configurations) which collectively provide a critical floor beneath linguistic monitoring, and would have mitigated recent sandbox exploits by frontier models.

References & Citations

...

Bibliographic Explorer

(What is the Explorer?) Connected Papers

(What is Connected Papers?) Litmaps

(What is Litmaps?) scite Smart Citations

(What are Smart Citations?) alphaXiv

(What is alphaXiv?) CatalyzeX Code Finder for Papers

(What is CatalyzeX?) DagsHub

(What is DagsHub?) Gotit.pub

(What is GotitPub?) Hugging Face

(What is Huggingface?) ScienceCast

(What is ScienceCast?) Influence Flower

(What are Influence Flowers?) CORE Recommender

(What is CORE?) IArxiv Recommender

(What is IArxiv?) arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

── more in #ai-safety 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-implications-of-…] indexed:0 read:2min 2026-09-18 ·