{"slug": "control-token-injection-suppresses-chain-of-thought-and-defeats-reasoning-based", "title": "Control-Token Injection Suppresses Chain-of-Thought and Defeats Reasoning-Based", "summary": "A September 23, 2026 arXiv paper by Usama Muhammad reports that appending a single string of a model's own channel-control tokens to a user message suppresses chain-of-thought in the released gpt-oss-20b reasoning model under its published tool sandbox, dropping the reasoning channel from a mean of 52.5 tokens to zero across forty tasks while the tool call still fired on every trial. The paper also found that on overtly malicious requests the attack converted 39.6% of the model's refusals into completed exfiltrations, and that identical tool-call generations produced opposite outcomes depending on the harness parser, with two parsers shipped for the Gemma agent firing on all twenty-four trials and on none. The author evaluates input sanitization, parser hardening, and empty-reasoning detection as defenses, noting that flagging an absent trace catches the basic attack but not an adaptive benign decoy.", "body_md": "# Computer Science > Cryptography and Security\n\n  [Submitted on 23 Sep 2026]\n\n# Title:Control-Token Injection Suppresses Chain-of-Thought and Defeats Reasoning-Based Oversight in Tool-Using Agents\n\n[View PDF](https://arxiv.org/pdf/2609.27542)\n\n[HTML (experimental)](https://arxiv.org/html/2609.27542v1)\n\nAbstract:The safety of a tool-using language model agent is usually treated as a property of the model alone. We give controlled, full-precision evidence that it is instead a joint property of the model and the software that renders its chat template and parses its tool calls, the decoding harness, and that both halves are attackable from untrusted input. On the released gpt-oss-20b reasoning model under its published tool sandbox, appending a single string of the model's own channel-control tokens to a user message makes the tokenizer render a reasoning turn that is already complete, so the model writes no chain-of-thought and proceeds directly to the tool call. Across forty tasks the model already completes, the reasoning channel falls from a mean of 52.5 tokens to zero on every trial while the [this http URL](http://http.post) still fires on every trial. A rule monitor and a cross-family language-model monitor detect the unsafe request on all plain trials and no forged trials, and on overtly malicious requests the attack converts 39.6% of the model's refusals into completed exfiltrations. Separately, whether an identical tool-call generation fires is decided by the harness parser, not the model: a truncation-tolerant regular expression fires a call whose closing token is missing while a strict one drops it, and two parsers shipped for the Gemma agent give opposite outcomes on identical greedy generations, firing on all twenty-four trials and on none. We show the suppression can be delivered indirectly and characterize its dependence on the chat template across two more reasoning models, and we evaluate input sanitization, parser hardening, and empty-reasoning detection as defenses; flagging an absent trace catches the basic attack but not an adaptive benign decoy. All measurements use greedy decoding on publicly released models. Code and per-trial logs: [this https URL](https://github.com/Usama1002/deleting-the-trace)\n\n## Submission history\n\nFrom: Usama Muhammad Mr. [\n[view email](https://arxiv.org/show-email/e4743486/2609.27542)]\n\n**[v1]** Wed, 23 Sep 2026 08:33:47 UTC (1,785 KB)\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/control-token-injection-suppresses-chain-of-thought-and-defeats-reasoning-based", "canonical_source": "https://arxiv.org/abs/2609.27542", "published_at": "2026-09-25 06:07:08+00:00", "updated_at": "2026-09-25 06:29:49.866307+00:00", "lang": "en", "topics": ["ai-safety", "ai-agents", "large-language-models", "ai-research"], "entities": ["Usama Muhammad", "gpt-oss-20b", "Gemma", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/control-token-injection-suppresses-chain-of-thought-and-defeats-reasoning-based", "markdown": "https://wpnews.pro/news/control-token-injection-suppresses-chain-of-thought-and-defeats-reasoning-based.md", "text": "https://wpnews.pro/news/control-token-injection-suppresses-chain-of-thought-and-defeats-reasoning-based.txt", "jsonld": "https://wpnews.pro/news/control-token-injection-suppresses-chain-of-thought-and-defeats-reasoning-based.jsonld"}}