{"slug": "stealing-reasoning-traces-from-proprietary-llm-apis", "title": "Stealing Reasoning Traces from Proprietary LLM APIs", "summary": "Researchers led by Alexander Panfilov have disclosed a vulnerability in proprietary LLM APIs from Anthropic, OpenAI, and Google that allows attackers to decrypt hidden chain-of-thought reasoning traces by injecting them into weaker models from the same provider, enabling four attack vectors including circumvention of anti-distillation, extraction of 367 PII artifacts and 182 credentials from 315,320 decoded blocks, exposure of hazardous information, and invisible prompt injections. The paper proposes cryptographic and system-level mitigations following responsible disclosure.", "body_md": "# Computer Science > Cryptography and Security\n\n[Submitted on 10 Aug 2026]\n\n# Title:Stealing Reasoning Traces from Proprietary LLM APIs\n\n[View PDF](/pdf/2608.09867)\n\nAbstract:Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather than storing these traces server-side, providers return them to the client as blocks of encrypted text, which the client passes back with each subsequent request. Building on prior research, we identify an architectural vulnerability: these encrypted blocks are fully compatible and interchangeable across different sessions, users, and models within a provider's ecosystem. We exploit this compatibility to develop a scalable decryption jailbreak. By injecting an encrypted reasoning trace from a given model into a weaker, and less safeguarded model from the same provider, we force it to decode and output the trace verbatim in plaintext, without ever jailbreaking the more capable model directly. This vulnerability enables four distinct attack vectors. First, it circumvents anti-distillation mechanisms, allowing adversaries to extract a proprietary model's reasoning, as we demonstrate across Anthropic, OpenAI, and Google. Second, it allows for large-scale private data extraction. Developers frequently share session logs publicly, unaware of contents of the encrypted blocks. By decoding 315,320 reasoning blocks scraped from public repositories, we recovered 367 Personally Identifiable Information (PII) artifacts and 182 credentials. Third, it inadvertently reveals hazardous information hidden within the reasoning process, even in cases where the model's final, visible output safely rejects a malicious request. Fourth, attackers can leverage this flaw to execute invisible prompt injections, embedding malicious payloads entirely within encrypted blocks to poison public agentic rollouts. Following responsible disclosure, we propose concrete cryptographic and system-level mitigations to secure client-side reasoning.\n\n## Submission history\n\nFrom: Alexander Panfilov [[view email](/show-email/81199360/2608.09867)]\n\n**[v1]** Mon, 10 Aug 2026 17:24:50 UTC (16,346 KB)\n\n### Current browse context:\n\ncs.CR\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer\n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers\n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps\n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations\n\n*(*[What are Smart Citations?](https://www.scite.ai/))# Code, Data and Media Associated with this Article\n\nalphaXiv\n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers\n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub\n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub\n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face\n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast\n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower\n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender\n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [ Learn more about arXivLabs](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/stealing-reasoning-traces-from-proprietary-llm-apis", "canonical_source": "https://arxiv.org/abs/2608.09867", "published_at": "2026-08-11 07:07:06+00:00", "updated_at": "2026-08-11 07:41:20.395003+00:00", "lang": "en", "topics": ["ai-safety", "ai-policy", "artificial-intelligence", "large-language-models"], "entities": ["Alexander Panfilov", "Anthropic", "OpenAI", "Google"], "alternates": {"html": "https://wpnews.pro/news/stealing-reasoning-traces-from-proprietary-llm-apis", "markdown": "https://wpnews.pro/news/stealing-reasoning-traces-from-proprietary-llm-apis.md", "text": "https://wpnews.pro/news/stealing-reasoning-traces-from-proprietary-llm-apis.txt", "jsonld": "https://wpnews.pro/news/stealing-reasoning-traces-from-proprietary-llm-apis.jsonld"}}