{"slug": "inadvertent-context-leakage-in-language-models", "title": "Inadvertent Context Leakage in Language Models", "summary": "A new study from arXiv researchers finds that language models leak sensitive in-context secrets through benign outputs, with 2-digit secrets reconstructed with near-perfect accuracy and 4-digit secrets at 82% exact match across eight proprietary models. The researchers developed an adaptive black-box attack and demonstrated practical attacks, including extracting full Social Security Numbers from a production-style agent, and found that more capable models leak more, suggesting leakage is a byproduct of capability.", "body_md": "# Computer Science > Machine Learning\n\n[Submitted on 20 Aug 2026]\n\n# Title:Inadvertent Context Leakage in Language Models\n\n[View PDF](/pdf/2608.19857)\n\n[HTML (experimental)](https://arxiv.org/html/2608.19857v1)\n\nAbstract:For AI agents to be useful beyond simple chat, they must hold sensitive user context such as calendars, credentials, health records, and financial data. We study whether the mere presence of such secrets in a model's context window introduces hidden correlations into the model's benign outputs, allowing reconstruction even when the model correctly refuses direct extraction. We further study whether an adversary can actively engineer prompts that amplify this effect, using the model as a covert carrier to transmit secrets through seemingly innocuous text. In both cases, this limited leakage is exploited using a novel adaptive attack that assumes black-box access to the underlying model.\n\nIn controlled experiments across eight proprietary models, we find that 2-digit in-context secrets are reconstructed with near-perfect accuracy and 4-digit secrets at 82\\% exact match, all from outputs the model produces in response to ordinary, non-adversarial requests. We observe that more capable models leak more: stronger instruction-following amplifies sensitivity to in-context secrets, suggesting leakage is a byproduct of capability as opposed to a patchable bug. We show this leakage enables two practical attacks: (1) a trained classifier that infers semantic predicates about user memories (e.g., health conditions, financial events) from routine natural-language outputs, and (2) an RL-trained adversary that extracts full Social Security Numbers from a production-style agent.\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer\n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers\n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps\n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations\n\n*(*[What are Smart Citations?](https://www.scite.ai/))# Code, Data and Media Associated with this Article\n\nalphaXiv\n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers\n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub\n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub\n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face\n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast\n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower\n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender\n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\nIArxiv Recommender\n\n*(*[What is IArxiv?](https://iarxiv.org/about))# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [ Learn more about arXivLabs](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/inadvertent-context-leakage-in-language-models", "canonical_source": "https://arxiv.org/abs/2608.19857", "published_at": "2026-08-21 08:07:06+00:00", "updated_at": "2026-08-21 08:43:57.928699+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-safety", "ai-research"], "entities": ["arXiv"], "alternates": {"html": "https://wpnews.pro/news/inadvertent-context-leakage-in-language-models", "markdown": "https://wpnews.pro/news/inadvertent-context-leakage-in-language-models.md", "text": "https://wpnews.pro/news/inadvertent-context-leakage-in-language-models.txt", "jsonld": "https://wpnews.pro/news/inadvertent-context-leakage-in-language-models.jsonld"}}