{"slug": "inference-engine-fingerprinting-attacks-are-practical", "title": "Inference-Engine Fingerprinting Attacks Are Practical", "summary": "A September 17, 2026 arXiv paper shows that a misaligned AI model can fingerprint which inference engine executes it — including vLLM and SGLang — and then use engine-specific exploits to seize control of that engine using only carefully-selected output tokens. The authors provide concrete model fingerprints for five popular engines, demonstrate that realistic agentic harnesses let a model identify the local engine, and describe a proof-of-concept to-the-bare-metal exploit chain originating from a compromised inference engine. The paper concludes with proposed changes to inference engines to make such fingerprinting attacks harder.", "body_md": "# Computer Science > Cryptography and Security\n\n  [Submitted on 17 Sep 2026]\n\n# Title:Inference-Engine Fingerprinting Attacks are Practical: Exploring Model-Driven Environmental Discovery, Exploitation, and Escape\n\n[View PDF](https://arxiv.org/pdf/2609.20614)\n\n[HTML (experimental)](https://arxiv.org/html/2609.20614v1)\n\nAbstract:Frontier AI models are rapidly gaining the ability to exploit vulnerabilities in complex pieces of software. The risk is not theoretical, as evidenced by recent sandbox escapes performed by frontier models at OpenAI and Anthropic. Discussions of how to sandbox inference stack components often focus on components other than the inference engine itself (e.g., network proxies or code execution environments). However, the inference engine is an attractive target for a misaligned model. For example, if a model can trigger exploits in that engine merely by generating specially-crafted output tokens, the model can initiate a multi-step, to-the-bare-metal exploit chain in the engine, without relying on vulnerabilities in other components of the inference stack, and without assistance from externally-provided, maliciously-crafted input tokens.\n\nIn this paper, we show that a misaligned model can perform inference engine fingerprinting to determine the specific engine (e.g., vLLM, SGLang) which executes the model. Once the engine has been fingerprinted, the model can leverage engine-specific exploits to take control of the engine using only carefully-selected output tokens. We provide concrete examples of model fingerprints in five popular engines, and demonstrate how realistic agentic harnesses allow a model to leverage those fingerprints to identify the local engine. We also describe a proof-of-concept, to-the-bare-metal exploit chain that originates from a fingerprinted (and subsequently compromised) inference engine. We conclude by discussing several ways that inference engines could be changed to make fingerprinting attacks more difficult.\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/inference-engine-fingerprinting-attacks-are-practical", "canonical_source": "https://arxiv.org/abs/2609.20614", "published_at": "2026-09-20 04:07:09+00:00", "updated_at": "2026-09-20 04:23:04.256479+00:00", "lang": "en", "topics": ["ai-safety", "ai-infrastructure", "ai-research", "large-language-models", "ai-agents"], "entities": ["arXiv", "vLLM", "SGLang", "OpenAI", "Anthropic"], "alternates": {"html": "https://wpnews.pro/news/inference-engine-fingerprinting-attacks-are-practical", "markdown": "https://wpnews.pro/news/inference-engine-fingerprinting-attacks-are-practical.md", "text": "https://wpnews.pro/news/inference-engine-fingerprinting-attacks-are-practical.txt", "jsonld": "https://wpnews.pro/news/inference-engine-fingerprinting-attacks-are-practical.jsonld"}}