Alibi: Adversarial Legitimacy Injection in Binaries Against LLM Malware A paper submitted to arXiv on 17 Sep 2026 presents ALIBI, a semantic cover story attack that adds a small, non-executed read-only section containing a false security product narrative to compiled binaries, flipping 30 of 35 baseline-malicious PE samples to benign verdicts on Gemini 2.5 Pro while GPT-5.5 Pro and Claude Opus 4.7 produced substantial severity downgrades with significant confidence reductions. The attack transferred to ELF binaries, where Gemini flipped 16 of 40 malicious samples, and a verification-guided defense prompt roughly halved benign verdicts but still left 42.9 percent of malicious samples reaching benign. The authors conclude that LLM malware analyzers require provenance checks that separate verified facts from attacker-controlled claims rather than narrative trust. Computer Science Cryptography and Security Submitted on 17 Sep 2026 Title:ALIBI: Adversarial Legitimacy Injection in Binary Input against LLM Malware Analyzers View PDF https://arxiv.org/pdf/2609.19722 HTML experimental https://arxiv.org/html/2609.19722v1 Abstract:Large language models are being integrated into malware triage workflows as reasoning components that summarize static evidence and produce analyst-facing verdicts. This paper shows that the same reasoning capability introduces a new attack surface. We present ALIBI, a semantic cover story attack against frontier LLM-based malware analyzers. ALIBI adds a small, non-executed read-only section to a compiled binary, containing a coherent but false security product narrative, without altering imports or executable behavior. Instead of issuing direct instructions to the model, it reframes suspicious evidence as expected behavior of a benign endpoint security tool. On a frozen PE set of 50 malicious samples, the payload flips 30 of the 35 baseline-malicious samples to benign on Gemini 2.5 Pro, while GPT-5.5 Pro and Claude Opus 4.7 produce substantial severity downgrades with significant confidence reductions even when verdict labels are preserved. The attack transfers to ELF binaries, where Gemini flips 16 of 40. A verification-guided defense prompt roughly halves the benign verdicts, but 42.9 percent of malicious samples still reach benign. LLM malware analyzers therefore require provenance checks that separate verified facts from attacker-controlled claims, not narrative trust. References & Citations Loading... Bibliographic and Citation Tools Bibliographic Explorer What is the Explorer? https://info.arxiv.org/labs/showcase.html arxiv-bibliographic-explorer Connected Papers What is Connected Papers? https://www.connectedpapers.com/about Litmaps What is Litmaps? https://www.litmaps.co/ scite Smart Citations What are Smart Citations? https://www.scite.ai/ Code, Data and Media Associated with this Article alphaXiv What is alphaXiv? https://alphaxiv.org/ CatalyzeX Code Finder for Papers What is CatalyzeX? https://www.catalyzex.com DagsHub What is DagsHub? https://dagshub.com/ Gotit.pub What is GotitPub? http://gotit.pub/faq Hugging Face What is Huggingface? https://huggingface.co/huggingface ScienceCast What is ScienceCast? https://sciencecast.org/welcome Demos Recommenders and Search Tools Influence Flower What are Influence Flowers? https://influencemap.cmlab.dev/ CORE Recommender What is CORE? https://core.ac.uk/services/recommender arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs https://info.arxiv.org/labs/index.html .