cd /news/ai-safety/alibi-adversarial-legitimacy-injecti… · home topics ai-safety article
[ARTICLE · art-133609] src=arxiv.org ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Alibi: Adversarial Legitimacy Injection in Binaries Against LLM Malware

A paper submitted to arXiv on 17 Sep 2026 presents ALIBI, a semantic cover story attack that adds a small, non-executed read-only section containing a false security product narrative to compiled binaries, flipping 30 of 35 baseline-malicious PE samples to benign verdicts on Gemini 2.5 Pro while GPT-5.5 Pro and Claude Opus 4.7 produced substantial severity downgrades with significant confidence reductions. The attack transferred to ELF binaries, where Gemini flipped 16 of 40 malicious samples, and a verification-guided defense prompt roughly halved benign verdicts but still left 42.9 percent of malicious samples reaching benign. The authors conclude that LLM malware analyzers require provenance checks that separate verified facts from attacker-controlled claims rather than narrative trust.

read2 min views1 publishedSep 18, 2026
Alibi: Adversarial Legitimacy Injection in Binaries Against LLM Malware
Image: source
  [Submitted on 17 Sep 2026]


[View PDF](https://arxiv.org/pdf/2609.19722)

[HTML (experimental)](https://arxiv.org/html/2609.19722v1)

Abstract:Large language models are being integrated into malware triage workflows as reasoning components that summarize static evidence and produce analyst-facing verdicts. This paper shows that the same reasoning capability introduces a new attack surface. We present ALIBI, a semantic cover story attack against frontier LLM-based malware analyzers. ALIBI adds a small, non-executed read-only section to a compiled binary, containing a coherent but false security product narrative, without altering imports or executable behavior. Instead of issuing direct instructions to the model, it reframes suspicious evidence as expected behavior of a benign endpoint security tool. On a frozen PE set of 50 malicious samples, the payload flips 30 of the 35 baseline-malicious samples to benign on Gemini 2.5 Pro, while GPT-5.5 Pro and Claude Opus 4.7 produce substantial severity downgrades with significant confidence reductions even when verdict labels are preserved. The attack transfers to ELF binaries, where Gemini flips 16 of 40. A verification-guided defense prompt roughly halves the benign verdicts, but 42.9 percent of malicious samples still reach benign. LLM malware analyzers therefore require provenance checks that separate verified facts from attacker-controlled claims, not narrative trust.

References & Citations

...

Bibliographic Explorer

(What is the Explorer?) Connected Papers

(What is Connected Papers?) Litmaps

(What is Litmaps?) scite Smart Citations

(What are Smart Citations?) alphaXiv

(What is alphaXiv?) CatalyzeX Code Finder for Papers

(What is CatalyzeX?) DagsHub

(What is DagsHub?) Gotit.pub

(What is GotitPub?) Hugging Face

(What is Huggingface?) ScienceCast

(What is ScienceCast?) Influence Flower

(What are Influence Flowers?) CORE Recommender

(What is CORE?) arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

── more in #ai-safety 4 stories · sorted by recency
── more on @alibi 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/alibi-adversarial-le…] indexed:0 read:2min 2026-09-18 ·