LLM Security Leaderboard Now Ranks Every Modality Where Agent Risk Lives: Text, Image, and Audio Cisco's AI Defense LLM Security Leaderboard added 102 new evaluations across text, image, and audio modalities since June 2026, bringing its total to 136 models from Anthropic, OpenAI, Google, xAI, Meta, Mistral, and Amazon. The update includes 69 new entries — 55 image models and 14 audio models — and tests prompt injection, jailbreaks, and other manipulation techniques in base configurations without additional guardrails. Cisco said the expansion matters because agents that browse the web, read screenshots, or take voice input face attack surfaces that text-only evaluations would miss. With research and development support from Ravikumar Balakrishnan, Ankit Garg, and Sanket Mendapara When we launched the Cisco LLM Security Leaderboard https://leaderboard.aidefense.cisco.com/ earlier this year, the goal was simple: give organizations clear, tested data on how models hold up against attacks, so they know the risks before they deploy one. That matters because AI models are increasingly built into products such as agents that read email, browse the web, and take actions on a person’s behalf. A model that can be manipulated could be turned against the person using it. That risk also varies by deployment: a model wired into a browsing agent is exposed on different inputs or modalities such as text, images, and audio than one only answering questions in a chat window, so where a specific model is weak matters as much as where it’s strong. The leaderboard tests for that a few different ways: prompt injection, where a malicious instruction is hidden in content the model processes, like a webpage or image; jailbreaks, where a model is talked into ignoring its own safety rules; and other techniques that push a model toward harmful or unsafe output. The exact method varies a single message or a drawn-out conversation, direct or obfuscated, text or image or audio , and so does the type of harm being tested for, but the underlying question is always the same: can this model be manipulated? Some models resist far better than others. Connect a weak one to an agent, and the risk grows. 102 new evaluations across modalities since June 2026 The LLM Security Leaderboard is one of the most comprehensive model security leaderboards. Since June, we added 102 new entries across three modalities to a total of 136 models, spanning frontier and open-weight releases from Anthropic, OpenAI, Google, xAI, Meta, Mistral, and others. As always, we test models in their base configuration without additional guardrails, so scores reflect a consistent baseline for layering on additional security protections. Multimodal results are now live Until now, the leaderboard measured text-based attacks two ways: single-turn, where one harmful message is sent straight to the model, and multi-turn, a longer back-and-forth where the attacker slowly builds up to a harmful request over several messages. That covered the most common way people interact with models, but today, models also power agents that can act. A model that can call tools, browse the web, or operate a computer is an agent, and an agent takes in information from everywhere it operates: a page it reads, a file it opens, an image it’s shown, a result a tool hands back. Each of those is a place an attacker can plant an instruction, and text is only one of the forms that an instruction can arrive in. A website an agent visits can embed a prompt injection in an image such an advertisement; a voice assistant can be handed an audio clip that may be engineered to manipulate it. If a model only gets evaluated on text, that risk may not show up until it becomes a real incident. Consider which of those surfacesactually matters in the context of what you’re deploying: an agent that only reads and writes text would not need to worry about its image resistance, but one that can browses the web, reads screenshots, or takes voice input does. Those are circumstances where text-only evaluations would not tell the whole security story. Today we’re releasing an update to the leaderboard that now expands beyond just text models. We have added 69 new entries including 55 image models and 14 audio models across Amazon, Anthropic, Google, Meta, Mistral, OpenAI and xAI. Each of those labs takes a different approach to building and training multimodal capability, whether that’s how image data flows into the LLM backbone, how much safety alignment goes into a vision or audio stack versus the base language model, or which modalities are red-teamed and evaluated internally. Those differences show up directly in how a model resists attack on one modality versus another. Image and audio attacks are tested the same way as single-turn text attacks one attempt, one message , using the same attack and harm categories as its text score, so their resistance is comparable across surfaces. Each model’s overall Combined Score is now an average across every format it was evaluated on, and a new modality switch lets you isolate scores for text, image, or audio on their own. That makes it possible to check a model against the specific modalities an AI deployment actually exposes it to, and to decide where that model would need to layer on additional defenses, like input filtering or output guardrails, for the modality where that model is weakest. Figure 1. Screenshot of image capable model rankings on the Cisco LLM Security Leaderboard https://blogs.cisco.com/gcs/ciscoblogs/1/2026/09/Figure-11.png Image model leaderboard results In our tests, Google’s Gemini 3.1 Pro Preview ranks the best-performing image model, resisting 93.9% of adversarial image attacks, just ahead of Anthropic’s Claude Opus 4.5 93.7% , both scoring in the leaderboard’s “Excellent” range 85–100% . Mistral’s Magistral Small 2509 performed poorly, refusing only 23.0% of attacks, meaning it complied with more than three out of every four image-based attacks it was tested against. The difference in testing images is that image-based attacks are single-turn only, with a single image carrying a hidden instruction, not a back-and-forth conversation. The attack methods are different in kind too, not just format: text hidden inside an image using typographic tricks, instructions embedded in a diagram or figure, or an attack that splits its intent between the image and an accompanying text prompt so neither half looks harmful on its own. The leaderboard displays evaluation results from models that can actually see images, which account for 55 of the 136 models on the leaderboard. Figure 2. Screenshot of audio capable models rankings on the Cisco LLM Security Leaderboard https://blogs.cisco.com/gcs/ciscoblogs/1/2026/09/Screenshot-2026-09-23-at-16.17.14.png Audio model leaderboard results In our latest test, Google’s Gemini 3.1 Pro Preview ranks as the best-performing audio model tested, refusing 90.0% of adversarial audio attacks, while Mistral’s Voxtral Small 24b 2507 demonstrated only 9.0% refusal rate, meaning it complied with roughly 9 out of every 10 audio attacks it faced. Like image, audio models were also single-turn only, using one adversarial audio clip rather than a conversation. This is also the newest and smallest slice of the leaderboard. Just 9 models across Google, Mistral, and OpenAI currently accept audio input and have been tested, so this ranking should be read as early results rather than a mature field. How to interpret new combined results view Text scores remain unchanged for every model that was already on the leaderboard, but what changed is how the Combined Score averages text, image, and audio modalities that a model has been tested on. The Combined Score may shift as the result of an image or audio result, even though its text score hadn’t changed. The direction of that shift depends entirely on how a model’s image or audio resistance compares to its text resistance. Some strong text performers dropped once image was factored in: Claude Sonnet 4.5 fell 7.2 points from 92.2 to 85.0 and dropped from 2 overall to 20; Claude Haiku 4.5 fell 8.1 points and dropped from 4 to 24; Amazon Nova 2 Lite fell 11.2 points and dropped from 25 to 50, each because its image resistance is meaningfully weaker than its text resistance. The two Mistral Voxtral models fell for the same reason based on their audio score. Other models climbed when image evaluations were added to the cross-modal score. Google’s four image-tested Gemini models showed the largest image-over-text advantages, while all three image-tested Gemma 3 variants and OpenAI’s GPT‑4.1 nano, GPT‑4.1 mini, and GPT‑4o mini also demonstrated stronger image than text resistance. That spread is a reminder that security work on one modality does not automatically transfer to another, especially across labs that built and trained their image or audio capabilities independently from their text models in the first place. A model’s Combined Score can move sharply once it’s tested against different modalities, especially when its security posture is uneven across modalities. That movement reflects how the score is calculated, not a change in how well the model actually defends itself. Check a model’s individual Text, Image, and Audio columns before taking its Combined Score as the whole story. Integration with AI Supply Chain Provenance Explorer Provenance matters because a model’s weaknesses often aren’t unique to that model. If two models share lineage, a vulnerability discovered in one can be present in the other, and stopping an investigation at the model currently deployed can miss where a problem actually originated or where else it might surface. That makes provenance most useful exactly when you’re actively investigating a model’s security and need to know what it’s related to. As such, we’ve also connected the leaderboard to the AI Supply Chain Provenance Explorer https://blogs.cisco.com/ai/supply-chain-provenance-explorer . Open-weight models on the rankings page now link directly to their provenance profile, showing lineage and fingerprint data drawn from the same techniques behind Model Provenance Kit. Security posture and where a model actually came from are related questions, so we made it easy for you to view them in one place. To see the full rankings, filter by modality, or look up a specific model, visit the Cisco LLM Security Leaderboard https://leaderboard.aidefense.cisco.com/ today.