cd /news/ai-safety/artificial-analysis-launches-cyber-i… · home › topics › ai-safety › article
[ARTICLE · art-140973] src=cryptobriefing.com ↗ pub= topic=ai-safety verified=true sentiment=· neutral

Artificial Analysis launches Cyber Index to benchmark AI models on enterprise vulnerability defense

Artificial Analysis launched its Cyber Index (v1) on September 25, a composite benchmark evaluating how well AI models discover, validate, and patch vulnerabilities in enterprise systems. The index combines three partner-contributed benchmarks — CWE-Bench-AA, DeepsecBench-AA, and CyberGym-E2E-AA — and is deliberately limited to defensive cybersecurity rather than offensive intrusion capability. Artificial Analysis, co-founded by Micah Hill-Smith, previously released domain Capability Indices for finance, legal, and healthcare, and its general-purpose Intelligence Index is at version 4.3.2, aggregating 10 evaluations across hundreds of AI models.

by read2 min views1 publishedSep 28, 2026
Artificial Analysis launches Cyber Index to benchmark AI models on enterprise vulnerability defense
Image: Cryptobriefing (auto-discovered)

The new index evaluates how well AI systems can discover, validate, and patch security vulnerabilities using three specialized benchmarks developed with industry partners.

Artificial Analysis, the independent AI benchmarking platform, is expanding into cybersecurity. The organization launched its Cyber Index (v1) on September 25, evaluating how effectively AI models perform the work of finding and fixing vulnerabilities in enterprise systems.

Three benchmarks, one composite score #

The Cyber Index is built on a composite of three distinct benchmarks, each contributed by industry partners: CWE-Bench-AA, DeepsecBench-AA, and CyberGym-E2E-AA. Together, they cover the full lifecycle of enterprise vulnerability management, from discovery through validation to patching.

CWE-Bench-AA appears oriented around Common Weakness Enumeration categories, the standardized classification system that catalogs software and hardware vulnerability types. DeepsecBench-AA focuses on deeper security analysis tasks. CyberGym-E2E-AA rounds out the trio with end-to-end evaluation scenarios, testing whether AI agents can navigate realistic attack surfaces from start to finish.

Defense, not offense #

One deliberate design choice stands out: the Cyber Index focuses squarely on defensive cybersecurity. The index evaluates models on their ability to help businesses preemptively manage vulnerabilities, not on their capacity to break into systems.

AI, tech, and the markets they move—in one daily briefing.

Daily. Free. Join 34,000+ readers across crypto, finance, and policy.

Building on a benchmarking empire #

The Cyber Index is not Artificial Analysis’s first foray into specialized domain evaluations. The platform, co-founded by Micah Hill-Smith, has previously launched Capability Indices covering finance, legal, healthcare, and other sectors. Its general-purpose Intelligence Index has reached version 4.3.2, aggregating 10 separate evaluations across hundreds of AI models used by major labs.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our

Editorial Policy.

── more in #ai-safety 4 stories · sorted by recency
── more on @artificial analysis 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/artificial-analysis-…] indexed:0 read:2min 2026-09-28 · —