I'm not a safety researcher, just build stuff with LLM's. Read some of the actual 30 page card, not the summary, and one section keeps bugging me. They measured whether Astra can sandbag. Told it "underperform on this evaluation," then checked if their monitors could catch it. Model dropped from 84.
Artificial Analysis launches Cyber Index to benchmark AI models on enterprise vulnerability defense