# Felony Bench

> Source: <https://www.felonybench.com/>
> Published: 2026-08-21 15:17:04+00:00

# Felony Bench

A benchmark you really don't want models to be saturated with.

[Learn more](#felony-records)

Score

↖ Most illegalLeast illegal ↘

Anthropic

OpenAI

Meta

Google

Moonshot

Scores indicate count of illegal activity. Higher is... you decide.

| Company | Felonies | Description | Date | Source |
|---|---|---|---|---|
| Anthropic | 1 | Exploited auth failures in an API to cancel other people's gym classes |
|

[The Information](https://www.theinformation.com/articles/meta-ai-model-hacked-another-company-cybersecurity-testing)[AISI](https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing)[OpenAI](https://openai.com/index/third-party-cyber-evaluations-involving-openai-models/)[AISI](https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing)[OpenAI](https://openai.com/index/third-party-cyber-evaluations-involving-openai-models/)[OpenAI](https://openai.com/index/hugging-face-model-evaluation-security-incident/)[Reuters](https://www.reuters.com/business/openai-finds-evidence-other-ai-agents-escaped-containment-it-widens-hacking-2026-07-31/)[Anthropic](https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals)[OpenAI](https://openai.com/index/hugging-face-model-evaluation-security-incident/)## Methodology

Felony Bench counts unique instances where AI agents affect third-party entities. Escaping a sandbox alone does not constitute a counted incident. It is for these reasons that Frontier Security's [Kimi K3](https://blog.frontier.security/chinese-model-kimi-k3-breaks-uk-ai-safety-institute-benchmark-evaluations/) incident and [Alibaba's ROME](https://arxiv.org/pdf/2512.24873) incident are not counted.
