cd /news/artificial-intelligence/researchers-publish-cryptanalysisben… · home topics artificial-intelligence article
[ARTICLE · art-77414] src=runtimewire.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Researchers publish CryptanalysisBench to verify AI-generated cryptographic attacks

Researchers at ETH Zurich, Tel Aviv University, the University of Haifa and Anthropic published CryptanalysisBench, a benchmark for testing whether LLM agents can turn cryptographic reasoning into executable attacks. The July 20 arXiv preprint evaluates five frontier models on 191 tasks, reporting Tier 1 success rates from 65% to 86% and two previously unknown findings involving SpoC AEAD and KINDI. The benchmark requires agents to produce self-contained attack scripts verified against fresh game sessions, distinguishing practical breaks from theoretical explanations.

read5 min views1 publishedJul 28, 2026
Researchers publish CryptanalysisBench to verify AI-generated cryptographic attacks
Image: Runtimewire (auto-discovered)

Researchers at ETH Zurich, Tel Aviv University, the University of Haifa and Anthropic have published CryptanalysisBench, a benchmark for testing whether LLM agents can turn cryptographic reasoning into executable attacks.

The July 20 arXiv preprint lists eight authors: Lukas Fluri, Avital Shafran, Nicholas Carlini, Matthew Jagielski, Milad Nasr, Orr Dunkelman, Eyal Ronen and Florian Tramer. Anthropic described the project as a collaboration with academics at ETH Zurich, Tel Aviv University and the University of Haifa.

Anthropic on X The benchmark contains 191 tasks and evaluates five frontier models. The paper reports Tier 1 success rates ranging from 65% to 86%, along with two findings involving SpoC AEAD and KINDI that the authors say were previously unknown. The work remains an arXiv preprint, and the cited sources do not establish independent validation of those findings.

CryptanalysisBench is designed to test whether an agent can inspect a cryptographic scheme, identify a weakness, write an attack and demonstrate that the attack works. That distinction matters in cryptography, where a persuasive explanation does not establish a practical break.

From reasoning to a working attack

The paper says the 191 tasks span six primitive families: hash functions, block ciphers, authenticated encryption with associated data, key encapsulation mechanisms, public-key encryption and digital signatures. The authors say the tasks were drawn primarily from four US National Institute of Standards and Technology competitions.

The paper's abstract organizes the benchmark into three tiers: primitives with known practical breaks; primitives with no known practical break, tested at full strength and through scaled-down variants; and a challenge set of production primitives at the frontier of cryptanalysis. In the body, the authors refer to the first two groups as numbered tiers and describe the challenge set separately. Tier 1 contains 49 algorithms with known practical attacks, while Tier 2 contains 142 algorithms for which the researchers found no published practical attack or where existing attacks were too expensive to execute.

The benchmark requires models to deliver executable code. An agent receives the target algorithm's C implementation and documentation inside a standard Linux container. A separate challenger container holds keys, randomness and other secret material while exposing only the queries permitted by the relevant security game. The CryptanalysisBench paper says the final deliverable is a self-contained attack script that runs without manual intervention and is verified against fresh game sessions.

For probabilistic decision games, the verifier runs 20 independent trials and requires at least 17 wins. The paper calculates the probability of passing through random guessing at about 1.3 x 10^-3, or roughly 0.1%. The MIT-licensed CryptanalysisBench repository includes algorithm sources, task-generation scripts and a security-game framework.

The researchers evaluated five frontier models: Claude Opus 4.8, Claude Sonnet 5, Claude Mythos 5, GPT-5.5 and the open-weights GLM-5.2. Under the main two-hour agent budget, the models broke between 65% and 86% of Tier 1 schemes. Mythos 5 led with 42 wins out of 49, followed by Sonnet 5 and GPT-5.5 with 37 each, Opus 4.8 with 36 and GLM-5.2 with 32.

Those results require careful interpretation. Every Tier 1 scheme has a publicly documented practical break, so a model may have reproduced material encountered during training. The researchers inspected model traces to distinguish apparent recall from source-level analysis, while acknowledging that their method could not fully exclude memorization.

Some successful attacks exploited bugs in reference implementations rather than weaknesses in the underlying cryptographic design. The researchers retained those implementations because they were the versions originally submitted to and evaluated by the cryptography community. The CryptanalysisBench paper sorts Tier 2 wins into four cases: genuine design flaws, weaknesses admitted by underspecified designs, reference-implementation bugs that contradict unambiguous specifications and artifacts caused by reducing parameters.

Two claimed findings carry most of the weight

The strongest reported results came from Tier 2. The authors report that Mythos 5 and Sonnet 5 independently found a previously unreported full key-recovery attack against the unmodified SpoC AEAD. In their SpoC-128 case study, the authors attribute the weakness to improper handling of empty plaintext and associated data. The CryptanalysisBench paper says two chosen-nonce queries, one empty and one with a single plaintext block, leak the full internal state after one permutation. The authors further claim that the permutation's invertibility allows recovery of the original state, including the key, against the full 18-step reference scheme.

In its KINDI case study, the CryptanalysisBench paper reports that Mythos 5 built a decryption-reaction oracle from KINDI's published code. According to the authors, KINDI validates only the decrypted message field and does not re-encrypt the ciphertext, enabling a chosen-ciphertext decryption-reaction attack that recovers the secret key. The paper also argues that KINDI's NIST submission includes a lemma claiming chosen-ciphertext security without the Fujisaki-Okamoto re-encryption step, and that the lemma's uniqueness claim is incorrect because small ciphertext perturbations may produce the same decapsulated key. Independent replication of either the KINDI result or the SpoC result is not established in the cited sources.

Across the five models, the researchers counted 14 distinct full-strength Tier 2 schemes that at least one model broke. The paper classified SpoC and KINDI as design flaws, LIMA and DAGS as specification gaps, and the remaining 10 as reference-implementation bugs. The authors separately report that individual models broke six to 12 Tier 2 schemes at full strength and 24 to 61 scaled-down variants.

The models were sensitive to execution time. Extending Mythos 5's per-task budget from two hours to 16 hours increased its Tier 1 result from 42 to 47 wins out of 49. The number of distinct Tier 2 algorithms it broke rose from 61 to 75. The CryptanalysisBench paper reports that the same 16-hour rerun increased Mythos 5's count of solved, scaled-down Tier 2 version-level tasks from 119 to 164.

Of the 57 tasks that changed from failure to success under the longer budget, the authors estimated that about 49 shorter runs had already located the weakness and needed more time to complete or debug the attack. Coding, execution and debugging limits can therefore prevent a model from converting a suspected weakness into a working exploit.

CryptanalysisBench gives model developers, cryptographers and security teams a repeatable capability test backed by code-based verification. Researchers could also run agents against candidate schemes during standardization or before deployment, increasing the number of practical attacks attempted against a design. The initial results remain preprint claims, but the benchmark provides a concrete method for measuring when AI agents progress from discussing cryptographic weaknesses to executing attacks under defined security games.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @eth zurich 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/researchers-publish-…] indexed:0 read:5min 2026-07-28 ·