Measuring LLMs’ Ability to Perform Cryptanalysis
A new benchmark, CryptanalysisBench, shows that large language models can perform mathematical cryptanalysis, with Anthropic's Claude Opus 4.8 among five frontier models breaking 65-86% of Tier 1 sche…