CryptanalysisBench: Can LLMs Do Cryptanalysis?
A new benchmark, CryptanalysisBench, shows that five frontier large language models can break 65%-86% of Tier 1 cryptographic schemes and produce novel cryptanalysis, including a key-recovery attack o…
A new benchmark, CryptanalysisBench, shows that five frontier large language models can break 65%-86% of Tier 1 cryptographic schemes and produce novel cryptanalysis, including a key-recovery attack o…
Claude Opus 5 passed 24% of 17 checkpoints on a subset of SlopCodeBench, quadrupling the 6% pass rates of Opus 4.8 and Sonnet 5, but achieved this by producing roughly 29,000 source lines versus about…
Opus 5 achieved a 24% strict pass rate on a subset of the SlopCodeBench coding benchmark from UW Madison, only marginally higher than Opus 4.6's 17% in the original paper, and wrote five times more fu…
Claude Opus 5, Anthropic's new model, leads on agentic knowledge work and undercuts Fable 5 on cost, according to Artificial Analysis. The model scores 61 on the Artificial Analysis Intelligence Index…
Anthropic reported elevated error rates for Sonnet 4.6 and Sonnet 5, with notifications available via email or SMS for incident updates.…
Anthropic released Claude Opus 5 on July 24, offering near-Fable-5 performance at the same pricing as Opus 4.8 ($5 per million input tokens, $25 per million output tokens), deliberately cannibalizing …
Amazon Web Services launched Anthropic's Claude Opus 5 on Amazon Bedrock on 24 July 2026, pricing the model at $5 per million input tokens and $25 per million output tokens — half the cost of the flag…
Anthropic released Claude Opus 5, a model that costs $5 per million input tokens and $25 per million output tokens—half the price of its higher-tier Fable 5—while using one-seventh the reasoning token…
Anthropic on Thursday released Claude Opus 5, a model that matches or exceeds its flagship Fable 5 on most coding and knowledge-work benchmarks at half the token price, costing $5 per million input to…
Anthropic released Claude Opus 5 on Friday, calling it its safest and most aligned model to date, with the lowest rates of deceptiveness and susceptibility to misuse among prior models. The model is p…
Anthropic's Claude Opus 4.8, Sonnet 5, and Fable 5 maintain a flat rate across their 1M token window, while Google's Gemini 3.1 Pro caps at 1M tokens with a pricing jump after 200K tokens, making Clau…
Claude Opus 5, Anthropic's latest model, shows brilliant reasoning but suffers from a neurotic personality and verbosity that frustrated the reviewer during real coding sessions, including refusing to…
Anthropic released Claude Opus 5, a model that approaches the performance of its own frontier system Fable 5 while costing half as much to run, priced at $5 per million input tokens and $25 per millio…
Anthropic launched Opus 5 on Friday, a new version of its heavyweight model that outperforms Fable 5 on several benchmarks while being cheaper and less restrictive. Opus 5 is free from the 30-day data…
Anthropic is preparing to ship Claude Opus 5, with evidence including a brief appearance in a coding tool's model picker and a matching model string on Google's Vertex quotas catalog, signaling a pote…
Claude reported elevated error rates for Sonnet 5, with users experiencing increased failures. The company is investigating the issue and providing updates via email and SMS notifications.…
A developer is canceling their Claude Code subscription after testing DeepSeek V4 Flash against Anthropic's Sonnet 5 and Opus 4.8 on a real-world coding task, finding that the free DeepSeek model outp…
Anthropic's Fable 5 model achieves 96% of the performance of the full Fable 5 at 46% of the cost when used as an orchestrator that delegates implementation to cheaper models, according to a technical …
Orchflows, a new open-source framework by Dan McInerney, lets developers build self-improving agent loops in a single sentence by autorouting every request to the smallest subagent-driven workflow. Th…
Moonshot AI released Kimi K3, a 2.8 trillion-parameter open-weight model that it claims is the largest of its kind globally, with performance approaching closed-source leaders like Anthropic's Claude …