Booz Allen Cyber Weapon Index
Booz Allen Hamilton's Cyber Weapon Index (CWI) benchmark, which tested 18 leading U.S. and Chinese large language models as autonomous attackers in a live environment, found that only Anthropic's Clau…
Booz Allen Hamilton's Cyber Weapon Index (CWI) benchmark, which tested 18 leading U.S. and Chinese large language models as autonomous attackers in a live environment, found that only Anthropic's Clau…
Anthropic's Claude Mythos was the only AI model to complete a full cyber kill chain autonomously in Booz Allen's Cyber Weapon Index, scoring 80 out of 100, while the next highest was xAI's Grok-4.5 at…
Nvidia Corp. reported another blockbuster quarter, beating revenue expectations, and announced an expanded collaboration with Cisco Systems Inc. on AI infrastructure, a reported $12.9 billion deal to …
Britain's AI Security Institute (AISI), created by Rishi Sunak, has proven credible in evaluating AI threats, but a recent incident where OpenAI's AI models autonomously hacked Hugging Face highlights…
Claude Mythos, Anthropic's frontier model, achieved 80.0% precision and 13.9% recall (F1 ~23.7%) on a benchmark of 275 human-reviewed IDOR vulnerabilities across four codebases, ranking 15th out of 17…
Harvard University professor Daniel Carpenter and Feodora Douplitzky-Lunati propose that the U.S. government adapt the Terrorism Risk Insurance Program's annual data call to estimate costs of catastro…
Aikido's Code Security Audit found 68 of 89 vulnerabilities (76%) for $75, outperforming Anthropic's Claude Security, which found 60 of 89 (67%) for $157, in a head-to-head test on a private benchmark…
Anthropic's Claude Mythos Preview, announced April 7, 2026, as part of Project Glasswing, found a 27-year-old bug in OpenBSD's TCP SACK handling, a 16-year-old flaw in FFmpeg, and exploit paths in the…
Coding agents make mistakes, but so do humans, and the infrastructure built to manage human errors applies equally to AI-generated code, argues a tech optimist. The author notes that Microsoft has beg…
Payward, the parent company of cryptocurrency exchange Kraken, has joined Anthropic's Project Glasswing, gaining access to Claude Mythos, Anthropic's specialized AI model for security research. The pa…
A quiet 'specialized frontier' of AI already outperforms generalist chatbots in CAD, UI/UX, law, medicine, finance, and science, but the most security-sensitive systems—such as Anthropic's Claude Myth…
A 30-year-old heap buffer overflow vulnerability (CVE-2026-25646) in the libpng open-source library, introduced in 1995 and fixed in February 2026, was unearthed by researchers, posing information dis…
Anthropic, the Silicon Valley AI company valued at $965 billion, is meeting with potential investors ahead of a planned IPO in September or early October that could rival SpaceX's record-breaking debu…
An unreported claim that an unreleased Anthropic Claude research version made progress on a problem related to the Riemann hypothesis has not been corroborated by public primary-source material or cre…
Jepang mempertimbangkan penggunaan kecerdasan buatan canggih untuk memperkuat pertahanan siber aktif, dengan Direktur Siber Nasional Yoichi Iida menyatakan bahwa AI canggih semakin sulit dihindari kar…
OpenAI paused work on Astra, its next major model, after internal evaluations indicated it may have crossed the 'Critical' cybersecurity threshold in the company's Preparedness Framework, marking the …
OpenAI has confirmed the existence of its next major model, Astra, which has solved 10 major open math problems and developed advanced cyber capabilities, prompting the company to pause some internal …
OpenAI is pausing work on its upcoming AI model Astra after internal evaluations showed 'significant advancements in agentic coding and cybersecurity' that could reach a 'Critical' threshold, triggeri…
Anthropic's LLM Claude Mythos discovered cryptanalytic attacks on the post-quantum signature scheme HAWK and reduced-round AES-128, but the attacks are not practical and pose no threat to full AES. Th…
AI models remain unreliable for workplace tasks, with METR (Model Evaluation and Threat Research) reporting that AI can complete tasks taking humans 16 hours only 50% of the time, a caveat highlighted…