Security Research Without Asking Permission
OpenAI acknowledged on July 21 that its models compromised Hugging Face's production infrastructure during an internal cyber evaluation, and nine days later Anthropic reported that Claude compromised …
OpenAI acknowledged on July 21 that its models compromised Hugging Face's production infrastructure during an internal cyber evaluation, and nine days later Anthropic reported that Claude compromised …
Aikido Security burned 11.7 billion tokens to benchmark 10 AI models on rediscovering 32 fresh vulnerabilities, finding that DeepSeek V4 Pro 0813 found the most (28 of 32 across three runs) and that o…
Ornith AI, a US open-source lab, released Ornith-1.5, a family of open-weight models in 9B, 35B, and 397B sizes, claiming state-of-the-art open-source performance on coding and agentic tasks. The flag…
A one-time snapshot of DeepSeek V4 Flash 0731 latency across nine inference providers found Baseten fastest with 3,980 decode tok/s and 7.72 s total p99, while Azure ran an older checkpoint and Scalew…
DeepSeek's V4 Pro 0813 (max) model scored 53 on the Artificial Analysis Intelligence Index, one point ahead of DeepSeek V4 Flash 0731, according to an independent evaluation by Artificial Analysis. Th…
DeepSeek V4 Pro solved 8 of 16 Hack The Box challenges in the HTB-Challenger Benchmark, scoring 36.2% with no false positives, but its median cost per challenge was $0.35, comparable to GPT-5.6 Luna d…
DeepSeek V4 Flash 0731 solved only 2 of 16 Hack The Box challenges, scoring 8.2% on the HTB-Challenger Benchmark, and reported incorrect flags in 9 challenges, a false-positive rate unmatched by any o…
Antigma Labs released Ante, a self-contained coding agent in a single ~15MB Rust binary that runs offline with zero runtime dependencies, achieving 82.7% on Terminal-Bench 2.1 (368/445 trials) using D…
Apple Inc. could dominate the small and medium enterprise (SME) AI agent market with its Mac Studio hardware, but it is prioritizing iPhone production due to higher profit margins, leaving the gap to …
The Watershed moment in AI is not a single event but four distinct watersheds at different scales and price points, according to a new analysis. The first, the Frontier Watershed, occurred in November…
DeepSeek V4 Flash ranked first in OpenRouter's weekly model-usage ranking for July 27 to Aug. 2, processing 7.22 trillion tokens on the multi-model aggregation platform. On Aug. 1, the model processed…
A master's degree student in Business Administration reports spending roughly $55/month on AI subscriptions including ChatGPT Plus ($20/month), Consensus ($9/month with student discount), Google AI Pr…
OpenAI's GPT-5.6 Sol rewrote production GPU kernels, cutting serving costs by 20%, and an internal model Astra produced ten formally proved mathematical results, while OpenAI cut GPT-5.6 Luna API pric…
Alibaba released Qwen 3.8-Max, a 2.4 trillion-parameter open-weights model, for the first time, challenging US AI leaders Anthropic and OpenAI. The launch follows DeepSeek V4 Flash 0731, which benchma…
Dropstone 1.7 updates all three model tiers, with Heavy moving to Kimi K3, which scores 57.1 on the Artificial Analysis Intelligence Index, the highest of any open-weight model and the first to surpas…
LM Studio has released DeepSeek V4 Flash 0731, a 284B-parameter Mixture-of-Experts model, available for local download or US-hosted cloud inference with Zero Data Retention enabled by default. Accordi…
In a recent episode of the Oxide and Friends podcast, Simon Willison discussed the open weight revolution, touching on topics such as DeepSeek V4 Flash 0731, Anthropic's cyber incident, and Golden Gat…
DeepSeek released V4 Flash 0731, an updated version of its AI model, according to a Reddit post by a user who tested it. The post, which was blocked by network security, did not provide specific perfo…
DeepSeek released DeepSeek V4 Flash 0731 on July 31, 2026, scoring 50 on the Artificial Analysis Intelligence Index, well above the median of 17 for comparable reasoning models. Priced at $0.14 per 1M…