We Raced Seven AI Models to RCE
In a controlled race, OpenAI's GPT 5.6 Sol achieved verified remote code execution in 9 of 11 vulnerable-target runs, the most among seven AI models tested, while Zhipu AI's GLM 5.3 verified 8 runs. T…
In a controlled race, OpenAI's GPT 5.6 Sol achieved verified remote code execution in 9 of 11 vulnerable-target runs, the most among seven AI models tested, while Zhipu AI's GLM 5.3 verified 8 runs. T…
A new census of open-weight model hosting reveals that the same model name can arrive at different quantizations, context lengths, and output caps depending on the provider, with some hosts serving GL…
Tencent released Hy4 preview on August 28, a 770B parameter open-weight model under Apache 2.0, with only 49B active parameters per token and a 1 million token context window, targeting software engin…
Apple's upcoming M5 Ultra, with 512GB memory and 1.2TB/s bandwidth, is projected to cost $15,000–17,000 when it ships in October, making local AI inference for models like DeepSeek V4 Pro, Kimi K3, an…
Opus 5 topped a new benchmark called Lockstep, scoring highest among 10 leading LLMs tested on their ability to run logic circuits in their heads, with GPT 5.5 second and Kimi K3 and DeepSeek V4 Pro t…
DeepSeek V4 Flash scored 30/42 on IMO 2026 problems in Cline, clearing the 29-point gold medal cutoff for just $0.12, roughly 140 times cheaper than Claude Fable 5. The benchmark, run by Cline, tested…
NVIDIA's Vera Rubin NVL72 systems deliver up to 30x higher throughput per megawatt than the GB300 NVL72 on agentic AI workloads, according to new measured performance data from NVIDIA using the SemiAn…
A benchmark of 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun shows Fable 5, run via claude-code at high effort, achieved the best validated result of 2,726 tokens wit…
OpenAI cut API prices for its flagship GPT-5.6 Sol model by more than 20% for three months starting August 21, reducing input tokens from $5 to $4 per million and output tokens from $30 to $20 per mil…
A developer has forked OpenCode's harness to create an open-source project that enables node-based AI workflows, allowing multiple agents with designated roles to collaborate instead of using one prom…
Dromeas, a code review tool, compared its LLM council-based review against Claude Code's ultrareview on a large, real pull request from the open-source project openclaw. Dromeas's three-model council …
In a benchmark of 10 leading AI models reviewing enterprise test case documentation, 9 out of 10 fell into 'Abstract Complacency,' merely summarizing headings without verifying claims. Only Kimi K3 ac…
On August 14, Cloudflare added DeepSeek V4 Pro and DeepSeek V4 Flash to Workers AI, both with a 1,048,576-token context window, the first models on the platform to cross the 1M mark. The Flash model (…
Qwen 3.8 27B scored 52 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Luna and trailing GLM-5.2 (753B) by one point despite being 27 times smaller. The model's performance enables dep…
Alibaba's Qwen3.8-27B, a 27-billion-parameter dense model released under Apache 2.0, scored 52 on Artificial Analysis' Intelligence Index v4.1.1, tying with OpenAI's GPT-5.6 Luna at max reasoning effo…
Open-source developer Tiger380 published a benchmark report claiming its J-Space text harness improved DeepSeek V4 Pro's performance on reasoning and agent tasks, with scores rising from 60.0 to 67.7 …
In Kai, a Liar's Dice game, an engineer found that large language models sometimes challenge a bid they know is true, losing the round. The issue was traced to the action schema, where the 'challenge'…
Microsoft pushed MAI-Code-1.1-Flash into GitHub Copilot on August 11, 2026, with a 73% price cut, native vision support, and benchmark gains, and will retire MAI-Code-1-Flash from all Copilot surfaces…
DeepSeek released DeepSeek Harness, an agentic coding framework in developer preview alongside the DeepSeek V4 Pro model, which uses a plugin architecture and a local web app interface. Installation r…
State-of-the-art AI models are two-thirds smarter than last November, with two new models released every three days, yet 84% of tokens on OpenRouter are not state of the art, and the six most-used mod…