NanoGPT Speedrun Frontier
A benchmark of 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun shows Fable 5, run via claude-code at high effort, achieved the best validated result of 2,726 tokens wit…
A benchmark of 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun shows Fable 5, run via claude-code at high effort, achieved the best validated result of 2,726 tokens wit…
OpenAI cut API prices for its flagship GPT-5.6 Sol model by more than 20% for three months starting August 21, reducing input tokens from $5 to $4 per million and output tokens from $30 to $20 per mil…
A developer has forked OpenCode's harness to create an open-source project that enables node-based AI workflows, allowing multiple agents with designated roles to collaborate instead of using one prom…
Dromeas, a code review tool, compared its LLM council-based review against Claude Code's ultrareview on a large, real pull request from the open-source project openclaw. Dromeas's three-model council …
In a benchmark of 10 leading AI models reviewing enterprise test case documentation, 9 out of 10 fell into 'Abstract Complacency,' merely summarizing headings without verifying claims. Only Kimi K3 ac…
On August 14, Cloudflare added DeepSeek V4 Pro and DeepSeek V4 Flash to Workers AI, both with a 1,048,576-token context window, the first models on the platform to cross the 1M mark. The Flash model (…
Qwen 3.8 27B scored 52 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Luna and trailing GLM-5.2 (753B) by one point despite being 27 times smaller. The model's performance enables dep…
Alibaba's Qwen3.8-27B, a 27-billion-parameter dense model released under Apache 2.0, scored 52 on Artificial Analysis' Intelligence Index v4.1.1, tying with OpenAI's GPT-5.6 Luna at max reasoning effo…
Open-source developer Tiger380 published a benchmark report claiming its J-Space text harness improved DeepSeek V4 Pro's performance on reasoning and agent tasks, with scores rising from 60.0 to 67.7 …
In Kai, a Liar's Dice game, an engineer found that large language models sometimes challenge a bid they know is true, losing the round. The issue was traced to the action schema, where the 'challenge'…
Microsoft pushed MAI-Code-1.1-Flash into GitHub Copilot on August 11, 2026, with a 73% price cut, native vision support, and benchmark gains, and will retire MAI-Code-1-Flash from all Copilot surfaces…
DeepSeek released DeepSeek Harness, an agentic coding framework in developer preview alongside the DeepSeek V4 Pro model, which uses a plugin architecture and a local web app interface. Installation r…
State-of-the-art AI models are two-thirds smarter than last November, with two new models released every three days, yet 84% of tokens on OpenRouter are not state of the art, and the six most-used mod…
Google and DeepSeek released major AI model upgrades on Thursday, with Google unveiling Gemini 3.7 Flash and China's DeepSeek launching the general-availability version of V4 Pro, as both companies pu…
Google's Made by Google 2026 hardware event delivered an awkward keynote, while new AI models from xAI (Grok 4.6) and DeepSeek (DeepSeek V4 Pro) were released, posing competitive pressure on Google. T…
DeepSeek V4 Pro, the 0813 release from DeepSeek, costs 43 cents per million input tokens and 87 cents per million output tokens via API, making it roughly 57 times cheaper than Gemini 3 Pro's $10.50 p…
Nvidia is developing Nemotron 4, an open-weight model family with at least one trillion parameters, aiming to compete with the best freely available models, according to The Information. The company h…
DeepSeek V4 Pro solved 8 of 16 Hack The Box challenges in the HTB-Challenger Benchmark, scoring 36.2% with no false positives, but its median cost per challenge was $0.35, comparable to GPT-5.6 Luna d…
A red-team audit of Pipe's AI agent sandbox by an unnamed developer found that all 23 escape vectors across five layers failed, but a follow-up round discovered three real ratchet escapes (empty white…
DeepSeek V4 Flash ranked first in OpenRouter's weekly model-usage ranking for July 27 to Aug. 2, processing 7.22 trillion tokens on the multi-model aggregation platform. On Aug. 1, the model processed…