OpenAI Just Made Analytics 10x Cheaper
OpenAI's price cut for GPT 5.6 Luna by 80% has made AI-powered analytics answers cost less than half a penny each, with Luna achieving 99.8% accuracy on an agentic SQL benchmark at a 5x lower price th…
OpenAI's price cut for GPT 5.6 Luna by 80% has made AI-powered analytics answers cost less than half a penny each, with Luna achieving 99.8% accuracy on an agentic SQL benchmark at a 5x lower price th…
Google has introduced Gemini 3.7 Flash, a new AI model that costs half the original price of its predecessor, Gemini 3.6 Flash, at $0.75 per million input tokens and $3.75 per million output tokens th…
Google shipped Gemini 3.7 Flash on August 13, three weeks after Gemini 3.6 Flash, with benchmark gains in agentic coding tasks such as DeepSWE v1.1 jumping from 49.0% to 65.3%. The launch price is $0.…
A new AI model scored 65% on the DeepSWE benchmark, but it is not the model Google promised, according to The New Stack. The article highlights the discrepancy between the announced model and the one …
DeepSWE, a new long-horizon software engineering benchmark, reports that Grok 4.6 /medium outperforms /high effort on its leaderboard, which measures frontier coding agents on original tasks across 91…
DeepSeek released the production version of its flagship model, DeepSeek-V4-Pro-0813, on August 13, ending a preview that began April 24, with the API alias now resolving to the new build. The model s…
DeepSeek released V4-Flash-0731 on July 31, 2025, and the smaller model now outperforms the larger V4-Pro-Preview across all nine published agentic benchmarks, including Terminal Bench 2.1 (82.7 vs. 7…
LoopTroop, an open-source GUI for long AI coding tasks, has evolved its architecture to distinguish between discarding state that may belong to a failure and preserving state that is still plausibly v…
Meta launched Muse Code, a terminal-based coding agent, and Muse Spark 1.2, a code-focused model, with Muse Spark 1.2 scoring 59.3% on the DeepSWE benchmark. The score places it behind GPT 5.6 Turbo a…
DeepSeek-V4 Flash 0731, at $0.10 per rollout, is the cheapest model on the DeepSWE board and, when used first with escalation to GPT-5.6 Luna on failure, solves 78.9% of tasks at $0.385 each—more accu…
Meta has shipped Muse Code, its first coding agent, in beta, alongside Muse Spark 1.2, a co-trained model with a 1 million-token context window. The agent, available via the Meta Model API, is priced …
Bespoke Labs is hiring a contract researcher to design and evaluate reinforcement learning environments and benchmarks for long-horizon agentic tasks, which require hours, days, or weeks of coherent m…
DeepSeek released the V4-Flash-0731 model on July 31, a retrained 13-billion-active-parameter Mixture-of-Experts model that outperforms its own V4-Pro-Preview on all nine agent benchmarks, including a…
DeepSeek shipped the public beta of V4 Flash on July 31 with no architecture changes — same 284B MoE, 1M-token context — but a retraining-only round boosted Terminal Bench 25.8 points to 82.7, Toolath…
DeepSeek has launched the public beta of its V4-Flash API, featuring upgrades for agent tasks, scoring 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE. The update adds Responses API support and Codex a…
Kimi K3 edges GPT-5.6 Sol on pass@4 (89.4% vs 85.8%) and costs 64% less per rollout ($4.65 vs $8.37), but GPT-5.6 Sol leads on pass@1 (72.7% vs 68.5%) and reliability (61 tasks solved on all four trie…
Together AI announced that Kimi.ai's K3 model will be hosted on its platform starting July 27, 2026, claiming the model achieves near-flagship coding performance at about 35% of the price of Anthropic…
Kimi K3 matches Claude Fable 5 on DeepSWE benchmark quality with a 68.5% pass@1 versus Fable's 69.9%, but costs $4.65 per rollout compared to Fable's $13.41, delivering 2.8x more solved tasks per doll…
Google shipped Gemini 3.6 Flash on July 21, 2026, offering 17% fewer output tokens than its predecessor and available immediately in GitHub Copilot for Pro, Pro+, Max, Business, and Enterprise users. …
Google released two production models Tuesday: Gemini 3.6 Flash, priced at $1.50 per million input tokens and $7.50 per million output tokens (a 17% reduction from Gemini 3.5 Flash's output price), an…