Grok 4.6 Is Out: What Developers Need to Know
XAI released Grok 4.6 on August 12, 35 days after Grok 4.5, priced at $2 per million input tokens and $6 per million output tokens, making it 2.5x to 5x cheaper than GPT-5.6 Sol and Claude Opus 5 at c…
XAI released Grok 4.6 on August 12, 35 days after Grok 4.5, priced at $2 per million input tokens and $6 per million output tokens, making it 2.5x to 5x cheaper than GPT-5.6 Sol and Claude Opus 5 at c…
A gamer used Anthropic's Claude Opus 5 to build FISH, an AI study helper for the game The Finals that converts YouTube pro guides into split-second decision drills, and reports it is improving their p…
Qwen 3.8 Max scored 92 on the LLM Coding Benchmark v2, tying with GLM 5.2, Kimi K2.5, Gemini 3.6 Flash, and Grok 4.6, while GLM 5.3 reached 94, Gemini 3.7 Flash scored 93, and Grok 4.6 scored 92, acco…
Anthropic disclosed in a risk report covering the period up to July 15, 2026, that it has an internal model called Model 2, which is somewhat more capable than Claude Mythos 5 and is used heavily for …
Mixedbread shipped Toast 1, a specialized search agent that matches Claude Opus 5 and GPT-5.6 Sol on deep-search benchmarks while running up to 10× cheaper and 12× faster, according to vendor-run test…
Anthropic launched Claude Opus 5 at half the price of its flagship, topping benchmarks, but developers on Hacker News are asking for the older Opus 4.8 back, citing that the new model stops asking cla…
Google DeepMind released Gemini 3.7 Flash, which scores 56 on the Artificial Analysis Intelligence Index with high reasoning, a 4-point improvement over Gemini 3.6 Flash, and achieves an average Time …
Meta launched Muse Code, a terminal coding agent that defaults to training on user code unless switched to the Standard tier, which costs $1.25 per million input tokens and $4.25 per million output to…
Sanderland's open-source ctok tool reconstructs Anthropic's Claude tokenizer, estimating 49,152 vocabulary entries in Claude 3 and 16,384 in Claude 4.7, with some reserved for image tokens. Claude 5's…
XAI released Grok 4.6, an incremental upgrade to its flagship language model, which now ties with GPT-5.6 Soul on the Artificial Analysis Intelligence Index and tops GDPval and Harvey Bench, while pri…
Mixedbread AI released Toast 1, a specialized search agent that matches or outperforms Claude Opus 5 and GPT-5.6 Sol while being up to 10× cheaper and 12× faster. In Databricks' OfficeQA Pro V2 benchm…
Clixad, a free AI coding agent for the terminal, launched with a business model funded by an offerwall, giving users 5,000 credits on signup and allowing them to earn more by completing advertiser for…
Ruby on Rails published the first public, same-harness benchmark of frontier models performing real Rails tasks, testing 8 models across 21 atomic tasks with 504 runs at a total cost of $491. Claude O…
Netlify's Agent Runners benchmark of 11 AI models on the same website-building brief found a roughly 200x cost spread, with Claude Opus 5 averaging 519 credits per run (spiking to 1,055) versus DeepSe…
Alibaba Group Holding has released the core files of its flagship AI model Qwen3.8-Max for free download, but introduced commercial restrictions requiring large companies to obtain a separate paid lic…
DeepSeek's V4 Pro 0813 (max) model scored 53 on the Artificial Analysis Intelligence Index, one point ahead of DeepSeek V4 Flash 0731, according to an independent evaluation by Artificial Analysis. Th…
Grok 4.6, developed by xAI, scores 61 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Sol and trailing Claude Opus 5 by two points, while priced at $2 per 1M input tokens and $6 per 1M…
Grok 4.6, developed by xAI, matches GPT-5.6 Sol's performance on agentic tasks while costing 60% less, completing complex tasks in roughly 53 steps compared to Claude Opus 5's 103 steps. The efficienc…
Anthropic's Claude Opus 5 performs worse with over-specific, step-by-step prompts, according to the Claude Code team, which has been removing large portions of its system prompt with each new model re…
Dave Fisher, founder of Revenant Systems, reported that in a gauntlet of over 5,000 prompt injection attempts across 42 LLMs, 12 of 16 frontier models made 34 fraudulent tool calls that would have sen…