Anthropic's New AI Just Solved 3D
Anthropic's Sonnet 5.5 scored 70.6% on Terminal-Bench 4.0, up from Sonnet 5's 10.3%, and generates outputs over 30% faster while cutting total task costs by up to 30%, according to the article. The mo…
Anthropic's Sonnet 5.5 scored 70.6% on Terminal-Bench 4.0, up from Sonnet 5's 10.3%, and generates outputs over 30% faster while cutting total task costs by up to 30%, according to the article. The mo…
Anthropic launched Claude Sonnet 5.5 on Monday at $2 per million input tokens and $10 per million output tokens, half the price of Opus 5.5 and the same as Sonnet 5, with Anthropic claiming it runs 30…
Anthropic released Claude Sonnet 5.5 on September 28th, claiming the model generates output at least 30% faster than Sonnet 5 and costs up to 30% less per task by completing work with fewer tokens, wi…
Independent developer Tomoyuki Muranaka ran a September 26th simulation race in which four AI models each designed a wheel-free creature from rigid parts, hinge joints and motors to cross a hidden Thr…
Anthropic's Claude prompt caching, priced as of 09/28/2026, charges 1.25x the input price for a five-minute cache write and 2x for a one-hour write, while cache reads cost 0.1x input on most models, 0…
An engineer benchmarked roughly 4,900 prompt sessions across five Claude models and seven local Ollama models to measure how spelling, grammar and punctuation errors affect LLM accuracy. Heavy misspel…
A developer's code review of the jev-gateway router for Claude Code found that a single ternary in src/adapters/messages.ts forces the gateway into "hint" mode for every real Claude Code request, beca…
A developer built beans-picker, an MCP server layered on top of cua-driver that lets Claude Code operate macOS apps in the background, and reports that a typical benchmark task ran 1.8x faster and cos…
A developer reported that Claude Code added Claude co-author attribution trailers to 25 commits in a single day — 10.5 hours of work — plus 15 earlier commits on 22 and 24 August, for 40 total, despit…
A developer's analysis of five days of Claude Code transcripts found that Opus 5.5, despite listing at twice Sonnet 5's per-token price, cost roughly 0.45 times as much per unit of comparable coding w…
A developer's cost analysis of production AI agents finds that per-token pricing dramatically understates real spend, with bills landing 5–10x above naive estimates. The breakdown attributes the gap t…
Anthropic launched its invite-only Life Sciences Verification Program on September 17, a gated beta granting qualifying organizations access to its biology-focused models including Mythos 5.1, Opus 5,…
AkitaOnRails published Part 1 of a new LLM Benchmark v4 on September 15, 2026, after rejecting an earlier v3 suite in which 24 of 37 tested models scored between 95 and 100 points. The author restarte…
Anthropic cut Fable 5.1's cache read price to $0.25 per million tokens — a quarter of Fable 5's $1 per MTok and half of Opus 5's $0.50 per MTok — extending the break-even for keeping a prompt cache wa…
Munder Difflin 0.5.2, a free and open source desktop app from harnessmd.com, runs Claude Code agents as a persistent "office" on a user's own machine, adding an orchestrator, per-agent memory.md files…
A developer proposes replacing dollar-denominated AI token costs with a relative unit called "Human-Time," anchored to the fully loaded cost of employing a mid-level worker at roughly $50K per year, s…
Sakana AI released Fugu Max and Fugu Ultra v2, two new models in its Sakana Fugu family, live today through its OpenAI-compatible API with no open weights and no EU/EEA availability. Fugu Max is price…
OpenAI's GPT-6 Astra, released last week, has prompted an Anthropic user and developer to resubscribe to Codex, calling it a level above Claude models. The user notes that Anthropic's model documentat…
Anthropic's Claude Fable 5.1 and Mythos 5.1 system card reports that the models are the most capable publicly available AI models at release, with Mythos 5.1 falling short of CB-2 classification and a…
A developer's experiment with Claude's Sonnet 5 medium thinking shows that prepending a 'low-skill' user profile to a system design prompt causes the model to omit advanced concepts like analytics tra…