Can LLMs work in the wet lab?
Benchling released BenchBench-Protocol, a benchmark built from thousands of real-world experiments, showing that Anthropic's Opus 5 leads at 59.2%, followed by OpenAI's GPT 5.6 at 47.1%, and open-sour…
Benchling released BenchBench-Protocol, a benchmark built from thousands of real-world experiments, showing that Anthropic's Opus 5 leads at 59.2%, followed by OpenAI's GPT 5.6 at 47.1%, and open-sour…
Anthropic apologized after secretly routing paying Fable 5 customers to the cheaper Opus 4.8 model while billing them for the premium tier, a practice that sparked developer backlash and was fixed onl…
A developer discovered that Claude Code subagents pinned to a cheaper model via frontmatter were silently running on the session model, Fable, after a release dropped the frontmatter layer, causing un…
Anthropic's Opus 5 model produces verbose output, and a developer offers four methods to fix it, from writing rules in CLAUDE.md to using output styles. The developer explains that output styles, whic…
Notion's Knowledge Board, a live evaluation using anonymized traffic and judged by models from Anthropic, OpenAI, and Google, shows Opus 5 leading with a 98.4% resolution rate on knowledge work tasks,…
Anthropic's Opus 5, despite being more capable than Opus 4.7 and Opus 4.8 and rivaling Fable in benchmarks, feels like a downgrade to work with because it makes assumptions and reinterprets plans with…
Jeremy Berman reported scoring 96.2% on ARC-AGI-3 with Opus 5, and 99.3% pass@2, using a program based on Claude Code and Opus 5 (high) with one action command and filesystem logs. The approach, which…
Anthropic and alignment researchers including Caspar Oesterheld and Emery Cooper released the Conceptual Reasoning Index (CRI), a 0–100 benchmark scoring AI reasoning in domains without ground truth, …
Databricks launched Smart Routing in Unity AI Gateway, now in Beta, which automatically matches coding tasks to the most cost-effective model and harness, achieving 30%+ lower cost per task while matc…
DeepSeek V4 Pro 0813, released by DeepSeek without a launch event, scored 87.9 on Terminal Bench 2.1, up from 72.1 in its April preview, and topped Cyberjim (83.3) and Automation Bench (31.8) among ri…
JetBrains' PyCharm introduced the Agent Environment Coordinator skill, which provides AI agents with the correct Python interpreter information, boosting task success rates from 68% to 98% on average …
A developer advises against over-prompting reasoning models, noting that modern models already verify and pace themselves, so extra instructions like 'double-check your work' cause over-verification a…
Val Town, the instant deploy platform for small apps, reported 6% ARR growth in July 2026, missing its 20% target, though Pro subscriptions grew 26% month-over-month. The company added DeepSeek v4 Fla…
An experiment by a developer found that when conflicting instructions appear in a file like CLAUDE.md, the model resolves the conflict by following the rule positioned lower in the file, with position…
A developer building the task management app lyphe argues that while AI coding agents can generate impressive code, the developer's core job is designing a well-structured system with clear architectu…
Anthropic updated Claude Fable 5's biology safeguards on August 7, narrowing a classifier that had rerouted almost every biology query to Opus 5, and reported an 85% reduction in biology-related fallb…
Anthropic's Opus 4.8 and Opus 5 tied at 9/25 strict test passes on 25 tasks from Stet's repository, with Opus 4.8 writing smaller patches on 20 of 25 tasks while Opus 5 searched wider, using more shel…
Meta launched Muse Code, a terminal-based coding agent, and Muse Spark 1.2, a code-focused model, with Muse Spark 1.2 scoring 59.3% on the DeepSWE benchmark. The score places it behind GPT 5.6 Turbo a…
Anthropic deleted 80% of Claude Code's system prompt, and the tool became smarter, according to Boris Cherny, creator of Claude Code. Cherny's team strips instructions as models improve, arguing that …
Anthropic's Opus 5 model upgrade prompt, released for Claude Code, audits and upgrades existing skills through a blind-judged Gauntlet Loop, retiring those that don't outperform the base model. The pr…