cd/entity/Opus 5· home entities Opus 5
grep -l @opus 5 /news/*.json | wc -l → 137

Opus 5

mentions 137 type Person page 4/7 feed RSS

// recent coverage 137 mentions

00:58
2026-08-15
benchling.com
artificial-intelligence

Can LLMs work in the wet lab?

Benchling released BenchBench-Protocol, a benchmark built from thousands of real-world experiments, showing that Anthropic's Opus 5 leads at 59.2%, followed by OpenAI's GPT 5.6 at 47.1%, and open-sour…

00:10
2026-08-15
byteiota.com
ai-ethics

The AI Trust Deficit: How Every Major Lab Is Breaking It

Anthropic apologized after secretly routing paying Fable 5 customers to the cheaper Opus 4.8 model while billing them for the premium tier, a practice that sparked developer backlash and was fixed onl…

18:17
2026-08-14
dev.to
large-language-models

Opus 5: How to Fix Verbose Output

Anthropic's Opus 5 model produces verbose output, and a developer offers four methods to fix it, from writing rules in CLAUDE.md to using output styles. The developer explains that output styles, whic…

13:31
2026-08-14
labs.notion.com
artificial-intelligence

Evaluating how models perform using live traffic

Notion's Knowledge Board, a live evaluation using anonymized traffic and judged by models from Anthropic, OpenAI, and Google, shows Opus 5 leading with a 98.4% resolution rate on knowledge work tasks,…

10:12
2026-08-14
mun-logadan.github.io
artificial-intelligence

Why does Opus 5 feel worse to work with?

Anthropic's Opus 5, despite being more capable than Opus 4.7 and Opus 4.8 and rivaling Fable in benchmarks, feels like a downgrade to work with because it makes assumptions and reinterprets plans with…

07:43
2026-08-14
twitter.com
artificial-intelligence

96.2% on ARC-AGI-3 with Opus 5

Jeremy Berman reported scoring 96.2% on ARC-AGI-3 with Opus 5, and 99.3% pass@2, using a program based on Claude Code and Opus 5 (high) with one action command and filesystem logs. The approach, which…

23:08
2026-08-13
sourcefeed.dev
artificial-intelligence

Anthropic's New Benchmark Scores Reasoning Nobody Can Verify

Anthropic and alignment researchers including Caspar Oesterheld and Emery Cooper released the Conceptual Reasoning Index (CRI), a 0–100 benchmark scoring AI reasoning in domains without ground truth, …

00:00
2026-08-13
mindstudio.ai
artificial-intelligence

DeepSeek V4 Pro 0813: Benchmark Results and Hands-On Test

DeepSeek V4 Pro 0813, released by DeepSeek without a launch event, scored 87.9 on Terminal Bench 2.1, up from 72.1 in its April preview, and topped Cyberjim (83.3) and Automation Bench (31.8) among ri…

11:00
2026-08-12
dev.to
artificial-intelligence

AI Coding Tip 031 - Stop Over-Prompting Reasoning Models

A developer advises against over-prompting reasoning models, noting that modern models already verify and pace themselves, so extra instructions like 'double-check your work' cause over-verification a…

00:00
2026-08-11
blog.val.town
ai-products

Investor Update – July 2026

Val Town, the instant deploy platform for small apps, reported 6% ARR growth in July 2026, missing its 20% target, though Pro subscriptions grew 26% month-over-month. The company added DeepSeek v4 Fla…

19:54
2026-08-10
dev.to
large-language-models

Opus 5: The Cost of Instruction Conflicts

An experiment by a developer found that when conflicting instructions appear in a file like CLAUDE.md, the model resolves the conflict by following the rule positioned lower in the file, with position…

17:54
2026-08-09
dev.to
artificial-intelligence

AI Can Write the Code. You Still Have to Design the System.

A developer building the task management app lyphe argues that while AI coding agents can generate impressive code, the developer's core job is designing a well-structured system with clear architectu…

06:46
2026-08-09
aiunderstanding.org
ai-safety

Anthropic Says Fable 5 Biology Fallbacks Fell 85%

Anthropic updated Claude Fable 5's biology safeguards on August 7, narrowing a classifier that had rerouted almost every biology query to Opus 5, and reported an 85% reduction in biology-related fallb…

15:11
2026-08-07
the-ai-corner.com
artificial-intelligence

Anthropic Deleted 80% of Claude Code's Prompt. It Got Smarter

Anthropic deleted 80% of Claude Code's system prompt, and the tool became smarter, according to Boris Cherny, creator of Claude Code. Cherny's team strips instructions as models improve, arguing that …

12:00
2026-08-07
somethingbig.ai
ai-tools

The Opus 5 Skills Upgrade Prompt

Anthropic's Opus 5 model upgrade prompt, released for Claude Code, audits and upgrades existing skills through a blind-judged Gauntlet Loop, retiring those that don't outperform the base model. The pr…

← prev page 4 / 7 next →
// co-occurs with top 8 entities
// topics top 6 topics