cd /news/artificial-intelligence/llm-benchmark-opus-5-e-bom · home topics artificial-intelligence article
[ARTICLE · art-105471] src=akitaonrails.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

LLM Benchmark: Opus 5 é bom?

Anthropic released Claude Opus 5 yesterday, which scored 95/100 (Tier A) in the author's benchmark, the highest score seen so far. However, the author cautions that this does not prove it is the best LLM overall, citing a previous article explaining why a higher score does not mean the best model.

read1 min views3 publishedJul 25, 2026

A Anthropic lançou o Claude Opus 5 ontem. A pergunta óbvia é:

“É bom?”

Resposta curta: sim, é muito bom. No meu benchmark fez 95/100, Tier A, com a engenharia mais completa que apareceu até agora em qualquer um dos harnesses que testei.

Agora a resposta que interessa: não, isso não prova que virou “o melhor LLM do mundo”. Nem prova que é melhor que Fable 5, Opus 4.8, GPT 5.6 Sol ou Kimi K3 em qualquer trabalho que você jogar neles. Semana passada publiquei um artigo inteiro explicando por que a maior nota não significa o melhor modelo. O Opus 5 chegou a tempo de produzir um belo estudo de caso praquele texto.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/llm-benchmark-opus-5…] indexed:0 read:1min 2026-07-25 ·