{"slug": "llm-benchmark-opus-5-e-bom", "title": "LLM Benchmark: Opus 5 é bom?", "summary": "Anthropic released Claude Opus 5 yesterday, which scored 95/100 (Tier A) in the author's benchmark, the highest score seen so far. However, the author cautions that this does not prove it is the best LLM overall, citing a previous article explaining why a higher score does not mean the best model.", "body_md": "A Anthropic lançou o **Claude Opus 5** ontem. A pergunta óbvia é:\n\n“É bom?”\n\nResposta curta: **sim, é muito bom**. No meu benchmark fez 95/100, Tier A, com a engenharia mais completa que apareceu até agora em qualquer um dos harnesses que testei.\n\nAgora a resposta que interessa: não, isso não prova que virou “o melhor LLM do mundo”. Nem prova que é melhor que Fable 5, Opus 4.8, GPT 5.6 Sol ou Kimi K3 em qualquer trabalho que você jogar neles. Semana passada publiquei um artigo inteiro explicando [por que a maior nota não significa o melhor modelo](https://www.akitaonrails.com/2026/07/19/llm-benchmark-devo-usar-o-que-tem-nota-maior/). O Opus 5 chegou a tempo de produzir um belo estudo de caso praquele texto.", "url": "https://wpnews.pro/news/llm-benchmark-opus-5-e-bom", "canonical_source": "https://www.akitaonrails.com/2026/07/25/llm-benchmark-opus-5-e-bom/", "published_at": "2026-07-25 12:00:00+00:00", "updated_at": "2026-08-21 04:15:39.188616+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-products"], "entities": ["Anthropic", "Claude Opus 5", "Fable 5", "Opus 4.8", "GPT 5.6 Sol", "Kimi K3"], "alternates": {"html": "https://wpnews.pro/news/llm-benchmark-opus-5-e-bom", "markdown": "https://wpnews.pro/news/llm-benchmark-opus-5-e-bom.md", "text": "https://wpnews.pro/news/llm-benchmark-opus-5-e-bom.txt", "jsonld": "https://wpnews.pro/news/llm-benchmark-opus-5-e-bom.jsonld"}}