What Independent Benchmarks Say About Opus 5.5
Two independent evaluations of Anthropic's Opus 5.5 disagree on its standing: Artificial Analysis's Intelligence Index places it first at 58 points, ahead of OpenAI's GPT-6 Astra and Anthropic's Fable…
Two independent evaluations of Anthropic's Opus 5.5 disagree on its standing: Artificial Analysis's Intelligence Index places it first at 58 points, ahead of OpenAI's GPT-6 Astra and Anthropic's Fable…
Mistral Large 4 Preview, released October 6, 2026, scored 38 on the Artificial Analysis Intelligence Index, above the median of 26 for comparable models, according to Artificial Analysis. The model is…
Mistral released Mistral Large 4, a 1 trillion-parameter open-weights multimodal reasoning model it calls "Le Chonk," with weights expected on Hugging Face and other repositories within the month. The…
A developer argues that hard-coding model names as pinned dependencies is a mistake, citing GPT-6.1 Sol replacing GPT-6 Sol after just seven days and Google's Gemini 4 Argon raising its output token l…
Mistral launched a public preview of Mistral Large 4, a 1-trillion-parameter natively multimodal model with 49 billion active parameters, on Mistral Studio, with weights to be released by the end of t…
Reflection announced Beam, a text-only 501B-total / 23B-active mixture-of-experts model for coding, agentic and scientific work, trained from scratch on 23.8T pretraining tokens with full weights unde…
Google's Gemini 4 Argon, run in Google's Antigravity CLI, scored 64 on the Artificial Analysis Coding Agent Index (v1.5), one point ahead of OpenAI's GPT-6.1 Sol (xhigh) running in Codex at 63, accord…
Independent benchmarking by Artificial Analysis found that xAI's Grok 4.7, despite keeping Grok 4.6's $2 per million input and $6 per million output token pricing, roughly doubled cost per task from $…
Top AI models answer roughly 1 in 4 specific-fact questions incorrectly when they cannot search, according to Artificial Analysis data cited by Randy Olson. The AA-Omniscience dataset covers 6,000 que…
Microsoft launched MAI-Transcribe-2-Streaming on October 1, a WebSocket-based streaming speech-to-text model that ranked first for both Final Transcript and First Partial Transcript accuracy across 38…
Anthropic's Claude Sonnet 5.5 and Opus 5.5 both landed in the same roughly $6 to $8 cost-per-task range on Artificial Analysis's blended intelligence-versus-cost tracking, despite Sonnet 5.5's lower p…
Microsoft AI released MAI-Transcribe-2-Streaming, a real-time transcription model that Microsoft says ranks first for accuracy on Artificial Analysis, transcribes 60 languages, and returns first parti…
Google announced Gemini 4 Argon on September 30, 2026, a frontier model with a 1-million-token output limit, up from the previous Gemini ceiling of 64,000, and introductory pricing of $2 per million i…
Self-hosting the open-weight Qwen3.8 27B model costs $619 per full Artificial Analysis Intelligence Index run through the cheapest zero-data-retention provider, versus $67 for OpenAI's GPT-6 Luna, acc…
Google announced Gemini 4 Argon on Wednesday, an invite-only model priced at $2 per million input tokens and $10 per million output, doubling to $4/$20 after the introductory period, with access limit…
Google DeepMind announced Gemini 4 Argon, a frontier model for long-horizon coding and vulnerability patching with a 1M-token context limit, with access starting only for vetted defenders in the Fairw…
Google's Gemini 4 Argon scored 53 on the Artificial Analysis Intelligence Index, tying OpenAI's flagship for second place among AI labs behind Anthropic's top models, according to Artificial Analysis.…
Google DeepMind released Gemini 4 Argon, a model it says takes first place on 13 of 19 published benchmarks against GPT-6 Astra and Claude Opus 5.5, with access initially limited to government users a…
Artificial Analysis's AA-Omniscience benchmark found Google's Gemini 4 Argon, launched September 30, posts a 15% hallucination rate — the lowest of any model scoring 45 or above on the firm's Intellig…
Google's Gemini 4 Argon scored 53 on Artificial Analysis's Intelligence Index (v4.3.2), tying OpenAI's GPT-6 Astra (max) and Claude Fable 5.1 (max with fallback) for third place, according to Artifici…