AA-Briefcase: a tougher test for agents
Artificial Analysis released AA-Briefcase, a new agentic benchmark for long-horizon knowledge work, on June 18, 2026. Claude Fable 5 leads the leaderboard with 1587 Elo at $31 per task, while open-wei…
Artificial Analysis released AA-Briefcase, a new agentic benchmark for long-horizon knowledge work, on June 18, 2026. Claude Fable 5 leads the leaderboard with 1587 Elo at $31 per task, while open-wei…
A new architecture using a nano-model router that dispatches knowledge-worker tasks to either a frontier model or a small, cheap model achieves near-frontier quality at a fraction of the cost, ranking…
New benchmarks reveal that larger AI models like GPT-5.5 and DeepSeek V4 Pro hallucinate significantly more than smaller open-weight models, with GLM-5.2 achieving a 28% hallucination rate compared to…
A new benchmark from Artificial Analysis, the AA-Briefcase, reveals that even the best AI model, Claude Fable 5, fully solves only 3 percent of realistic knowledge work tasks. The benchmark tests mode…
Zhipu AI's GLM-5.2 open-weight model has passed multiple real-world 'vibe checks,' with Jeremy Howard and Artificial Analysis rating it comparable to GPT 5.5 and Opus 4.8, marking a shift from benchma…
The gap between open-source and proprietary AI models has narrowed from nearly 10 months in December 2024 to just 2–3.5 months as of early 2025, according to an analysis of the open-source Pareto fron…
A new analysis using Epoch's ECI metric shows that open-weight AI models continue to trail closed models on the frontier, with the gap persisting over time. The analysis, based on item response theory…
Zhipu AI released GLM-5.2, a large language model with a 1M-token context window, flexible effort levels, and an MIT license, targeting long-horizon coding tasks. The model introduces IndexShare, an a…
Z.ai released GLM-5.2, an open-weight model that the author calls the best open-weight model available. The model introduces IndexShare, a cross-layer reuse trick for DeepSeek Sparse Attention that re…
NVIDIA released Nemotron 3 Ultra on June 4, 2026, a 550-billion-parameter open-weights model that achieves the highest intelligence score among US open models. The model uses a hybrid Mamba-Transforme…
Chinese AI lab Z.ai released GLM-5.2, a 753B-parameter open-weights text-only LLM with a 1 million token context window, under an MIT license. The model leads the Artificial Analysis Intelligence Inde…
A developer launched Models Pie, a tool that ranks LLMs by user-defined tradeoffs between speed, cost, and quality using data from BenchLM, OpenRouter, and Artificial Analysis. Users adjust priorities…
Z ai's GLM-5.2, a 744B-parameter open weights model with 40B active parameters, achieved a score of 51 on the Artificial Analysis Intelligence Index v4.1, surpassing competitors MiniMax-M3 and DeepSee…
Amazon Bedrock announced the availability of Gemma 4 models, a family of open-weight AI models from Google DeepMind, including dense and mixture-of-experts variants with built-in reasoning, function c…
Artificial Analysis released AgentPerf, the first benchmark for agentic-AI infrastructure, which replays recorded multi-step agent trajectories instead of single chat completions. NVIDIA reported that…
NVIDIA released Nemotron 3 Ultra, a 550-billion-parameter open-weights reasoning model, on June 4, 2026. It is the best US open model by Artificial Analysis's scoring but trails Chinese leader Kimi K2…
Artificial Analysis released its Frontier Language Model Intelligence index, tracking performance, cost, and execution time of leading AI models over time. The index evaluates models on agentic tasks,…
NVIDIA's Blackwell platform achieved up to 20x more agents per megawatt than the previous generation in the first AgentPerf benchmark, a new test from independent firm Artificial Analysis designed for…
Artificial Analysis, an independent AI benchmarking platform, launched its Coding Agent Benchmarks and Index at a June 11 event in San Francisco featuring speakers from Cognition, Cursor, and NVIDIA. …
Artificial Analysis released initial results for its AA-AgentPerf benchmark, showing NVIDIA's Blackwell systems outperforming AMD's Instinct MI355X GPUs on power-efficient agentic inference using Deep…