cd/entity/Artificial Analysis· home› entities› Artificial Analysis
grep -l @artificial analysis /news/*.json | wc -l → 460

Artificial Analysis

mentions 460 type Person page 22/23 feed RSS

// recent coverage 460 mentions

00:00
2026-06-20
runagentrun.co.uk
artificial-intelligence

AA-Briefcase: a tougher test for agents

Artificial Analysis released AA-Briefcase, a new agentic benchmark for long-horizon knowledge work, on June 18, 2026. Claude Fable 5 leads the leaderboard with 1587 Elo at $31 per task, while open-wei…

23:09
2026-06-19
mukulsingh105.github.io
artificial-intelligence

Knowledge workers don't need frontier models

A new architecture using a nano-model router that dispatches knowledge-worker tasks to either a frontier model or a small, cheap model achieves near-frontier quality at a fraction of the cost, ranking…

16:11
2026-06-19
arrowtsx.dev
large-language-models

GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

New benchmarks reveal that larger AI models like GPT-5.5 and DeepSeek V4 Pro hallucinate significantly more than smaller open-weight models, with GLM-5.2 achieving a 28% hallucination rate compared to…

15:30
2026-06-18
yaroslavvb.github.io
artificial-intelligence

How far behind is open-source AI?

The gap between open-source and proprietary AI models has narrowed from nearly 10 months in December 2024 to just 2–3.5 months as of early 2025, according to an analysis of the open-source Pareto fron…

11:01
2026-06-18
lesswrong.com
large-language-models

How far do open weights trail the frontier?

A new analysis using Epoch's ECI metric shows that open-weight AI models continue to trail closed models on the frontier, with the gap persisting over time. The analysis, based on item response theory…

10:16
2026-06-18
dev.to
large-language-models

What GLM-5.2 Changes for Long-Horizon Coding

Zhipu AI released GLM-5.2, a large language model with a 1M-token context window, flexible effort levels, and an MIT license, targeting long-horizon coding tasks. The model introduces IndexShare, an a…

09:16
2026-06-18
sebastianraschka.com
large-language-models

GLM-5.2 and IndexShare for Long-Context Sparse Attention

Z.ai released GLM-5.2, an open-weight model that the author calls the best open-weight model available. The model introduces IndexShare, a cross-layer reuse trick for DeepSeek Sparse Attention that re…

07:57
2026-06-18
dev.to
large-language-models

Nemotron 3 Ultra went live June 4. Here's the call that works.

NVIDIA released Nemotron 3 Ultra on June 4, 2026, a 550-billion-parameter open-weights model that achieves the highest intelligence score among US open models. The model uses a hybrid Mamba-Transforme…

23:58
2026-06-17
simonwillison.net
large-language-models

GLM-5.2 is probably the most powerful text-only open weights LLM

Chinese AI lab Z.ai released GLM-5.2, a 753B-parameter open-weights text-only LLM with a 1 million token context window, under an MIT license. The model leads the Artificial Analysis Intelligence Inde…

20:24
2026-06-15
aws.amazon.com
artificial-intelligence

Introducing Gemma 4 models on Amazon Bedrock

Amazon Bedrock announced the availability of Gemma 4 models, a family of open-weight AI models from Google DeepMind, including dense and mixture-of-experts variants with built-in reasoning, function c…

00:00
2026-06-15
runagentrun.co.uk
large-language-models

Nemotron 3 Ultra: America's best open model

NVIDIA released Nemotron 3 Ultra, a 550-billion-parameter open-weights reasoning model, on June 4, 2026. It is the best US open model by Artificial Analysis's scoring but trails Chinese leader Kimi K2…

00:04
2026-06-14
artificialanalysis.ai
large-language-models

Frontier Language Model Intelligence, over Time

Artificial Analysis released its Frontier Language Model Intelligence index, tracking performance, cost, and execution time of leading AI models over time. The index evaluates models on agentic tasks,…

00:00
2026-06-13
runagentrun.co.uk
ai-infrastructure

NVIDIA Blackwell tops the first agentic AI benchmark

NVIDIA's Blackwell platform achieved up to 20x more agents per megawatt than the previous generation in the first AgentPerf benchmark, a new test from independent firm Artificial Analysis designed for…

← prev page 22 / 23 next →
// co-occurs with top 8 entities
// topics top 6 topics