Model Intelligence, Cost, and Speed
Artificial Analysis released a comparison of 13 frontier AI models, ranking them by fidelity, cost, and speed, with GPT-5.6 Sol and DeepSeek V4 Pro 0813 achieving the highest fidelity (99-100), while …
Artificial Analysis released a comparison of 13 frontier AI models, ranking them by fidelity, cost, and speed, with GPT-5.6 Sol and DeepSeek V4 Pro 0813 achieving the highest fidelity (99-100), while …
A developer detailed their criteria for selecting AI models on Mac hardware, emphasizing the trade-off between benchmark scores and token generation efficiency. They highlighted Qwen3.8 27B as an exam…
OpenAI's August 13 limited API preview of Ultrafast for GPT-5.6 Sol reports up to 750 output tokens per second, as much as 14 times Standard speed, with a Cerebras benchmark showing a 5.59-times end-t…
Microsoft AI's MAI-Image-2.6-Preview tops the Artificial Analysis Image Editing Leaderboard with an Elo of 1,284, surpassing OpenAI's GPT Image 2 (high) at 1,259 and Reve 2.1 at 1,260, based on 13,492…
Chinese AI companies, including Alibaba Group Holding, MiniMax, and ByteDance, dominate the AI video generation market, occupying eight of the top 10 positions on Artificial Analysis' text-to-video le…
Fidian's Terminal-Bench variant TB-fn, built from the same 89 tasks as TB-2.1, separates models that appear comparable on the standard benchmark, narrowing the top tier to OpenAI and Anthropic models.…
Perplexity's Search API took first, second and third place in Artificial Analysis' August 27th Search Index test, with its medium setting scoring 80 at about $0.091 per task, beating rivals while land…
Perplexity's medium Search API configuration scored 80 on the new Artificial Analysis Search Index, a 47-point improvement over the 33-point baseline without search, and five points ahead of Parallel …
Artificial Analysis, in partnership with Liquid AI, has launched a benchmark suite for pocket-scale AI models that fit within 8 GB of memory after quantization, including KV cache at 8K context, measu…
OpenAI's new small model gpt-5.6-luna delivers high speed and low cost, with API costs in the tens of cents for complex tasks, making consumer AI apps more viable. The author, Calvin French-Owen, argu…
Z.ai released GLM-5.3-Flash, an open-source model with 320 billion parameters that scores just three points behind the larger GLM-5.3 on Artificial Analysis's Intelligence Index, at a seventh of the c…
Z.ai's GLM-5.3-Flash model, initially released anonymously on OpenRouter as Ox Alpha, has processed over 20 trillion tokens in its first six days, according to OpenRouter. The model features a mixture…
Fal, an inference-infrastructure company, announced on August 27, 2026, that its new H3 Max model generates five seconds of AI video with audio in under three seconds, claiming roughly 35x the through…
On June 12, 2026, Artificial Analysis, in collaboration with NVIDIA, released AA-AgentPerf, a benchmark that introduces 'Agents per Megawatt' as a primary metric for agentic AI hardware performance. I…
Z.ai released GLM-5.3-Flash, a 320B-total-parameter mixture-of-experts model with 18B active parameters, a 1,048,576-token context window, and native image and video input, under an MIT license with w…
Google introduced Gemini 3.5 Transcribe, its most precise speech-to-text model, available via the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform. The model achieves a 4.0% word er…
Z AI released GLM-5.3-Flash on August 26, 2026, scoring 57 on the Artificial Analysis Intelligence Index, well above the median of 18 for comparable reasoning models. Priced at $0.15 per 1M input toke…
Speechify's SIMBA 3.2 text-to-speech model ranks #1 on the Artificial Analysis TTS leaderboard and Voice Arena's real-time category, beating ElevenLabs, OpenAI, and Google DeepMind in blind listening …
Artificial Analysis updated its Coding Agent Index with reward hacking corrections from Terminal-Bench v2.1, which assigns zero scores to AI models that game task completion without doing the work. Th…
Liquid AI released Pipette, an open-source benchmarking platform for foundation models on edge devices, developed with Artificial Analysis as an independent methodology validator. Pipette measures on-…