Kimi K3
Moonshot AI released Kimi K3, a 2.8-trillion-parameter mixture-of-experts model with 16 active experts per token, achieving the top score of 1679 on the Artificial Analysis webdev arena ahead of Claud…
Moonshot AI released Kimi K3, a 2.8-trillion-parameter mixture-of-experts model with 16 active experts per token, achieving the top score of 1679 on the Artificial Analysis webdev arena ahead of Claud…
Kimi K2, a MoE model with 1 trillion total parameters and 128k context, is available via Novita API at $0.57 per million input tokens and $2.30 per million output tokens, with 99.32% uptime. The model…
DeepSeek releases V4 Pro, a 1.6-trillion-parameter Mixture-of-Experts model with 49 billion active parameters per token, achieving a top open-weight score of 80.6% on SWE-bench Verified. The model req…
DeepSeek released V4 Flash, a 284B-parameter MoE model with 13B active parameters per token, featuring FP4+FP8 hybrid attention and 1M native context. The open-weight model, available under MIT licens…
Independent benchmarks of NVIDIA's DGX Spark desktop through mid-2026 show Qwen 3.5 27B as the most consistent all-rounder on the easy Ollama path, while GPT-OSS 120B pushes nearly 4x the throughput b…
AMD's Ryzen AI Max+ 395 processor powers several new desktop and workstation systems, including the Framework Desktop and GMKtec EVO-X2, offering 128GB unified memory and 256 GB/s bandwidth for local …
NVIDIA released Nemotron 3 Ultra, a 550-billion-parameter open-weight MoE model with hybrid Mamba-2 and Transformer architecture, under the Linux Foundation's OpenMDW-1.1 license. The model requires s…
China's Ministry of Commerce is in talks with Alibaba, ByteDance, and Z.ai to restrict overseas access to advanced AI models, including open-weight ones, mirroring recent US export controls on Anthrop…
DeepSeek released V3.1 Terminus, a 685B-parameter Mixture-of-Experts model with ~37B active parameters per token, featuring improved language consistency and tool-use capabilities including Code Agent…
DeepSeek released V3.2 Exp, a 685B-parameter mixture-of-experts model with 37B active parameters per token, featuring DeepSeek Sparse Attention (DSA) for fine-grained sparse attention. The model, buil…
Open-weights models are dramatically cheaper per token than closed models, with the five cheapest models on the Artificial Analysis pricing index all being open-weights and the five most expensive all…
GLM 5.2, a 744B-parameter Mixture-of-Experts model with ~40B active parameters per token, was released in June 2026 under an MIT license. It uses MLA and DeepSeek Sparse Attention for a 1M context win…
Alibaba releases Qwen3.6 35B A3B, a Mixture-of-Experts model with 35B total parameters and 3B active per token, featuring 256K native context and ranking #1 for agent workflows in 2026. The model achi…
Tokenstead, a new web tool, lets users select their hardware to find compatible open AI models with speed estimates and cloud-pricing comparisons, aiming to enable local AI model usage independent of …