Artificial Analysis Endpoint Accuracy Index
Artificial Analysis launched its Endpoint Accuracy Index, measuring how much of an open weights model's accuracy each serverless API endpoint preserves, with initial coverage of GLM-5.2, gpt-oss-120b,…
Artificial Analysis launched its Endpoint Accuracy Index, measuring how much of an open weights model's accuracy each serverless API endpoint preserves, with initial coverage of GLM-5.2, gpt-oss-120b,…
DeepSeek's V4 Flash 0731 build corrupts integer fields in strict JSON schema outputs when thinking mode is enabled, failing 8 of 13 runs across two request paths, while disabling thinking fixes all ru…
Thinking Machines Lab released Inkling on July 15, a 975-billion-parameter Mixture-of-Experts model with 41B active parameters, open-weight under Apache 2.0, from Mira Murati's $12B startup. The model…
Researchers introduced SWE-Touch, a benchmark evaluating coding agents' ability to repair software after users directly edit code, finding that most models' performance drops significantly when user e…
DeepSeek V4 Pro, an open-source coding model, scored 80.6% on SWE-bench Verified, matching top closed-source models, but requires a server rack. The best open-source model that runs on a single 24GB G…
A test of five AI models — GPT-5.6 Sol, DeepSeek V4 Pro, Claude Fable 5, Grok-4.5, and Gemini 3.1 Pro — found that only two showed clear self-preference bias when scoring their own unlabeled essays al…
DeepSeek V4 Pro, a Mixture-of-Experts model with 1.6 trillion total parameters, scores 80.6% on SWE-bench Verified, tying Gemini 3.1 Pro and achieving the highest score for any open-weight model. Pric…
DeepSeek's April 2026 release introduced two distinct V4 models — V4 Pro and V4 Flash — with shared features including a 1M-token context window and 384K maximum output, but differing economics and ar…
A new analysis predicts that AI models will grow from 10 trillion parameters in 2026 to 1.4 quadrillion parameters by 2031, with inference costs remaining surprisingly low due to KV cache scaling. The…
Anthropic published a guide on building verification loops in Claude Code with skills, enabling AI agents to automatically check and fix their own work using tests, linters, and custom checks. The app…
Developer Nokka reviewed SCOTOMA, an abliterated version of Google's Gemma 4 31B Instruct model created by ReadyArt. SCOTOMA uses a novel J-space abliteration technique with Jacobian-lens projection t…
Claude Opus 5 (Adaptive Reasoning, Max Effort) leads the Artificial Analysis Intelligence Index with a score of 61, out of 170 models evaluated, according to Artificial Analysis. The top five models a…
A developer shared a configuration guide for integrating Myt AI Console's custom AI models into OpenCode. The setup involves appending an opencode-config.json file to the global OpenCode profile and a…
DeepSeek retired the deepseek-chat and deepseek-reasoner API aliases on July 24 at 15:59 UTC, shipping V4 Stable alongside, with migration traps including V4 Flash having thinking mode ON by default a…
DigitalOcean's new server-side model synthesis tool on its Inference Engine outperformed Fable 5 at half the cost, scoring 65.65% on the DRACO benchmark at $0.83 per task versus Fable 5's 62.21% at $1…
Anthropic researcher Levent Alpöge tweeted a counterargument to the 80-year-old Jacobian Conjecture, identified using Claude Fable 5 and empirically validated. When fed to LLMs, the counterargument cr…
Moonshot AI released Kimi K3 on July 16, a 2.8-trillion-parameter Mixture-of-Experts model that activates only ~50 billion parameters per token, making it the largest open-weight model ever built. K3 …
PenguinHarness, an open-source agent framework from Prism Shadow, claims to build agents at 100× the speed of LangChain with a zero-code CLI and Web UI connected to 1000+ models, achieving best accura…
Simon Willison's informal benchmark asking AI models to generate an SVG of a pelican riding a bicycle has become a widely discussed test for large language models. A new experiment tested 1,008 SVGs a…
LM Studio released Bionic on July 16, a new agent application for Mac and Windows that marks the company's first metered product with a checkout page, pivoting from its free local LLM desktop app to a…