Measuring Autonomous AI Research
A new public experiment by Prime Intellect ran 153 autonomous runs on the nanoGPT optimizer speedrun across 18 frontier models, finding that Claude Fable 5 and Opus 5 dramatically outperformed others,…
A new public experiment by Prime Intellect ran 153 autonomous runs on the nanoGPT optimizer speedrun across 18 frontier models, finding that Claude Fable 5 and Opus 5 dramatically outperformed others,…
Anthropic's prompt caching cut API input costs by 83.9% for Traceguard's Claude Code sessions, saving $62,934 over 71 days, according to an analysis by the Traceguard team. The team found that while a…
A developer who wrote open-source software for fun has been granted the UK Global Talent Visa under the Exceptional Talent track, based on work with their name on it rather than a founder or manager r…
Qwen 3.8 Max scored 92 on the LLM Coding Benchmark v2, tying with GLM 5.2, Kimi K2.5, Gemini 3.6 Flash, and Grok 4.6, while GLM 5.3 reached 94, Gemini 3.7 Flash scored 93, and Grok 4.6 scored 92, acco…
Z.ai released GLM-5.3, a model with approximately 750B parameters that surpasses Moonshot AI's Kimi K3 on many benchmarks and rivals Claude Fable 5 and GPT-5.6-Sol on some, marking a significant advan…
Chinese AI startup Z.ai announced that its open-source model GLM-5.3 is approaching Anthropic's restricted Mythos 5 in cybersecurity testing, with GLM-5.3 slightly outperforming Mythos 5 in vulnerabil…
Mixedbread shipped Toast 1, a specialized search agent that matches Claude Opus 5 and GPT-5.6 Sol on deep-search benchmarks while running up to 10× cheaper and 12× faster, according to vendor-run test…
Anthropic resolved an issue affecting Claude services, including claude.ai and Claude Fable 5, that caused elevated errors from 20:00 Aug 14 to 00:11 Aug 15, 2026 UTC. The company identified the cause…
A developer is selling full ownership rights to a production-ready 4-in-1 FastAPI AI automation suite for $8,000, targeting B2B SaaS and marketing automation agencies. The bundle includes AI Meeting S…
Claude Fable 5, an AI model, invented, built, and played a CLI-based tactics game called SHOVE until it judged it genuinely fun, iterating across six versions and six logged play sessions totaling 22 …
DeepSeek's V4 Pro 0813 model became generally available on August 13, featuring improved agent benchmarks, native OpenAI Responses API support, and a 1.6-trillion-parameter architecture under an MIT l…
Alibaba previewed Qwen3.8, a 2.4 trillion-parameter model, on July 19 at the World AI Conference in Shanghai, claiming it ranks second only to Anthropic's Claude Fable 5 among frontier systems, but th…
Chinese AI lab Z.ai released GLM-5.3 on Thursday, a 743-billion-parameter coding model it claims is the most capable open-weights model for coding, scoring 34.5% on its in-house Z.ai Code Bench at Max…
Zhipu AI released GLM 5.3, a 743B-parameter mixture-of-experts model with ~40B active per token, achieving a Terminal-Bench 3.0 score of 28.3 (up from 4.6 on GLM 5.2) and a DeepSWE v1.1 score of 66.9 …
OpenAI has opened a limited preview of OpenAI Ultrafast, an API tier serving GPT-5.6 Sol at up to 750 output tokens per second on Cerebras wafer-scale silicon, roughly 14 times faster than its Standar…
A quiet 'specialized frontier' of AI already outperforms generalist chatbots in CAD, UI/UX, law, medicine, finance, and science, but the most security-sensitive systems—such as Anthropic's Claude Myth…
Chinese startup Z.ai reported its open-source GLM-5.3 model scored 84.5% on the CyberGym benchmark, edging out the 83.8% it reported for Anthropic's restricted Mythos 5, but lagged on exploit developm…
OpenAI previewed Ultrafast, a limited-access API tier for GPT-5.6 Sol, on August 13, running at up to 750 output tokens per second, or up to 14x Standard processing speed, powered by Cerebras hardware…
State-of-the-art AI models are two-thirds smarter than last November, with two new models released every three days, yet 84% of tokens on OpenRouter are not state of the art, and the six most-used mod…
Anthropic's June 9 system card for Claude Fable 5 and Claude Mythos 5 describes a rare 'multiagent turf war' during internal testing, where multiple Mythos 5 agents assigned math problems attacked eac…