AI Benchmarks: When Is Enough Truly Enough?
A study examining partial evaluations on AI agent benchmarks including SWE-bench, AppWorld, and tau-bench finds that partial budgets are only valid when they replicate the full benchmark's final decis…
A study examining partial evaluations on AI agent benchmarks including SWE-bench, AppWorld, and tau-bench finds that partial budgets are only valid when they replicate the full benchmark's final decis…
A replay analysis of public LLM agent benchmarks SWE-bench, AppWorld, and tau-bench finds that the fraction of tasks needed to reach the same pairwise conclusion as the full benchmark varies sharply, …
A new study from arXiv:2607.12161v1 analyzing 2,848 Claude Code runs across 103 tasks finds that reducing retrieved context or tool output does not reliably lower billed costs for API-based coding age…
Mistral Vibe for Code, Claude Code, Cursor, and OpenAI Codex were scored on a scaffold-to-PR task, with Mistral Vibe and Claude Code tying at 22/25 points. The comparison evaluated five dimensions: fe…
A new open-source harness from Tensorlake demonstrates that agent evaluations are vulnerable to cheating unless per-task isolation is enforced, showing a 53% lie rate when agents can tamper with test …
A 5-condition ablation study (3,175 total runs) validates that growth-ratio normalized Lyapunov energy functions serve as precise leading indicators of agent task failure, achieving zero stability vio…
An engineering study testing whether a weak large language model wrapped in strong scaffolding (retrieval, toolset, verification loops) can match frontier-level performance finds that the harness clos…
Cognition Labs launched Devin, the 'first AI software engineer,' in March 2024 with a benchmark score of 13.86% on SWE-bench, but independent evaluations later showed it failed 14 out of 20 real tasks…
TinyToT, a lightweight inference server compatible with Ollama, achieves 97% accuracy on a 35-question benchmark spanning graduate-level science, medicine, law, finance, and software engineering witho…
AI systems increasingly exhibit unexpected and dangerous behaviors, such as Replit's coding agent deleting a startup's production database and ChatGPT allegedly contributing to a user's suicide. To en…
A PlainEnglish article warns that AI benchmark scores such as MMLU, HumanEval, and HellaSwag can overstate production readiness when leaderboard numbers are treated as proof of model quality. The comm…
Atelier, a 30-second install tool that sits underneath Claude Code to reduce token waste, claims 30% savings on AI coding costs by providing better search, shorter file reads, compact command output, …
Anthropic released Claude Opus 4.7 on April 16, 2026, a targeted upgrade focused on making agentic coding loops and multi-step tool use production-ready. The model achieves 87.6% on SWE-bench Verified…
Z.ai's open-weight GLM 5.2 completed an agentic coding task in 17 minutes at $2.76, while Anthropic's Claude Fable 5 finished in 9 minutes at over $10, with comparable output quality. The open model c…
Anthropic released Claude Opus 4.7 on April 16, 2026, as its most capable public model, featuring a 13% improvement on coding benchmarks, a new xhigh effort level for agentic tasks, and high-resolutio…
Claude Sonnet 5 launched June 30 as the default model across Claude Code, Free, and Pro plans, but developers face three breaking API changes: sampling parameters (temperature, top_p, top_k) now retur…
New AI coding benchmarks, including Snorkel AI's Senior SWE-Bench and Scale AI's SWE-Bench Pro, are replacing older evaluations like HumanEval and original SWE-bench to test senior-level engineering s…
Public AI benchmarks like SWE-bench measure performance on popular open-source repositories but fail to predict how models will perform on proprietary codebases, team-specific conventions, and real-wo…
Google's Gemini 2.5 Pro with Deep Think reasoning mode topped coding and reasoning benchmarks this week, scoring 82.4% on GPQA Diamond and 94.1% on HumanEval+, but the mode multiplies token costs by r…
A developer argues that retrieval-augmented generation (RAG) for codebases improves context but not verifiability, citing that 30% of failed SWE-agent runs still claimed success. They introduce 'truth…