cd/entity/SWE-bench· home› entities› SWE-bench
grep -l @swe-bench /news/*.json | wc -l → 106

SWE-bench

mentions 106 type Organization page 4/6 feed RSS

// recent coverage 106 mentions

07:39
2026-07-15
machinebrief.com
artificial-intelligence

AI Benchmarks: When Is Enough Truly Enough?

A study examining partial evaluations on AI agent benchmarks including SWE-bench, AppWorld, and tau-bench finds that partial budgets are only valid when they replicate the full benchmark's final decis…

04:00
2026-07-15
arxiv.org
artificial-intelligence

Token Reduction Is Not Cost Reduction

A new study from arXiv:2607.12161v1 analyzing 2,848 Claude Code runs across 103 tasks finds that reducing retrieved context or tool output does not reliably lower billed costs for API-based coding age…

05:53
2026-07-10
github.com
artificial-intelligence

TinyToT – Tree of Thoughts Inference Server

TinyToT, a lightweight inference server compatible with Ollama, achieves 97% accuracy on a 35-question benchmark spanning graduate-level science, medicine, law, finance, and software engineering witho…

00:27
2026-07-10
lesswrong.com
ai-safety

Toward A Public Science of Model Behavior

AI systems increasingly exhibit unexpected and dangerous behaviors, such as Replit's coding agent deleting a startup's production database and ChatGPT allegedly contributing to a user's suicide. To en…

06:27
2026-07-09
letsdatascience.com
artificial-intelligence

AI Benchmark Scores Overstate Model Performance

A PlainEnglish article warns that AI benchmark scores such as MMLU, HumanEval, and HellaSwag can overstate production readiness when leaderboard numbers are treated as proof of model quality. The comm…

22:03
2026-07-07
sourcefeed.dev
large-language-models

Claude Opus 4.7: Engineering the Agentic Loop

Anthropic released Claude Opus 4.7 on April 16, 2026, a targeted upgrade focused on making agentic coding loops and multi-step tool use production-ready. The model achieves 87.6% on SWE-bench Verified…

14:09
2026-07-07
byteiota.com
large-language-models

Claude Sonnet 5 Migration: Three Breaking API Changes

Claude Sonnet 5 launched June 30 as the default model across Claude Code, Free, and Pro plans, but developers face three breaking API changes: sampling parameters (temperature, top_p, top_k) now retur…

13:06
2026-07-02
sourcefeed.dev
artificial-intelligence

Beyond Bug Fixing: The Rise of Senior-Level AI Coding Benchmarks

New AI coding benchmarks, including Snorkel AI's Senior SWE-Bench and Scale AI's SWE-Bench Pro, are replacing older evaluations like HumanEval and original SWE-bench to test senior-level engineering s…

14:31
2026-07-01
devblogs.microsoft.com
artificial-intelligence

What AI benchmarks are not telling you

Public AI benchmarks like SWE-bench measure performance on popular open-source repositories but fail to predict how models will perform on proprietary codebases, team-specific conventions, and real-wo…

15:17
2026-06-30
byteiota.com
large-language-models

Gemini 2.5 Pro Deep Think: What the Benchmarks Mean

Google's Gemini 2.5 Pro with Deep Think reasoning mode topped coding and reasoning benchmarks this week, scoring 82.4% on GPQA Diamond and 94.1% on HumanEval+, but the mode multiplies token costs by r…

17:16
2026-06-29
dev.to
large-language-models

RAG for codebases is hard. Trusting the answer is harder.

A developer argues that retrieval-augmented generation (RAG) for codebases improves context but not verifiability, citing that 30% of failed SWE-agent runs still claimed success. They introduce 'truth…

← prev page 4 / 6 next →
// co-occurs with top 8 entities
// topics top 6 topics