cd/entity/SWE-bench· home entities SWE-bench
grep -l @swe-bench /news/*.json | wc -l → 82

SWE-bench

mentions 82 type Organization page 3/5 feed RSS

// recent coverage 82 mentions

05:53
2026-07-10
github.com
artificial-intelligence

TinyToT – Tree of Thoughts Inference Server

TinyToT, a lightweight inference server compatible with Ollama, achieves 97% accuracy on a 35-question benchmark spanning graduate-level science, medicine, law, finance, and software engineering witho…

00:27
2026-07-10
lesswrong.com
ai-safety

Toward A Public Science of Model Behavior

AI systems increasingly exhibit unexpected and dangerous behaviors, such as Replit's coding agent deleting a startup's production database and ChatGPT allegedly contributing to a user's suicide. To en…

06:27
2026-07-09
letsdatascience.com
artificial-intelligence

AI Benchmark Scores Overstate Model Performance

A PlainEnglish article warns that AI benchmark scores such as MMLU, HumanEval, and HellaSwag can overstate production readiness when leaderboard numbers are treated as proof of model quality. The comm…

22:03
2026-07-07
sourcefeed.dev
large-language-models

Claude Opus 4.7: Engineering the Agentic Loop

Anthropic released Claude Opus 4.7 on April 16, 2026, a targeted upgrade focused on making agentic coding loops and multi-step tool use production-ready. The model achieves 87.6% on SWE-bench Verified…

14:09
2026-07-07
byteiota.com
large-language-models

Claude Sonnet 5 Migration: Three Breaking API Changes

Claude Sonnet 5 launched June 30 as the default model across Claude Code, Free, and Pro plans, but developers face three breaking API changes: sampling parameters (temperature, top_p, top_k) now retur…

13:06
2026-07-02
sourcefeed.dev
artificial-intelligence

Beyond Bug Fixing: The Rise of Senior-Level AI Coding Benchmarks

New AI coding benchmarks, including Snorkel AI's Senior SWE-Bench and Scale AI's SWE-Bench Pro, are replacing older evaluations like HumanEval and original SWE-bench to test senior-level engineering s…

14:31
2026-07-01
devblogs.microsoft.com
artificial-intelligence

What AI benchmarks are not telling you

Public AI benchmarks like SWE-bench measure performance on popular open-source repositories but fail to predict how models will perform on proprietary codebases, team-specific conventions, and real-wo…

15:17
2026-06-30
byteiota.com
large-language-models

Gemini 2.5 Pro Deep Think: What the Benchmarks Mean

Google's Gemini 2.5 Pro with Deep Think reasoning mode topped coding and reasoning benchmarks this week, scoring 82.4% on GPQA Diamond and 94.1% on HumanEval+, but the mode multiplies token costs by r…

17:16
2026-06-29
dev.to
large-language-models

RAG for codebases is hard. Trusting the answer is harder.

A developer argues that retrieval-augmented generation (RAG) for codebases improves context but not verifiability, citing that 30% of failed SWE-agent runs still claimed success. They introduce 'truth…

23:20
2026-06-25
letsdatascience.com
ai-tools

GitHub evaluates Copilot agentic harness performance

GitHub announced two harness-level improvements to Copilot agentic sessions—prompt caching achieving 94% cache hit rates and deferred tool loading—plus a new Auto model selection feature using its HyD…

12:30
2026-06-24
andrewjesson.com
ai-agents

When Does Data Help Automated Context Engineering?

Claude Code can improve other AI agents without training data in four of seven tested applications, performing as well as with data. Data helps only where Claude Code's prior knowledge of the task run…

00:00
2026-06-22
epics.tech
ai-policy

The Outcry Grows, but the Capital Keeps Flowing

Institutional backlash against AI is mounting across public sector, media, and education, with New York City council members urging a classroom AI pause and a German newsroom caught using AI to write …

← prev page 3 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics