{"slug": "show-hn-benchmark-ai-doesn-t-find-bugs-unless-you-tell-it-what-s-wrong", "title": "Show HN: Benchmark: AI doesn't find bugs unless you tell it what's wrong", "summary": "Researchers at Meta, Stanford, Harvard and the University of Washington released SWE-Sweep, an MIT-licensed benchmark of 4,000 real-world GitHub bugs across 100 repositories in 22 languages, finding that the best model configuration tested fixed fewer than 5% of bugs when not told what was wrong. The top setup, Sol 5.6 (xhigh), resolved 4.7% of bugs at a cost of $7,230, while Luna 5.6 (high) resolved 1.4% for $28, and Gemini 3.5 Flash Lite resolved 0.1% for $6. The benchmark is open-sourced at github.com/facebookresearch/swe-sweep.", "body_md": "New benchmark from researchers at Meta, Stanford, Harvard, UW, including the researchers who've worked on SWE-bench, ProgramBench etc.\n\nMost benchmarks just test if AI can fix a problem you've already pointed out.\n\nBut obviously it would be much better to fix problems before you or any user runs into it. Like, isn't it crazy that we still have to wait for people to open tickets before a lot of obvious bugs get found?\n\nWe wanted to test that capability at scale. Turns out that models still are terrible at it (best setup we tested still fixed < 5% of bugs)\n\nWe have 100 repos of 22 languages and 4k bugs between. All the bugs are real-world bugs from github. We do a lot of filtering to ensure everything can be solved in this setting.\n\nModel Resolve Cost Sol 5.6 (xhigh) 4.7% $7,230 Luna 5.6 (xhigh) 2.5% $224 Terra 5.6 (xhigh) 1.5% $357 Luna 5.6 (high) 1.4% $28 Opus 5 (xhigh) 1.3% $5,363 Kimi K3 0.6% $2,451 Luna 5.6 0.5% $4 GPT-5.4 Mini (high) 0.5% $122 GPT-5.4 Mini 0.2% $5 Gemini 3.5 Flash Lite 0.1% $6\n\nAlso the best model is very expensive.\n\nWe have a lot more FAQ on the website [https://swesweep.com/](https://swesweep.com/)\nOh and we're all open-source (MIT license) at [https://github.com/facebookresearch/swe-sweep](https://github.com/facebookresearch/swe-sweep)\n\nCurious what you all think!\n\nComments URL: [https://news.ycombinator.com/item?id=49923102](https://news.ycombinator.com/item?id=49923102)\n\nPoints: 1\n\n# Comments: 0", "url": "https://wpnews.pro/news/show-hn-benchmark-ai-doesn-t-find-bugs-unless-you-tell-it-what-s-wrong", "canonical_source": "https://swesweep.com/", "published_at": "2026-10-01 15:36:14+00:00", "updated_at": "2026-10-01 15:48:33.172496+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "large-language-models", "ai-agents", "developer-tools"], "entities": ["Meta", "Stanford University", "Harvard University", "University of Washington", "SWE-Sweep", "SWE-bench", "Sol 5.6", "Luna 5.6"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/show-hn-benchmark-ai-doesn-t-find-bugs-unless-you-tell-it-what-s-wrong", "markdown": "https://wpnews.pro/news/show-hn-benchmark-ai-doesn-t-find-bugs-unless-you-tell-it-what-s-wrong.md", "text": "https://wpnews.pro/news/show-hn-benchmark-ai-doesn-t-find-bugs-unless-you-tell-it-what-s-wrong.txt", "jsonld": "https://wpnews.pro/news/show-hn-benchmark-ai-doesn-t-find-bugs-unless-you-tell-it-what-s-wrong.jsonld"}}