We found defects in 37 of DeepSWE's 113 tasks
An audit of DeepSWE v1.1 found defects or ambiguous requirements in 37 of the benchmark's 113 tasks (32.7%), based on a review of all 372 recorded failures for the Opus 5, Sol and Fable 5 models. The …
An audit of DeepSWE v1.1 found defects or ambiguous requirements in 37 of the benchmark's 113 tasks (32.7%), based on a review of all 372 recorded failures for the Opus 5, Sol and Fable 5 models. The …
Google released three Gemini models on the same day: Gemini 3.6 Flash, Gemini 3.5 Flash Light, and Gemini 3.5 Cyber, prioritizing token efficiency over raw benchmark chasing. Gemini 3.6 Flash cuts out…
DeepSWE v1.1 updates the benchmark for long-horizon engineering tasks with isolated verification and structured test reports, making results more reproducible and harder to game. Pass rates remain clo…
DataCurve released DeepSWE, a new benchmark for evaluating frontier coding agents on original, long-horizon software engineering tasks. The benchmark features contamination-free tasks written from scr…