We found defects in 37 of DeepSWE's 113 tasks
An audit of DeepSWE v1.1 found defects or ambiguous requirements in 37 of the benchmark's 113 tasks (32.7%), based on a review of all 372 recorded failures for the Opus 5, Sol and Fable 5 models. The defects included hid…