09:31
2026-09-11
danluu.com
ai-research
Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
A blog post examines three categories of benchmarks โ performance "napkin math" estimates, AI model evals including DeepSWE and Senior SWE-Bench, and winter tire comparisons โ arguing each is misused โฆ