16:39
2026-09-22
horizonanalyticslabs.com
ai-research
Benchmarks are more broken than we could have imagined
An audit by Horizon of 20 public task datasets in the Harbor hub found 29 confirmed broken tasks out of 5,241 scanned, with failures that often made models look better rather than worse, according to …