15:29
2026-07-14
forum.effectivealtruism.org
ai-safety
Good Benchmarks
METR contributor Ivan Bercovich argues that most AI benchmarks are flawed and that building good ones requires nuanced understanding, drawing on 18 months of experience with Terminal Bench. Good tasksβ¦