11:36
2026-10-02
mostlyright.md
large-language-models
A table of 1,200 model benchmarks since 2019, updated daily
Epoch AI's Benchmarking Hub held about 6,800 results covering roughly 1,200 model versions of 710 models as of 1 October 2026, spanning more than 80 benchmarks from GPQA Diamond and SWE-bench Verifiedβ¦