A table of 1,200 model benchmarks since 2019, updated daily Epoch AI's Benchmarking Hub held about 6,800 results covering roughly 1,200 model versions of 710 models as of 1 October 2026, spanning more than 80 benchmarks from GPQA Diamond and SWE-bench Verified to FrontierMath and Humanity's Last Exam. The table is rebuilt daily from Epoch AI's download and covers models released from October 2019 onward, with one row per published result so a single model can appear multiple times on one benchmark when shots, harness, reasoning effort or source differ. Epoch AI changed 36 of its benchmark tables between 29 September and 1 October 2026, and about 70 published rows were excluded for lacking a numeric score or model name. LLM benchmark scores by model, since 2019 Compare large language models on more than 80 benchmarks, from GPQA Diamond and SWE-bench Verified to FrontierMath and Humanity's Last Exam. The table gathers every result published in Epoch AI's Benchmarking Hub, with each score in its own unit, who ran the evaluation and a link to the source. Use it to rank models, follow how scores rose with each release, or set lab and paper numbers beside independent runs. On 1 October 2026 the hub held about 6,800 results for about 1,200 model versions of 710 models from OpenAI, Anthropic, Google DeepMind, Meta, Alibaba and about 60 other developers. Each row is one model version on one benchmark under one evaluation setup. Coverage - Window: models released from October 2019 to Epoch AI's latest update, most of them in 2025 and 2026. Dates are UTC calendar dates. - Grain: one row per published result, so a model can have several rows on one benchmark when the shots, harness, reasoning effort or source differ. - Cadence: the table is rebuilt every day from Epoch AI's download. Epoch AI changed 36 of its benchmark tables between 29 September and 1 October 2026. - New benchmarks: a benchmark Epoch AI adds to its Capabilities Index appears automatically, with one best score per model family, until its full table is added to the build. Columns - result id : Epoch AI's record id for the result, or a hash of the row where Epoch AI publishes none. - benchmark , benchmark release date : the benchmark as Epoch AI names it, and the date it was published. - model , model version : the model family and the exact version evaluated, usually the developer's API name plus Epoch AI's reasoning-setting suffix. - organization , country , model access , model release date : the developer, its country, how the model is released API or open weights and when. - training compute flop : Epoch AI's estimate of training compute, in floating-point operations. - score , score unit , higher is better : the score as published, its unit and its direction. Units are fraction, percent, points out of 10, Elo rating, rating, speedup factor, Brier score, index points or US dollars. - score fraction : the score as a fraction from 0 to 1, for benchmarks scored as a proportion. - stderr : the standard error the source reports, in the unit of score . - result date : the run, grading, addition or update date, where the source gives one. - reported by : Epoch AI for its own runs, Leaderboard, Model developer, Paper or report, or Not stated. - source , source url : the cited paper, leaderboard or site and its link, or Epoch AI's evaluation log for its own runs. - evaluation setup : the shots, harness, agent, reasoning effort or provider that tell several results for one model apart. - notes , epoch file : Epoch AI's note on the result and the file in its download the row comes from. Missing values An empty cell means the source did not publish that value. Rows without a numeric score or a model name are left out, about 70 of the published rows on 1 October 2026. Results for older models often lack an organisation, access type or release date. Most results carry no standard error or result date. Epoch AI does not flag lab self-reports, so labs' own numbers appear under Paper or report and Model developer, next to academic papers. Suitable for - Ranking models on one benchmark - Tracking how the best scores changed with model release dates - Comparing Epoch AI's independent runs with numbers from leaderboards, papers and labs - Relating benchmark scores to training compute and model access Source and rights Every row comes from the Epoch AI Benchmarking Hub at epoch.ai/benchmarks, read daily from its download. Epoch AI publishes its own evaluation runs under the Creative Commons Attribution 4.0 licence. Results Epoch AI compiled from other projects keep those projects' licences, so credit Epoch AI and the source linked in each row. Tables | Name | Rows est. | Updated | Get the data | |---|---|---|---| | Table overviewPublished rows6,843Columns22 rows22 cols Update detailsStatusLiveLast published2 Oct 2026 | Table overviewPublished rows6,843Columns22 | Update detailsStatusLiveLast published2 Oct 2026 | | Sources 1 publisher Epoch AIEpoch AI is a research institute investigating key trends and questions that will shape the trajectory and governance of Artificial Intelligence.2 endpoints - About - Epoch AI is a research institute investigating key trends and questions that will shape the trajectory and governance of Artificial Intelligence. - Website - epoch.ai ↗ https://epoch.ai/ - Usage rights - Licensed by contract. - Requests - 85 requests across 2 endpoints Details - Contents - 1 table · 6,843 rows est. · 22 columns - Updated - 2 October 2026 - Published - 2 October 2026 - Version - v1 - License - Not stated - Visibility - Public - Publisher - Mostly Right https://mostlyright.md/publishers/mostlyrightmd - Topics - llm benchmarks · large language models · ai evaluation +4 Activity 339 +339 in the last 30 days 0 +0 in the last 30 days 10 +10 in the last 30 days