We are releasing HLE-Diamond, a refined subset of Humanity’s Last Exam (HLE) question collection, following a year-long process of cleaning and refinement with input from research communities.
HLE-Diamond consists of 1,000 questions.
Main results. We compare the performance of current models on HLE-Diamond without tools.
All models are evaluated with reasoning high.
Dataset. HLE-Diamond consists of 500 reasoning and 500 knowledge questions, measuring reasoning and expert knowledge, respectively.
Evaluation with tools. HLE-Diamond questions are designed to be answerable in a closed-book setting, testing both reasoning and expert knowledge. Since HLE is also used to evaluate the capabilities of agentic systems, we outline our recommended settings for evaluating HLE-Diamond with tools here.
Acknowledgement. We extend our deepest gratitude to all participating question contributors and experts involved in creating and refining the dataset, and to the researchers whose inputs across HLE-Rolling shaped HLE-Diamond.
For any inquiries, please contact agibenchmark@safe.ai.
Citation #
@article{phan2025lastexam,
title = {A benchmark of expert-level academic questions to assess {AI} capabilities},
author = {{Center for AI Safety} and {Scale AI} and {HLE Contributors Consortium}},
journal = {Nature},
volume = {649},
pages = {1139--1146},
year = {2026},
doi = {10.1038/s41586-025-09962-4},
eprint = {2501.14249},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2501.14249}
}