cd /news/artificial-intelligence/introducing-hle-diamond-humanity-s-l… · home topics artificial-intelligence article
[ARTICLE · art-138445] src=lastexam.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Introducing HLE-Diamond – Humanity's Last Exam

The Center for AI Safety released HLE-Diamond, a refined 1,000-question subset of the Humanity's Last Exam (HLE) benchmark, following a year of cleaning and refinement with input from research communities. HLE-Diamond splits into 500 reasoning questions and 500 knowledge questions, and the release compares current models' performance on it without tools, with all models evaluated at reasoning effort set to high. The dataset is designed to be answerable closed-book, with recommended settings published for evaluating it with tools given HLE's use in testing agentic systems.

by read1 min views1 publishedSep 23, 2026
Introducing HLE-Diamond – Humanity's Last Exam
Image: source

We are releasing HLE-Diamond, a refined subset of Humanity’s Last Exam (HLE) question collection, following a year-long process of cleaning and refinement with input from research communities.

HLE-Diamond consists of 1,000 questions.

Main results. We compare the performance of current models on HLE-Diamond without tools.

All models are evaluated with reasoning high.

Dataset. HLE-Diamond consists of 500 reasoning and 500 knowledge questions, measuring reasoning and expert knowledge, respectively.

Evaluation with tools. HLE-Diamond questions are designed to be answerable in a closed-book setting, testing both reasoning and expert knowledge. Since HLE is also used to evaluate the capabilities of agentic systems, we outline our recommended settings for evaluating HLE-Diamond with tools here.

Acknowledgement. We extend our deepest gratitude to all participating question contributors and experts involved in creating and refining the dataset, and to the researchers whose inputs across HLE-Rolling shaped HLE-Diamond.

For any inquiries, please contact agibenchmark@safe.ai.

Citation #

@article{phan2025lastexam,
      title = {A benchmark of expert-level academic questions to assess {AI} capabilities},
      author = {{Center for AI Safety} and {Scale AI} and {HLE Contributors Consortium}},
      journal = {Nature},
      volume = {649},
      pages = {1139--1146},
      year = {2026},
      doi = {10.1038/s41586-025-09962-4},
      eprint = {2501.14249},
      archivePrefix = {arXiv},
      primaryClass = {cs.LG},
      url = {https://arxiv.org/abs/2501.14249}
}
── more in #artificial-intelligence 4 stories · sorted by recency
── more on @center for ai safety 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/introducing-hle-diam…] indexed:0 read:1min 2026-09-23 ·