# Introducing HLE-Diamond – Humanity's Last Exam

> Source: <https://lastexam.ai/blog/hle-diamond>
> Published: 2026-09-23 18:06:20+00:00

We are releasing **HLE-Diamond**, a refined subset of *[Humanity’s Last Exam (HLE)](https://lastexam.ai)* question collection, following a year-long process of cleaning and refinement with input from research communities.

HLE-Diamond consists of **1,000 questions**.

**Main results.** We compare the performance of current models on HLE-Diamond without tools.

All models are evaluated with reasoning *high*.

**Dataset.** HLE-Diamond consists of **500 reasoning** and **500 knowledge** questions, measuring reasoning and expert knowledge, respectively.

**Evaluation with tools.** HLE-Diamond questions are designed to be answerable in a closed-book setting, testing both reasoning and expert knowledge. Since HLE is also used to evaluate the capabilities of agentic systems, we outline our recommended settings for evaluating HLE-Diamond with tools [here](https://github.com/centerforaisafety/hle/blob/main/docs/evaluation-with-tools.md).

**Acknowledgement.** We extend our deepest gratitude to all participating question contributors and experts involved in creating and refining the dataset, and to the researchers whose inputs across HLE-Rolling shaped HLE-Diamond.

For any inquiries, please contact [agibenchmark@safe.ai](mailto:agibenchmark@safe.ai).

## Citation

```
@article{phan2025lastexam,
      title = {A benchmark of expert-level academic questions to assess {AI} capabilities},
      author = {{Center for AI Safety} and {Scale AI} and {HLE Contributors Consortium}},
      journal = {Nature},
      volume = {649},
      pages = {1139--1146},
      year = {2026},
      doi = {10.1038/s41586-025-09962-4},
      eprint = {2501.14249},
      archivePrefix = {arXiv},
      primaryClass = {cs.LG},
      url = {https://arxiv.org/abs/2501.14249}
}
```


