cd/entity/Humanity's Last Exam· home entities Humanity's Last Exam
grep -l @humanity's last exam /news/*.json | wc -l → 16

Humanity's Last Exam

mentions 16 type Person feed RSS

// recent coverage 16 mentions

14:24
2026-08-16
gist.github.com
artificial-intelligence

Show HN: Reasoning prefills on a few open models

A developer's independent analysis of reasoning prefills on open models finds that Kimi K3 shows a significant accuracy increase on Humanity's Last Exam questions when prefilled with Opus 4.8 reasonin…

18:10
2026-08-13
cerebras.ai
artificial-intelligence

Accelerating GPT-5.6 Sol Ultrafast

Cerebras and OpenAI launched Ultrafast Mode, a new service tier in the OpenAI API powered by Cerebras that delivers GPT-5.6 Sol at up to 750 output tokens per second with no quality compromise. In Cer…

16:07
2026-08-04
abhishek-shankar.com
artificial-intelligence

The AI Math Boom Is a Trust Infrastructure Boom in Disguise

Three startups building AI mathematicians—Harmonic, Axiom Math, and Math Inc—raised over $580 million in the last twelve months, with Harmonic at a $1.45 billion valuation and Axiom Math jumping from …

09:32
2026-08-01
testingcatalog.com
artificial-intelligence

Thinking Machines launched open-weight Inkling-Small

Thinking Machines Lab released Inkling-Small, an open-weight Mixture-of-Experts transformer with 276 billion total parameters and 12 billion active, designed to deliver performance comparable to its l…

05:34
2026-07-25
news.ycombinator.com
artificial-intelligence

Companies are optimizing models for specific benchmarks

OpenAI is optimizing its GPT-5.6 model for the GPQA Diamond benchmark, while Anthropic is optimizing Opus 5 for the Humanity's Last Exam benchmark, with GPT-5.6 winning on GPQA Diamond and Opus 5 winn…

04:00
2026-07-23
machinebrief.com
artificial-intelligence

PoTRE: Test-Time Reasoning inspired by Cognitive Heterogeneity

Researchers introduced PoTRE (Poly-Topological Reasoning Ensembles), a heterogeneous framework that decouples inference into four agents to improve complex reasoning in large language models. PoTRE ac…

00:34
2026-07-23
poetiq.ai
artificial-intelligence

Benchmarks Are Dead (For Us)

Poetiq claims its Recursive Self-Improvement (RSI) Metasystem has autonomously set state-of-the-art results on six diverse benchmarks, including outperforming Muse Spark 1.1 within 48 hours of publica…

20:47
2026-07-18
humans-vs-hle.lizzie-siegle5086.workers.dev
artificial-intelligence

Humans vs. LLMs: Can You Beat Humanity's Last Exam?

A new benchmark called Humanity's Last Exam, designed to stump frontier AI models, is now available as a 10-question randomized quiz called The Adult SAT. The quiz lets users compare their scores agai…

12:00
2026-07-02
kdnuggets.com
artificial-intelligence

Humanity’s Last Exam is a Distraction

The Center for AI Safety, with world experts, created Humanity's Last Exam (HLE), a benchmark of over 2,500 expert-level questions across disciplines to test AI reasoning. Even top models like GPT, Ge…

15:10
2026-06-18
byteiota.com
large-language-models

Gemini 3.5 Flash Is Now GA: Three API Traps to Know

Google has made Gemini 3.5 Flash generally available, offering faster performance and improved agentic and coding benchmarks over its predecessor, but developers migrating from the preview version mus…

// co-occurs with top 8 entities
// topics top 6 topics