Glossary · Term

Humanity's Last Exam

← all terms

Definition

Plain language

A very hard benchmark of expert-level questions across many fields used to push top AI systems.

As stated in the literature

A frontier evaluation suite of expert-level multi-domain questions used to probe ceiling performance of capable models.

Also called: HLE

Why it matters: It gives researchers a way to keep measuring progress once frontier models already score near-perfectly on easier benchmarks.

For example, a question might ask for a derivation in graduate-level quantum field theory or an obscure historical translation that only a domain specialist would normally attempt.

Heard on the show

“On Humanity's Last Exam, which is a brutal multi-domain expert question set, thirty-four point six versus thirty-two point nine.”
Episode 021 — Ten Thousand Examples Beat the Full Industrial Pipeline for Search Agents

Related terms