Glossary · Term

Humanity's Last Exam

← all terms

Definition

Plain language

A very hard benchmark of expert-level questions across many fields used to push top AI systems.

As stated in the literature

A frontier evaluation suite of expert-level multi-domain questions used to probe ceiling performance of capable models.

Also called: HLE

Why it matters: It gives researchers a way to keep measuring progress once frontier models already score near-perfectly on easier benchmarks.

For example, a question might ask for a derivation in graduate-level quantum field theory or an obscure historical translation that only a domain specialist would normally attempt.

Heard on the show

“M-M-L-U, G-P-Q-A, Humanity's Last Exam, OpenAI's FrontierScience — every one of them is answer-centric.”
Episode 240 — Frontier Models Designed Follow-Ups To Fraudulent Papers 93% Of The Time

Mentioned in 7 episodes

  1. 240
    Frontier Models Designed Follow-Ups To Fraudulent Papers 93% Of The Time
  2. 166
    A Router That Beats the Frontier Models It Calls
  3. 131
    Why Autonomous Research Agents Forget Their Own Lessons, and Arbor's Fix
  4. 090
    How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents
  5. 083
    Training the Translator: How a Small Communication Model Lets Agent Teams Outperform Themselves
  6. 082
    Training a Deep Research Agent on 8,000 Synthetic Tasks: The Rubric Tree Trick
  7. 021
    Ten Thousand Examples Beat the Full Industrial Pipeline for Search Agents

Related terms