Definition
Plain language
A very hard benchmark of expert-level questions across many fields used to push top AI systems.
As stated in the literature
A frontier evaluation suite of expert-level multi-domain questions used to probe ceiling performance of capable models.
Also called: HLE
Why it matters: It gives researchers a way to keep measuring progress once frontier models already score near-perfectly on easier benchmarks.
For example, a question might ask for a derivation in graduate-level quantum field theory or an obscure historical translation that only a domain specialist would normally attempt.
Heard on the show
“M-M-L-U, G-P-Q-A, Humanity's Last Exam, OpenAI's FrontierScience — every one of them is answer-centric.”Episode 240 — Frontier Models Designed Follow-Ups To Fraudulent Papers 93% Of The Time