Definition
Plain language
A benchmark of research-grade math problems used to test AI math reasoning.
As stated in the literature
An Epoch AI benchmark of advanced mathematics problems organized into difficulty tiers (Tier 4 being short-term research projects for PhD mathematicians), used to evaluate AI mathematical reasoning capabilities.
Also called: FrontierScience-Research
Why it matters: It pushes math benchmarks past competition-style problems into territory where solving anything is a meaningful capability signal.
For example, a Tier 4 FrontierMath problem might be the kind of question a math PhD student would spend a week or two working on as a self-contained research mini-project.
Heard on the show
“They can solve problems on a benchmark called FrontierMath that human PhDs struggle with for hours.”Episode 029 — Why Forty-Eight Percent on FrontierMath Isn't the Real Story in DeepMind's New Math Paper