Definition
Plain language
DeepSeek's family of math-focused language models, including a version that grades full mathematical proofs.
As stated in the literature
DeepSeek's math-specialized model series, including DeepSeekMath-V2, used as a proof-grading judge in olympiad-level RL training.
Also called: DeepSeekMath-V2
Why it matters: Reliable proof graders are the bottleneck for training models on olympiad-level math, since otherwise the RL signal is too noisy to learn from.
For example, DeepSeekMath-V2 reads a multi-step proof and outputs a structured grade that an RL trainer uses as the reward signal.
Heard on the show
“The other variant gets reinforcement learning — GRPO, the algorithm from the DeepSeekMath paper — and the only signal it gets is binary correct or incorrect on the same two-hundred problems.”Episode 011 — When RL Actually Teaches Agents Something New, And When It Doesn't