Definition
Plain language
A benchmark of math word problems where you can dial up how many reasoning steps are required.
As stated in the literature
A procedurally generated grade-school math benchmark with controllable arithmetic depth, used to test how reasoning quality scales with sequential computation in long-context and hybrid models.
Why it matters: Controllable difficulty lets researchers separate raw math knowledge from the ability to chain many steps together, which is a key axis of reasoning capability.
For example, GSM-Infinite can generate the same kind of problem with 5 reasoning steps or 50, letting researchers see exactly where a model's accuracy starts to crack.
Heard on the show
“And then they go to GSM-Infinite, which is procedurally generated math word problems with controllable arithmetic depth.”Episode 085 — Why Long-Context Models Might Need Compute, Not Capacity, Before Eviction