Glossary · Term

reasoning token

← all terms

Definition

Plain language

Extra text a model writes to itself while working out an answer before replying.

As stated in the literature

Intermediate chain-of-thought output emitted at test time; functionally lets the model iterate over candidates and apply a rule serially, recovering much of the authoring gap.

Also called: reasoning tokens

Why it matters: These intermediate steps let a model walk through candidates one at a time instead of guessing a whole set at once, which recovers much of the accuracy it otherwise loses.

For example, before answering "which of these ten numbers are prime," the model writes out a short check for each number in turn, then gives its final list.

Heard on the show

“And when they look at why, it's spending only about a hundred-twenty reasoning tokens regardless — it isn't really deliberating its way back to the policy, it's answering fast.”
Episode 118 — Why the Best-Aligned AI Models Are the Easiest to Trick Into Producing Harm

Related concepts

Related terms