Definition
Plain language
The raw numbers a model produces before they get turned into probabilities over words.
As stated in the literature
The unnormalized scores output by a model's final layer, converted to a probability distribution via softmax.
Also called: logit
Why it matters: Operating on logits rather than probabilities — for sampling, calibration, or distillation — is standard practice and avoids precision issues.
For example, a model might output logits of (2.1, 0.4, -1.0) over three tokens, which the softmax then turns into probabilities of roughly (0.78, 0.14, 0.08).
Heard on the show
“By layer thirty — the second-to-last layer — the model's logit-lens probability of saying Cooperate has hit point-eight-four.”Episode 018 — Language Models Compute the Rational Move, Then Override It