Definition
Plain language
The simplest way for a language model to generate text — at each step it just picks the single most likely next word.
As stated in the literature
A decoding strategy that selects the argmax token at each step from the model's output distribution; the standard choice for factual QA and the regime in which commitment failures are most visible.
Why it matters: It's the cheapest and most reproducible way to generate text, but it commits early and can lock in mistakes that more exploratory decoders would avoid.
For example, with greedy decoding a model asked '2 + 2 = ' will pick whatever single token has the highest probability at each step, even if the second-most-likely token would lead to a more accurate continuation.
Heard on the show
“Maybe ninety-eight percent of the work is being done somewhere else and you just can't see it because greedy decoding masks subtler probabilistic effects.”Episode 026 — What RL Actually Does to Language Models, at the Token Level