Glossary · Term

greedy decoding

← all terms

Definition

Plain language

The simplest way for a language model to generate text — at each step it just picks the single most likely next word.

As stated in the literature

A decoding strategy that selects the argmax token at each step from the model's output distribution; the standard choice for factual QA and the regime in which commitment failures are most visible.

Why it matters: It's the cheapest and most reproducible way to generate text, but it commits early and can lock in mistakes that more exploratory decoders would avoid.

For example, with greedy decoding a model asked '2 + 2 = ' will pick whatever single token has the highest probability at each step, even if the second-most-likely token would lead to a more accurate continuation.

Heard on the show

“Maybe ninety-eight percent of the work is being done somewhere else and you just can't see it because greedy decoding masks subtler probabilistic effects.”
Episode 026 — What RL Actually Does to Language Models, at the Token Level

Related terms