Definition
Plain language
A benchmark that tests whether a model can predict the last word of a passage that requires understanding the whole context.
As stated in the literature
A long-context word prediction benchmark requiring broad discourse understanding; a standard zero-shot evaluation for pretrained language models.
Also called: Lambada
Why it matters: It's a classic check that a model is doing real discourse-level prediction rather than just local pattern completion.
For example, the model reads a short paragraph that ends mid-sentence and must predict the final word, which only makes sense if it has tracked the whole story.
Heard on the show
“And the absolute scores we're talking about — HellaSwag in the low forties, LAMBADA around forty — they're in a regime where the gaps between architectures are small and benchmark noise is real.”Episode 033 — Echo: The Paper Arguing You Never Needed a KV Cache for Retrieval