Definition
Plain language
A benchmark suite for testing how well models handle long documents.
As stated in the literature
A multi-task evaluation suite for long-context understanding spanning summarization, retrieval, and reasoning over extended inputs.
Why it matters: It provides a standard way to compare long-context models on tasks more realistic than synthetic needle-in-a-haystack tests.
For example, one LongBench task gives the model a long news article and asks for a faithful summary, while another asks it to find a specific fact buried mid-document.
Heard on the show
“Jessica, on most of the longer benchmarks — LongBench, RULER — the accuracy edge over FlashAttention is modest.”Episode 036 — Sparse Attention Was the Wrong Frame. Treat It as Geometry Instead.