Glossary · Term

LongBench

← all terms

Definition

Plain language

A benchmark suite for testing how well models handle long documents.

As stated in the literature

A multi-task evaluation suite for long-context understanding spanning summarization, retrieval, and reasoning over extended inputs.

Why it matters: It provides a standard way to compare long-context models on tasks more realistic than synthetic needle-in-a-haystack tests.

For example, one LongBench task gives the model a long news article and asks for a faithful summary, while another asks it to find a specific fact buried mid-document.

Heard on the show

“Jessica, on most of the longer benchmarks — LongBench, RULER — the accuracy edge over FlashAttention is modest.”
Episode 036 — Sparse Attention Was the Wrong Frame. Treat It as Geometry Instead.

Related terms