Glossary · Term

FlashAttention

← all terms

Definition

Plain language

A heavily optimized way to compute attention on GPUs that uses memory more carefully.

As stated in the literature

A fused-kernel implementation of exact attention that reduces HBM traffic by tiling and recomputation, dramatically lowering memory and improving throughput.

Also called: FlashAttention-2

Why it matters: It made long-context transformers practical by removing a memory wall, and is now the default attention kernel in most training stacks.

For example, swapping a standard attention implementation for FlashAttention can cut a long-context training run's memory use and speed it up noticeably without changing what the model computes.

Heard on the show

“Tried FlashAttention-2 — produced NaNs.”
Episode 027 — When AI Agents Build the Serving Stack: A Bet on Bespoke Infrastructure

Related terms