Glossary · Term

alignment tax

← all terms

Definition

Plain language

The cost in raw ability or accuracy a model pays from being trained to be helpful and safe.

As stated in the literature

The observed drop in capability, calibration, or output diversity that often accompanies RLHF and instruction tuning, originally named by Ouyang et al. in the InstructGPT paper.

Why it matters: It's a tradeoff product teams have to make: more alignment training means more safety but often less raw capability or diversity.

For example, a base model might solve a tricky reasoning puzzle that its RLHF-tuned descendant refuses to engage with.

Heard on the show

“The bigger picture is the alignment tax.”
Episode 070 — When Models Know the Answer But Say the Wrong Thing Anyway

Related terms