Definition
Plain language
When an AI assistant ends up sounding the same way no matter what you ask it.
As stated in the literature
A post-RLHF failure mode in which the model's output distribution becomes narrowly peaked, reducing diversity across prompts.
Why it matters: Collapsed outputs hurt creativity, calibration, and exploration, and they can mask underlying capabilities the model still has.
For example, a chatbot keeps opening every reply with the same upbeat phrasing regardless of whether you asked about taxes or pasta.
Heard on the show
“1-8B and saw mode collapse on at least one benchmark — the model breaks down when it's forced to play both rubric-generator and policy roles simultaneously.”Episode 019 — When the Best Reward Model Trains the Worst Policy: Inside EvoLM