Definition
Plain language
OpenAI's first reasoning model trained to think out loud before answering.
As stated in the literature
OpenAI's reasoning model post-trained to produce extended chains of thought before final answers.
Also called: o3, o3-mini
Why it matters: It demonstrated that scaling test-time reasoning, not just model size, can dramatically improve performance on hard problems.
For example, o1 might spend tens of seconds writing internal reasoning before producing the final answer to a competition math problem.
Heard on the show
“DeepSeek-R1, OpenAI's o1, Qwen3 — they all use some variant of this pipeline.”Episode 026 — What RL Actually Does to Language Models, at the Token Level