Definition
Plain language
Judging a system on one attempt rather than letting it try many times.
As stated in the literature
Evaluation with a single sampled response per problem; contrasts with best-of-n or iterative-feedback protocols, and can rank models differently from deployment settings that sample heavily.
Also called: single shot
Why it matters: It can rank models very differently from real deployments where systems sample many attempts, so a single-shot leaderboard may not predict which model is actually more useful.
For example, a model is asked each contest problem exactly once and scored on that one answer, with no retries allowed.
Heard on the show
“At k equals one — single shot — the base model is at thirty-six percent, RL is at thirty-four.”Episode 011 — When RL Actually Teaches Agents Something New, And When It Doesn't