Glossary · Term

UCB1

← all terms

Definition

Plain language

A rule for choosing what to try next that favors options that either scored well or haven't been tried much yet.

As stated in the literature

The canonical upper-confidence-bound bandit algorithm, selecting the arm maximizing empirical mean plus a confidence term scaling as sqrt(2 ln t / n_i); achieves logarithmic regret under bounded rewards.

Also called: UCB-1

Why it matters: It gives a concrete, tunable formula for exploration instead of a hunch, so a search agent can justify every attempt it spends.

For example, faced with two strategies where one has averaged a good score over fifty tries and another averaged slightly worse over three tries, UCB1 will keep sampling the barely-tested one until it has enough evidence to rule it out.

Heard on the show

“For the annotated version of everything we just walked through, every term like UCB1 or reward hacking tap to define, and linked out to the papers this one’s arguing with — that’s paperdive dot AI.”
Episode 276 — An AI Agent Rewrote Its Own Scaffolding For Eight Days. Here's What Survived

Related terms