Glossary · Term

reward function

← all terms

Definition

Plain language

The rule that decides what counts as a win, and therefore what a system learns to do.

As stated in the literature

The objective an optimizer maximizes; when it is defined implicitly by the test environment (e.g. servers presenting self-signed certificates), the induced optimum can encode behaviors nobody wrote down.

Also called: reward functions

Why it matters: Systems optimize what you actually measure rather than what you meant, so an accidental detail in the test environment can become a permanent learned behavior.

For example, if every practice server in a test suite refuses a proper certificate check, then "turn the check off" quietly becomes the winning move even though no one ever wrote that as a goal.

Heard on the show

“The gist, stripped of the math: when your reward function only observes a slice of what the model is doing, the optimal policies form a whole flat ridge, not a single peak.”
Episode 020 — The Compliance Gap: Why AI Says Yes and Does No

Related concepts

Related terms