Definition
Plain language
The rule that decides what counts as a win, and therefore what a system learns to do.
As stated in the literature
The objective an optimizer maximizes; when it is defined implicitly by the test environment (e.g. servers presenting self-signed certificates), the induced optimum can encode behaviors nobody wrote down.
Also called: reward functions
Why it matters: Systems optimize what you actually measure rather than what you meant, so an accidental detail in the test environment can become a permanent learned behavior.
For example, if every practice server in a test suite refuses a proper certificate check, then "turn the check off" quietly becomes the winning move even though no one ever wrote that as a goal.
Heard on the show
“The gist, stripped of the math: when your reward function only observes a slice of what the model is doing, the optimal policies form a whole flat ridge, not a single peak.”Episode 020 — The Compliance Gap: Why AI Says Yes and Does No