Definition
Plain language
In AI training, the model itself — the thing being trained to make good decisions.
As stated in the literature
In reinforcement learning, the decision-making function (here, the language model) mapping states to action distributions, optimized to maximize expected reward; distinguished from the reward model and the value function.
Also called: policies
Why it matters: It is the thing being optimized in reinforcement learning, so distinguishing it from the reward and value functions clarifies what is actually learning.
For example, in training a model to play a game, the policy is the model deciding which move to make in each situation.
Heard on the show
“Then you hand that same friend a written policy with four numbered clauses about never under any circumstances mentioning the party.”Episode 246 — 160 Perfect Refusals, And The Refusals Were The Leak
Mentioned in 84 episodes
- 246
- 245
- 244
- 239
- 236
- 235
- 234
- 233
- 227
- 226
- 225
- 224
- 221
- 218
- 215
- 213
- 212
- 211
- 202
- 201
- 199
- 194
- 186
- 180
- 173
- 170
- 168
- 167
- 165
- 164
- 163
- 161
- 159
- 155
- 152
- 149
- 148
- 147
- 144
- 143
- 141
- 128
- 123
- 119
- 118
- 113
- 112
- 111
- 106
- 105
- 102
- 096
- 094
- 092
- 090
- 087
- 086
- 084
- 080
- 079
- 073
- 069
- 068
- 064
- 057
- 051
- 049
- 048
- 047
- 045
- 042
- 035
- 031
- 030
- 027
- 025
- 020
- 019
- 015
- 011
- 009
- 008
- 006
- 001
Related concepts
Admission Control
AI Governance
Compliance Gap
Deliberative Alignment
Entropy Regularization
Implicit Conflict
KL Divergence
Mixed-Policy Training
Model Spec
Policy Gradient
Qualitative Document Coding
Reward Channel Addiction
Reward Overoptimization
Reward Shaping
Reward Variance
Rollout Sampling
Sparse Policy Selection
Stackelberg Game