Definition
Plain language
A tuning knob that controls how big each training step is.
As stated in the literature
The scalar coefficient on gradient updates in neural network training, governing step size in parameter space; analogized in SkillOpt to a bounded number of skill-document edits per round.
Why it matters: It's one of the most consequential single hyperparameters in training — too small wastes compute, too large diverges.
For example, with a learning rate of 0.01, each gradient update nudges the weights by 1% of the suggested direction, while 0.1 would take ten times bigger steps.
Heard on the show
“Then it kept going — adjusted learning rate schedules, tuned the depth further.”Episode 053 — An AI Agent Swapped In Focal Loss And Beat A Human-Tuned Training Script