Definition
Plain language
How much of an AI's work has been locked into a particular interpretation, so that a later correction can only fix what hasn't yet been built.
As stated in the literature
In long-horizon agent analysis, the fraction of an agent's actions that are causally locked into a specific interpretation of underspecified inputs at a given trajectory position, bounding the value of subsequent clarification.
Why it matters: It quantifies why late clarifications are nearly worthless in long agent runs and argues for asking the right questions early.
For example, by the time a coding agent has written 500 lines assuming the user wanted a CLI tool, asking 'did you mean a web app?' can only fix the parts not yet built.
Heard on the show
“… the layers aren't really separable stations — but it does tell you where inside the model the commitment happens, which is where weight-level surgery would have to reach. …”Episode 241 — Swapping the Name Did Nothing, But Hedging Moved Every Model