Definition
Plain language
The knife-edge point where a system is nearly tied between two answers and the smallest push tips it either way.
As stated in the literature
The surface in input or representation space where a classifier's predicted class flips; near it, small perturbations produce large changes in the output label while confidence is minimal.
Also called: decision boundaries
Why it matters: Behavior near this edge is where systems are least stable and most easily nudged, which is why both errors and attacks cluster there.
For example, a spam filter that rates one email at 50.1 percent spam and a nearly identical one at 49.9 percent is sitting right on the boundary, where a single extra word flips the verdict.
Heard on the show
“So before the agent ever attempts a full multi-step task, its judgment at decision boundaries has been sharpened directly.”Episode 066 — Why Giving an AI Agent More Tools Can Make It Worse at Using a Computer