Definition
Plain language
A test that checks whether a model's internal picture of a situation puts similar situations close together and different ones far apart.
As stated in the literature
A learned low-dimensional readout from hidden activations trained so that pairwise distances between mapped points match true task distances; unlike a per-feature classifier it tests the geometry of the representation, not just the presence of information.
Also called: distance probes
Why it matters: It reveals whether a model holds a genuinely map-like sense of how situations relate, which plain yes/no readouts of individual facts can miss entirely.
For example, if two puzzle positions are one move apart and a third is twenty moves away, a distance probe checks whether the model's internal numbers place the first two close together and the third far off.
Heard on the show
“And the honest bounding is: two models, both Qwen-derived, one puzzle, one size — four disks, eighty-one states, and the distance probe is fit on all eighty-one of them with nothing held out.”Episode 237 — The Model Built a Perfect Map of the Puzzle, Then Lost It