Glossary · Term

CEBRA

← all terms

Definition

Plain language

A tool borrowed from brain science that sorts a model's internal states by which line of reasoning they belong to, rather than by their loudest surface features.

As stated in the literature

A contrastive neural embedding method that maps high-dimensional activations into a latent space by pulling temporally adjacent states together, used here to make discrete reasoning modes visible in a model's activation trajectory.

Why it matters: It gives researchers a way to see the distinct modes of thinking hidden inside a model's activity, which is otherwise buried in high-dimensional noise.

For example, it can take a model's internal states and group together the moments when it was following the same chain of reasoning, even if those moments looked different on the surface.

Heard on the show

“They instead use an encoder from neuroscience called CEBRA that's trained to pull moments that are next to each other in the reasoning close together.”
Episode 225 — How a Frozen Model Went From Zero to Sixty Percent by Borrowing Another's Thinking

Mentioned in 1 episode

  1. 225
    How a Frozen Model Went From Zero to Sixty Percent by Borrowing Another's Thinking

Related terms