Glossary · Term

data attribution

← all terms

Definition

Plain language

Figuring out which training examples were most responsible for what a model ends up doing.

As stated in the literature

The problem of tracing a model's predictions or behaviors back to the training data that caused them; classical tools include influence functions and TracIn, and switchable-source architectures approximate it cheaply by toggling a single source's contribution on or off.

Why it matters: Knowing which data caused a behavior is essential for fixing harmful outputs, crediting sources, and understanding why a model does what it does.

For example, it can trace a model's tendency to give a certain biased answer back to the specific batch of training documents that taught it that pattern.

Heard on the show

“It's exactly TracIn — a self-influence estimator that's been sitting in the data attribution literature since 2020.”
Episode 025 — The Missing Gradient Term That Predicts Sycophancy in RLHF

Related terms