Glossary · Term

deduplication

← all terms

Definition

Plain language

Throwing away near-identical copies so the same thing doesn't show up over and over.

As stated in the literature

Near-duplicate detection and removal in a corpus or retrieval result set, typically via hashing or embedding similarity thresholds; a partial mitigation against repeated echoes of the same source dominating a context window.

Also called: dedup, deduplicate, deduplicated

Why it matters: Without it, one loud source can fill a model's entire working context and look like independent corroboration when it is really a single echo.

For example, if the same press release has been reposted on twenty sites, deduplication keeps one copy so the model doesn't read the same claim twenty times and treat it as twenty confirmations.

Heard on the show

“… of the four hundred twenty-one confirmed crashes — they collapse to three seventy-nine after deduplication — about sixty percent could also be reproduced by a fuzzer once you seeded it with the witness …”
Episode 014 — Why a Constrained Pipeline Beat a Full Coding Agent at Finding Bugs 30-to-1

Related terms