Definition
Data filtering is the removal or screening of training examples that match certain content, such as specific tokens, topics, or patterns, to keep unwanted traits out of a model. It is a standard defense against data poisoning and trait leakage in distillation pipelines. In subliminal-learning settings it proves insufficient: removing digits, task-related vocabulary, and candidate tokens did not fully block transfer, because the signal is carried by subtle statistical patterns in the outputs rather than by identifiable content.