Definition
Plain language
An automatic screener that checks text for harmful content before or after a model sees it.
As stated in the literature
A content classifier applied to inputs or outputs that flags categories like violence, hate, or self-harm; scores content per item, so it can miss harm that emerges only from the aggregate of individually benign statements.
Also called: moderation systems, moderation API, moderation endpoint, content moderation
Why it matters: Because these screeners judge one item at a time, harm that only emerges from the accumulation of individually harmless pieces of text can slip straight through them.
For example, a moderation classifier will flag a message asking how to hurt someone, but may pass a hundred separate messages that each contain one innocuous chemistry fact.
Heard on the show
“Healthcare, finance, AI governance, academic integrity, public health, content moderation, and so on.”Episode 044 — How One Sentence and a Forged History Flip the Most Aligned Models