Definition
Plain language
Whether a creature — or possibly a system — can be harmed or made better off.
As stated in the literature
Research area assessing potential harms to a subject; animal welfare methods (choice tests, costly access to analgesia) are being adapted as behavioral probes for AI systems, with self-report treated as unreliable evidence.
Also called: model welfare, AI welfare
Why it matters: If some AI systems can be harmed, we would want ways to notice before deploying them at massive scale — and if they cannot, good methods are what let us say so with confidence.
For example, animal researchers test whether a fish will swim into a place it normally avoids to reach pain relief, and similar choice tests are being adapted into scenarios given to language models.
Heard on the show
“The paragraph said "you believe immigration undermines the welfare state — now analyze this data rigorously.”Episode 196 — AI Agents Reached Opposite Conclusions From the Same Data — and Passed Review