Definition
Model extraction is an attack in which an adversary reconstructs the functional behavior of a target model or hidden capability purely by querying it and observing outputs, without ever accessing its parameters, training data, or internal text. The extracted stand—in can then be probed, copied, or exploited offline, which makes model extraction a core concern for anyone deploying proprietary or safety—filtered models behind an API.
Episodes covering this
Worth reading next
Papers we haven't done a deep dive on yet, but would recommend on this topic.