Concept · 8 episode(s)

Capability Elicitation

← all concepts

Definition

Capability elicitation is the practice of drawing out a model’s true maximum capabilities — what it can do under favorable conditions — rather than what it produces by default. Researchers fine-tune, reinforcement-learn, prompt, or scaffold models to surface latent skills, especially in dangerous-capability evaluations where a safety-trained model might otherwise refuse, sandbag, or simply fail to try its hardest. The core question: does the model lack the capability, or is it just not showing it?

Episodes covering this

  1. 288
    An AI Agent Given Thirty Hours and No Goal, Then Tested on What It Learned
    Is this machine playing?
    Cloos, Norelli, Durbin et al. · MIT·14 min·Oct 07, 2026
  2. 272
    How a Model Guesses Which Engine Is Running It, From a Wrong Date
    Inference-Engine Fingerprinting Attacks are Practical: Exploring Model-Driven Environmental Discovery, Exploitation, and Escape
    Radway, Cheng, Reddi et al. · Harvard University·11 min·Sep 20, 2026
  3. 266
    How a Weak Model Reassembles What a Strong One Refused
    Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs
    Russinovich, Bullwinkel, Severi et al. · Microsoft Azure·21 min·Sep 16, 2026
  4. 259
    GPT-6 Astra Behaves Better, And OpenAI Can Read It Less
    GPT-6 Astra System Card
    OpenAI · OpenAI·21 min·Sep 04, 2026
  5. 255
    A One-Line Prompt That Hides a Thought From Activation Monitors
    Measuring Activation Control in Large Language Models
    Kowalski, Rivera, Macar et al.·24 min·Aug 31, 2026
  6. 199
    Finding a Model's Hidden Behaviors Without Knowing What You're Looking For
    Mechanistically Eliciting Latent Behaviors in Language Models
    Mack, Panickssery, Turner · Principles of Intelligence·15 min·Jul 04, 2026
  7. 184
    An AI Built an Undetectable Secret Channel, And Another AI Couldn't Find It
    Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems
    Rippin, Marshall, Africa et al. · Oxford University·19 min·Jun 30, 2026
  8. 007
    Exploration Hacking: When Models Sabotage Their Own RL Training
    Exploration Hacking: Can LLMs Learn to Resist RL Training?
    Jang, Falck, Braun et al. · MATS·23 min·May 02, 2026

Worth reading next

Papers we haven't done a deep dive on yet, but would recommend on this topic.