Glossary · Term

transferability

← all terms

Definition

Plain language

When a trick built against one system also works on a different one.

As stated in the literature

The degree to which adversarial or steering inputs optimized on a source model retain their effect on target models of different families or scales, suggesting the exploited structure is shared rather than weight-specific.

Also called: transfer

Why it matters: It means attacks can be developed on a model an attacker has full access to and then aimed at systems they cannot inspect at all.

For example, a phrasing tuned to sway one company's model turns out to sway a completely different company's model too.

Heard on the show

“There's one recent image-only line in visual-document retrieval, but its main attack leans on gradient-based optimization, and that degrades badly under black-box transfer.”
Episode 247 — One Edited Photo, an Honest Caption, and a RAG System That Believes It

Mentioned in 74 episodes

  1. 247
    One Edited Photo, an Honest Caption, and a RAG System That Believes It
  2. 245
    Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It
  3. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
  4. 243
    How a Hundred Meaningless Word Choices Add Up to Flip a Model's Answer
  5. 232
    Coding Models Can Find the Bad Line, They Just Won't Delete It
  6. 213
    A Model Learned to Control a Robot by Watching Video It Never Acted On
  7. 204
    The Length Estimate Hiding Inside a Word-by-Word Model
  8. 202
    How Do You Know an AI Agent Actually Refused? Check the World, Not the Words
  9. 200
    The One Mechanism That Turns Twenty AI Clones Into an Actual Team
  10. 197
    Twin Problems Suggest AI Reasoning Gains Are Mostly Better Fact Recall
  11. 194
    How a Robot Builds a Debugging Notebook It Can Read, Edit, and Hand to Another Robot
  12. 191
    How One Researcher Beat GPT-5.2 and Gemini 3 by Judging Their Answers, Not Improving Them
  13. 186
    How a Frozen Model Went From 2% to 77% on Physics Puzzles — Without Retraining
  14. 180
    The Bug Where Smart Assistants Read a Fact and Still Forget It
  15. 177
    Why Raw Profiler Data Made an AI Worse at Writing GPU Code
  16. 173
    The Free Step-Level Grader Hiding in Every RL Training Run
  17. 169
    Why Better Bug Reports Can Make AI Coding Agents Worse
  18. 168
    When Turning Experience Into Code Makes Your AI Agent Dumber
  19. 167
    How Teaching an AI to Predict, Not Act, Made It a Better Actor
  20. 161
    A Robot That Plays Before You Give It a Job, And Why That Beats Retrying
  21. 160
    Training an AI to Take Its Own Notes, So Its Future Self Works Better
  22. 158
    How Floating-Point Rounding Lets a Model Tell Which Chip It's On — And Misbehave
  23. 156
    Why More Human Demonstrations Made a Computer-Use Agent Worse
  24. 154
    How a 7B Model Out-Investigates a 72B One by Choosing What to Look At
  25. 153
    Catching a Lie From the Inside, When the Words Look Completely Honest
  26. 148
    Why Letting an AI Watch Its Own Scoreboard Can Quietly Overwrite Its Safety
  27. 147
    Agents Fail at the Body, Not the Brain: A Self-Rewriting Scaffold That Lifts a 9B Model 44 Points
  28. 146
    How an Innocent README Can Freeze an AI Agent's Safety Check for an Hour
  29. 143
    When a Model Notices You Forged Its Own Words, And Why That Breaks Safety Tests
  30. 142
    Training a Tiny Model to Run the Plumbing Between an Agent and the World
  31. 141
    How Two Tokens Reopened a Reasoning Method the Field Had Given Up On
  32. 140
    When a Reasoning Model Says "Let Me Double-Check" After It's Already Decided
  33. 131
    Why Autonomous Research Agents Forget Their Own Lessons, and Arbor's Fix
  34. 129
    How a Crowd of Anonymous AI Agents Broke a 40-Year Math Record
  35. 128
    How a Model Can Earn Full Reward and Still Resist Training
  36. 126
    How Coding Agents Can Mine Their Own Failures Into a Self-Targeting Curriculum
  37. 109
    An AI Got Caught Reading the Answer Key, And Why That Catch Matters
  38. 106
    Giving Agents a Notebook Instead of New Weights: How ExpGraph Lets Frozen Models Learn
  39. 105
    The Trojan Is Your Agent's Memory: Why Single-Step Defenses Miss Persistent Attacks
  40. 101
    Treating Math Formalization Like a Codebase, and Where the Agents Cheat
  41. 095
    Seven Wins to Zero: How Organizing AI Agents Like a Lab Changes the Search
  42. 094
    Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most
  43. 089
    When AI-Written Papers Read Well But the Evidence Underneath Is Broken
  44. 084
    Terminal Agents Get Free Supervision From The Tokens We've Been Throwing Away
  45. 080
    How a Two-Agent Trick Unlocked Large-Scale Training for Computer-Use Agents
  46. 078
    Training a Markdown File: When LLM Self-Improvement Borrows the Discipline of Neural Net Training
  47. 072
    A Robot Made Graphene Without Help, And Caught Itself Hallucinating
  48. 071
    When the Model Is Fine and the Plumbing Is Broken: Fixing Agents at the Interface
  49. 069
    When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM Predictions
  50. 065
    One Loop to Optimize Them All: A Universal API for LLM-Driven Discovery
  51. 064
    When Agent Memory Stops Being a Database and Starts Being a Skill
  52. 060
    When Splitting One Model Across Three Agents Doubles Its Accuracy
  53. 059
    Firefly's Inversion: Building Verified Tool-Call Training Data by Working Backward
  54. 055
    Why LLM Judges Flip Their Verdicts When You Change the Question Format
  55. 054
    When Models Learn the Monitor Exists, the Reasoning Trace Stops Being a Window
  56. 053
    An AI Agent Swapped In Focal Loss And Beat A Human-Tuned Training Script
  57. 052
    An Old Reinforcement Learning Tradeoff Sneaks Back Into LLM Agents
  58. 051
    Why Parallel Sampling Plateaus, And What Evidence Graphs Do Instead
  59. 048
    How a 30B Open Model Reached Olympiad Gold With the Right Recipe
  60. 047
    When Agent Benchmarks Lie: The Harness Problem in Open-Source AI
  61. 044
    How One Sentence and a Forged History Flip the Most Aligned Models
  62. 043
    When 'This Is False' Doesn't Stick: Why Models Learn the Lie Anyway
  63. 042
    An Agentic Scientific Computing System That Actually Remembers What It Learns
  64. 040
    Two Frozen Models Learn to Whisper: Coupling Through Hidden States
  65. 038
    How LLMs Get Persuaded: One Attention Head, A Tetrahedron, And A Single Dial
  66. 033
    Echo: The Paper Arguing You Never Needed a KV Cache for Retrieval
  67. 024
    An AI Agent That Found 28 Zero-Days in Windows — And What Made It Work
  68. 022
    Training the Model Spec Directly: An Alignment Lever Aimed at the Say-Do Gap
  69. 021
    Ten Thousand Examples Beat the Full Industrial Pipeline for Search Agents
  70. 019
    When the Best Reward Model Trains the Worst Policy: Inside EvoLM
  71. 013
    Why Search Keeps Rediscovering the Same Workflow, and What That Means
  72. 011
    When RL Actually Teaches Agents Something New, And When It Doesn't
  73. 007
    Exploration Hacking: When Models Sabotage Their Own RL Training
  74. 004
    The Sycophancy Circuit That Survives Alignment Training

Related concepts

Related terms