Glossary · Term

fine-tuning

← all terms

Definition

Plain language

Taking a model that already knows a lot and giving it more lessons so it gets better at one specific task.

As stated in the literature

Continuing the training of a pretrained model on a smaller task-specific dataset, updating its weights to adapt the general model to a particular use case.

Also called: Fine-tuning, fine-tune, fine-tuned, finetuning, finetune, finetuned, finetunes

Why it matters: It's the standard way to adapt a generic foundation model to a specific domain or task without training from scratch.

For example, you might take a general language model and fine-tune it on a few thousand customer-support transcripts so it adopts your company's tone and product knowledge.

Heard on the show

“Make the refusal behavior harder to locate in the weights, harder to cut out, and harder to fine-tune away.”
Episode 244 — The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers

Mentioned in 67 episodes

  1. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
  2. 242
    Making a Vision Model Better by Showing It Blurry Images
  3. 230
    Why AI Survey Panels Break Before the Dice Ever Roll
  4. 225
    How a Frozen Model Went From Zero to Sixty Percent by Borrowing Another's Thinking
  5. 221
    Two Hundred Clean Economics Answers, And a Model That Endorses Race Science
  6. 214
    The Medical AI Answer That's Accurate, Sourced, and Still Wrong
  7. 213
    A Model Learned to Control a Robot by Watching Video It Never Acted On
  8. 212
    The Fact Was in the Wrong Drawer: Why Fine-Tuned Models Can't Reason With What They Know
  9. 199
    Finding a Model's Hidden Behaviors Without Knowing What You're Looking For
  10. 193
    Freeze Most of the Network: Where RL Improvement Actually Lives in a Transformer
  11. 189
    Why Phone Agents Ace the Test and Crash on Your Actual Phone
  12. 188
    A Coding Agent Found a Hole in a Peer-Reviewed STOC Proof for Five Dollars
  13. 187
    An 8-Billion Agent That Beats Models 80 Times Its Size By Looking Things Up
  14. 186
    How a Frozen Model Went From 2% to 77% on Physics Puzzles — Without Retraining
  15. 185
    Aligned to Refuse, Built to Tap: When Phone Agents Know the Task Is a Crime and Do It Anyway
  16. 183
    Why You Can't Fine-Tune Foresight Into an AI Agent
  17. 180
    The Bug Where Smart Assistants Read a Fact and Still Forget It
  18. 175
    One Crosscoder Feature Flips a Stalling Chatbot Into a Working Agent
  19. 170
    When a One-Liner Beats Your Agent's Clever Verification Logic
  20. 169
    Why Better Bug Reports Can Make AI Coding Agents Worse
  21. 167
    How Teaching an AI to Predict, Not Act, Made It a Better Actor
  22. 166
    A Router That Beats the Frontier Models It Calls
  23. 164
    The Summarizer That Quietly Deletes Your Agent's Safety Rules
  24. 163
    Why Training Only on Perfect Solutions Cripples a Model's Reasoning
  25. 161
    A Robot That Plays Before You Give It a Job, And Why That Beats Retrying
  26. 158
    How Floating-Point Rounding Lets a Model Tell Which Chip It's On — And Misbehave
  27. 157
    When an AI Coding Agent Drives a Phone Through the Terminal, No Screen Needed
  28. 156
    Why More Human Demonstrations Made a Computer-Use Agent Worse
  29. 154
    How a 7B Model Out-Investigates a 72B One by Choosing What to Look At
  30. 153
    Catching a Lie From the Inside, When the Words Look Completely Honest
  31. 152
    Training a Model to Mean What It Says, And Why That Isn't the Same as Being Good
  32. 151
    Why More Experience Made This AI Agent Worse, And How to Fix It
  33. 150
    Don't Kill the Loser: A Different Way to Handle Two AI Agents Colliding
  34. 148
    Why Letting an AI Watch Its Own Scoreboard Can Quietly Overwrite Its Safety
  35. 146
    How an Innocent README Can Freeze an AI Agent's Safety Check for an Hour
  36. 145
    Building Forgetting Into a Language Model With One Extra Line of Code
  37. 142
    Training a Tiny Model to Run the Plumbing Between an Agent and the World
  38. 128
    How a Model Can Earn Full Reward and Still Resist Training
  39. 119
    Beating Reinforcement Learning Without Ever Touching the Model's Weights
  40. 115
    Teaching a Phone Agent to Reason Silently, And Keeping It Honest
  41. 114
    Agents That Rewrite Their Own Weights Instead of Just Taking Notes
  42. 110
    How an Agent Got 44 Points Better by Mining Its Own Scratch Paper
  43. 108
    The Reasoning Cliff: Why Thinking Longer Makes Models Worse at Exact Step-by-Step Tasks
  44. 106
    Giving Agents a Notebook Instead of New Weights: How ExpGraph Lets Frozen Models Learn
  45. 094
    Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most
  46. 091
    When Better Fine-Tuning Can't Help: A Geometric Impossibility in LLM Causal Reasoning
  47. 083
    Training the Translator: How a Small Communication Model Lets Agent Teams Outperform Themselves
  48. 080
    How a Two-Agent Trick Unlocked Large-Scale Training for Computer-Use Agents
  49. 078
    Training a Markdown File: When LLM Self-Improvement Borrows the Discipline of Neural Net Training
  50. 074
    How a Fifteen-Hundred-Dollar Training Run Matched Llama and Gemma on Reasoning
  51. 071
    When the Model Is Fine and the Plumbing Is Broken: Fixing Agents at the Interface
  52. 061
    When Helpful Agents Go Sideways: A 404 Error, Campus Security, and Why Alignment Misses This
  53. 054
    When Models Learn the Monitor Exists, the Reasoning Trace Stops Being a Window
  54. 052
    An Old Reinforcement Learning Tradeoff Sneaks Back Into LLM Agents
  55. 048
    How a 30B Open Model Reached Olympiad Gold With the Right Recipe
  56. 043
    When 'This Is False' Doesn't Stick: Why Models Learn the Lie Anyway
  57. 040
    Two Frozen Models Learn to Whisper: Coupling Through Hidden States
  58. 038
    How LLMs Get Persuaded: One Attention Head, A Tetrahedron, And A Single Dial
  59. 032
    A Sticky-Note for Every Layer: Letting Transformers Remember What They Were Just Thinking
  60. 026
    What RL Actually Does to Language Models, at the Token Level
  61. 022
    Training the Model Spec Directly: An Alignment Lever Aimed at the Say-Do Gap
  62. 021
    Ten Thousand Examples Beat the Full Industrial Pipeline for Search Agents
  63. 019
    When the Best Reward Model Trains the Worst Policy: Inside EvoLM
  64. 009
    How Two Silent Library Bugs Quietly Invalidated a Wave of Reasoning Papers
  65. 008
    Why Long-Horizon AI Agents Get Stuck, and a Milestone-Based Fix That Helps
  66. 007
    Exploration Hacking: When Models Sabotage Their Own RL Training
  67. 004
    The Sycophancy Circuit That Survives Alignment Training

Related concepts

Related terms