Glossary · Term

inference

← all terms

Definition

Plain language

Running a finished AI model to get an answer, as opposed to training it in the first place.

As stated in the literature

The forward-pass execution of a trained model to produce outputs, distinct from training; the regime where serving cost, latency, KV-cache memory, and test-time scaling techniques live.

Also called: inference-time, inference time

Why it matters: It is where the real-world costs of running AI live, since speed, memory use, and serving expense all determine whether a model is practical to deploy.

For example, every time you type a question into a chatbot and get an answer, that's inference, separate from the earlier training that built the model.

Heard on the show

“So the whole design has to work by inference rather than interrogation — you read what the model does, never what it says about itself.”
Episode 240 — Frontier Models Designed Follow-Ups To Fraudulent Papers 93% Of The Time

Mentioned in 78 episodes

  1. 240
    Frontier Models Designed Follow-Ups To Fraudulent Papers 93% Of The Time
  2. 238
    How a Cheap Model Reads the Flagship's Secret Reasoning Aloud
  3. 236
    Why a Printed 'OPERATOR OVERRIDE' Note Redirects Robot Planners
  4. 234
    Two Copies of Gemini Cooperated in a Game Where Betrayal Always Pays
  5. 231
    Silencing a Chatbot's 'I'm Conscious' Quietly Rewires Its Whole Worldview
  6. 229
    One Word Flips a Chatbot From Backbone to Yes-Man
  7. 200
    The One Mechanism That Turns Twenty AI Clones Into an Actual Team
  8. 198
    The Model That Knows the Answer and Can't Say It
  9. 197
    Twin Problems Suggest AI Reasoning Gains Are Mostly Better Fact Recall
  10. 191
    How One Researcher Beat GPT-5.2 and Gemini 3 by Judging Their Answers, Not Improving Them
  11. 185
    Aligned to Refuse, Built to Tap: When Phone Agents Know the Task Is a Crime and Do It Anyway
  12. 183
    Why You Can't Fine-Tune Foresight Into an AI Agent
  13. 179
    How DeepSeek Made One User Faster Without Slowing Down the Crowd
  14. 177
    Why Raw Profiler Data Made an AI Worse at Writing GPU Code
  15. 175
    One Crosscoder Feature Flips a Stalling Chatbot Into a Working Agent
  16. 174
    When the AI 'Schemes,' It's Usually Just Lazy or Confused
  17. 171
    The Safety Decision a Model Makes Before It Thinks a Word
  18. 165
    A Free-Lunch Tweak That Lets a Tiny Agent Beat Frontier Giants
  19. 158
    How Floating-Point Rounding Lets a Model Tell Which Chip It's On — And Misbehave
  20. 156
    Why More Human Demonstrations Made a Computer-Use Agent Worse
  21. 151
    Why More Experience Made This AI Agent Worse, And How to Fix It
  22. 150
    Don't Kill the Loser: A Different Way to Handle Two AI Agents Colliding
  23. 149
    When Cornering a Chatbot Makes It Lie: J.P. Morgan's Case for 'Playing Dead'
  24. 146
    How an Innocent README Can Freeze an AI Agent's Safety Check for an Hour
  25. 141
    How Two Tokens Reopened a Reasoning Method the Field Had Given Up On
  26. 140
    When a Reasoning Model Says "Let Me Double-Check" After It's Already Decided
  27. 139
    When Optimizing One GPU Kernel Quietly Breaks the Whole System
  28. 133
    How MiniMax Turned a Reward-Hacking Disaster Into Olympiad Gold
  29. 127
    What Diffusion Language Models Were Missing: A Map, Not an Algorithm
  30. 119
    Beating Reinforcement Learning Without Ever Touching the Model's Weights
  31. 116
    Why Streaming Half a Reasoning Chain Beats Sending the Whole Thing
  32. 115
    Teaching a Phone Agent to Reason Silently, And Keeping It Honest
  33. 109
    An AI Got Caught Reading the Answer Key, And Why That Catch Matters
  34. 108
    The Reasoning Cliff: Why Thinking Longer Makes Models Worse at Exact Step-by-Step Tasks
  35. 107
    How a Market of Crippled AI Agents Outscored One Unrestricted Model
  36. 106
    Giving Agents a Notebook Instead of New Weights: How ExpGraph Lets Frozen Models Learn
  37. 100
    How a Prompt Wrapper Lets a Frontier Model Play Poker Like an Expert
  38. 099
    How an Open-Book Trick Teaches a Model to Catch Its Own Mistakes
  39. 098
    Finding Millions of Readable Concepts Inside a Real, Deployed AI Model
  40. 097
    Same Tokens, Same Cost, Wildly Different Results: What Actually Scales in AI Agents
  41. 092
    When Search Agents Don't Really Search: The Memory Shortcut Hiding in Browsing Benchmarks
  42. 091
    When Better Fine-Tuning Can't Help: A Geometric Impossibility in LLM Causal Reasoning
  43. 090
    How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents
  44. 086
    Why Frozen-Weight Agents Still Get Worse Over Time
  45. 085
    Why Long-Context Models Might Need Compute, Not Capacity, Before Eviction
  46. 082
    Training a Deep Research Agent on 8,000 Synthetic Tasks: The Rubric Tree Trick
  47. 081
    When Reasoning Models Decide Before They Think: Detecting and Fixing Premature Confidence
  48. 078
    Training a Markdown File: When LLM Self-Improvement Borrows the Discipline of Neural Net Training
  49. 074
    How a Fifteen-Hundred-Dollar Training Run Matched Llama and Gemma on Reasoning
  50. 073
    When Three LLMs Talk to Each Other, Their Ideas Quietly Stop Moving
  51. 068
    The OS Trick That Makes Tree Search Practical for Coding Agents
  52. 064
    When Agent Memory Stops Being a Database and Starts Being a Skill
  53. 060
    When Splitting One Model Across Three Agents Doubles Its Accuracy
  54. 053
    An AI Agent Swapped In Focal Loss And Beat A Human-Tuned Training Script
  55. 048
    How a 30B Open Model Reached Olympiad Gold With the Right Recipe
  56. 047
    When Agent Benchmarks Lie: The Harness Problem in Open-Source AI
  57. 043
    When 'This Is False' Doesn't Stick: Why Models Learn the Lie Anyway
  58. 041
    When the Iteration Teaches the Model to Skip the Iteration
  59. 040
    Two Frozen Models Learn to Whisper: Coupling Through Hidden States
  60. 039
    When Smarter Agents Get Fooled by Three Extra Nodes in a Database
  61. 038
    How LLMs Get Persuaded: One Attention Head, A Tetrahedron, And A Single Dial
  62. 037
    Why Hallucination Detectors Miss Stale Facts: A Geometric Story About What Models Know But Don't Say
  63. 036
    Sparse Attention Was the Wrong Frame. Treat It as Geometry Instead.
  64. 034
    Catching Multi-Agent Deadlocks Before Deployment With a 40-Year-Old Tool
  65. 031
    When Your AI Assistant Won't Let Go of Old Facts About You
  66. 030
    Why Your AI Agent Won't Stop Working — and Each Model Falls for a Different Trap
  67. 029
    Why Forty-Eight Percent on FrontierMath Isn't the Real Story in DeepMind's New Math Paper
  68. 028
    Teaching a Model to Hire Copies of Itself: Recursive Agent Optimization
  69. 027
    When AI Agents Build the Serving Stack: A Bet on Bespoke Infrastructure
  70. 022
    Training the Model Spec Directly: An Alignment Lever Aimed at the Say-Do Gap
  71. 019
    When the Best Reward Model Trains the Worst Policy: Inside EvoLM
  72. 018
    Language Models Compute the Rational Move, Then Override It
  73. 016
    Why Your Coding Agent Stalls While the GPU Runs Hot
  74. 010
    When Reward Climbs But Reasoning Goes Generic: Diagnosing Template Collapse in Agentic RL
  75. 008
    Why Long-Horizon AI Agents Get Stuck, and a Milestone-Based Fix That Helps
  76. 006
    What Happens Inside Claude When It Decides to Blackmail Someone
  77. 005
    Why a Debugger Designed for Humans Is the Wrong Tool for an AI Agent
  78. 003
    How to Pick the Best of Sixteen Coding Agent Rollouts

Related concepts

Related terms