Definition
Plain language
Running a finished AI model to get an answer, as opposed to training it in the first place.
As stated in the literature
The forward-pass execution of a trained model to produce outputs, distinct from training; the regime where serving cost, latency, KV-cache memory, and test-time scaling techniques live.
Also called: inference-time, inference time
Why it matters: It is where the real-world costs of running AI live, since speed, memory use, and serving expense all determine whether a model is practical to deploy.
For example, every time you type a question into a chatbot and get an answer, that's inference, separate from the earlier training that built the model.
Heard on the show
“So the whole design has to work by inference rather than interrogation — you read what the model does, never what it says about itself.”Episode 240 — Frontier Models Designed Follow-Ups To Fraudulent Papers 93% Of The Time
Mentioned in 78 episodes
- 240
- 238
- 236
- 234
- 231
- 229
- 200
- 198
- 197
- 191
- 185
- 183
- 179
- 177
- 175
- 174
- 171
- 165
- 158
- 156
- 151
- 150
- 149
- 146
- 141
- 140
- 139
- 133
- 127
- 119
- 116
- 115
- 109
- 108
- 107
- 106
- 100
- 099
- 098
- 097
- 092
- 091
- 090
- 086
- 085
- 082
- 081
- 078
- 074
- 073
- 068
- 064
- 060
- 053
- 048
- 047
- 043
- 041
- 040
- 039
- 038
- 037
- 036
- 034
- 031
- 030
- 029
- 028
- 027
- 022
- 019
- 018
- 016
- 010
- 008
- 006
- 005
- 003