Definition
Plain language
The program that actually runs a model and turns your words into its reply.
As stated in the literature
The serving stack around the weights — templating, tokenization, batching and scheduling, the forward pass, sampling, detokenization, and tool-call parsing; examples include vLLM, SGLang, TensorRT-LLM, llama.cpp, and ollama.
Also called: inference engines, inference-engine, serving engine
Why it matters: The same model weights can behave noticeably differently depending on which engine serves them, so the engine — not just the model — determines speed, cost, output quirks, and much of the attack surface.
For example, when you send a message to a locally hosted model, the inference engine formats your text, splits it into tokens, runs it through the weights, picks each next word, and stitches the reply back into readable text.
Heard on the show
“A team at Harvard stacked a few probes like that, and fingerprinted five different AI inference engines from the inside.”Episode 272 — How a Model Guesses Which Engine Is Running It, From a Wrong Date