Definition
Plain language
All the software wrapped around an AI model that takes your request, preps it, and sends the answer back.
As stated in the literature
The layers between client and model weights — request preprocessing, templating, batching, safety filters, routing — any of which can alter behavior independently of the checkpoint itself.
Also called: serving-stack, serving stacks
Why it matters: A model's visible behavior can shift because of a change anywhere in this plumbing, so results attributed to "the model" may really be results about the wrapper around it.
For example, between typing a question and seeing an answer, your text may be reformatted into a template, bundled with other users' requests, and screened by a filter.
Heard on the show
“It says modern LLM serving stacks — vLLM, TGI, the whole open-source set — were designed for a world that doesn't exist anymore.”Episode 016 — Why Your Coding Agent Stalls While the GPU Runs Hot