Definition
Plain language
A popular lightweight program for running AI models on ordinary personal computers.
As stated in the literature
A C/C++ inference implementation with its own tokenizer, GGUF weight format, and quantization support, targeting CPU and consumer GPU execution.
Also called: llama cpp
Why it matters: It is a large part of why running capable models locally became practical, which moves model use out of a handful of company data centres and onto ordinary hardware.
For example, a hobbyist with a gaming laptop and no cloud account can use llama.cpp to run a chat model entirely offline.
Heard on the show
“… vLLM, SGLang, TensorRT-LLM, llama.cpp, and ollama are five separate open-source projects, written by different teams, in different languages: …”Episode 272 — How a Model Guesses Which Engine Is Running It, From a Wrong Date