Glossary · Term

llama.cpp

← all terms

Definition

Plain language

A popular lightweight program for running AI models on ordinary personal computers.

As stated in the literature

A C/C++ inference implementation with its own tokenizer, GGUF weight format, and quantization support, targeting CPU and consumer GPU execution.

Also called: llama cpp

Why it matters: It is a large part of why running capable models locally became practical, which moves model use out of a handful of company data centres and onto ordinary hardware.

For example, a hobbyist with a gaming laptop and no cloud account can use llama.cpp to run a chat model entirely offline.

Heard on the show

“… vLLM, SGLang, TensorRT-LLM, llama.cpp, and ollama are five separate open-source projects, written by different teams, in different languages: …”
Episode 272 — How a Model Guesses Which Engine Is Running It, From a Wrong Date

Related terms