Concept · 1 episode(s)

LLM Routing

← all concepts

Definition

LLM routing is the practice of using a lightweight decision layer to choose which language model should handle each incoming query, typically picking between a cheap, fast model and a costly, more capable one before any generation happens. It matters because it can cut inference spend substantially while preserving quality on most traffic, but it also introduces a new component whose misjudgments—sending hard queries to weak models or easy ones to expensive models—can be audited and measured independently of the models themselves.

Episodes covering this