Definition
Plain language
The part of a system that turns recorded sound into numbers the rest of the model can work with.
As stated in the literature
The front-end module of a speech-capable multimodal model that maps raw waveform or spectrogram input into a sequence of continuous representations consumed by the language model; its representation space can already blur categories before fusion with text.
Also called: audio encoders
Why it matters: Whatever the encoder throws away or muddles at this first step can never be recovered later, so its choices quietly set a ceiling on what the whole system can hear.
For example, when you speak into a voice assistant, the audio encoder is the piece that converts your recorded voice into a long list of numbers before anything else can interpret it.
Heard on the show
“They examine the model’s internal audio representations — the numerical patterns produced by its audio encoder before the language model reasons about them.”Episode 262 — Raise the Pitch Nine Percent and the Model Cries Sarcasm