Definition
Plain language
Software that turns spoken audio into written words.
As stated in the literature
Automatic transcription of an audio signal into text; used here as an intelligibility check, with character error rate on filtered or manipulated audio verifying whether lexical content survives a signal manipulation.
Also called: speech-recognition, automatic speech recognition, ASR
Why it matters: Beyond its everyday uses, it doubles as a check on whether audio has been altered so much that the words are no longer there to be heard.
For example, dictating a text message and watching the words appear on screen is speech recognition doing its job.
Heard on the show
“… We're skipping the streaming speech-recognition case and the MacBook JSON-decoding case in detail — both show similar wins, around one-point-seven …”Episode 027 — When AI Agents Build the Serving Stack: A Bet on Bespoke Infrastructure