Speech-to-text (STT)
Software that converts spoken audio into written text, the first stage of most voice bots, IVRs and chained AI voice agents.
Speech-to-text, also called automatic speech recognition (ASR), is software that converts spoken audio into written text.
How it is used. In a phone or voice channel, speech-to-text is the first step: the caller’s words become text that a dialog system or language model can act on. In a chained AI voice agent it feeds the model, and text-to-speech speaks the reply. IVR platforms can drive a recognizer on a separate server through MRCPv2 (RFC 6787). In browsers, the W3C Web Speech API’s SpeechRecognition interface returns results as a list of hypotheses, each with a confidence score between 0 and 1, and can deliver interim results while the person is still speaking.
Output is a best guess. A recognizer’s output is its best guess. Confidence scores and alternative hypotheses let the system confirm uncertain parts, for example by reading back a number and asking the caller to confirm it.
Measuring accuracy. The usual measure is word error rate (WER). NIST’s sclite tool, part of its Scoring Toolkit, aligns the recognizer’s output with a correct reference transcript and counts substitutions, deletions and insertions. WER is their sum divided by the number of words in the reference, times 100.
Illustrative example. The reference is “order four four one seven ships friday” (six words). The recognizer hears “order four four seven one ships friday”. Alignment finds two substitutions, so WER is 2 / 6, about 33%. Four of six words are right, yet the order number is wrong, and that is enough to look up the wrong order.
Why it matters between agents. Identifiers such as order numbers, names, addresses and amounts are the parts of a request that matter most, and one misrecognized character changes them. When the customer’s side is itself software, sending those values as structured fields, such as A2A data parts, skips speech recognition entirely.
Neighbouring terms. Text-to-speech is the reverse step. A voice agent chains the two with a model in between.
Sources
- Web Speech API, W3C Draft Community Group Report (accessed )
- NIST SCTK: sclite documentation (accessed )
- RFC 6787: Media Resource Control Protocol Version 2 (MRCPv2) (accessed )
- OpenAI API docs: Voice agents guide (accessed )