Glossary · Customer service

Text-to-speech (TTS)

Software that turns written text into spoken audio, used for IVR prompts, screen readers and the voice of AI voice agents.

Text-to-speech, also called speech synthesis, is software that converts written text into spoken audio.

Where it is used. IVR systems use it for prompts that cannot be prerecorded, such as account balances and dates. Screen readers use it to read interfaces aloud. AI voice agents use it to speak every reply: in a chained design, the language model’s text goes to a text-to-speech engine, and in a speech-to-speech design the model produces audio itself.

Controlling how it sounds. The W3C Speech Synthesis Markup Language (SSML) 1.1, a Recommendation since September 2010, is an XML language for guiding synthesis. It lets authors control pronunciation, volume, pitch and rate. Its main elements include phoneme for exact pronunciation, prosody for pitch, rate and volume, break for pauses, emphasis, say-as for how to read constructs such as dates, sub for substitute text, and voice for choosing a voice. An illustrative prompt that slows down for a confirmation code:

<speak version="1.1" xmlns="http://www.w3.org/2001/10/synthesis" xml:lang="en-US">
  Your confirmation code is <prosody rate="slow">A 4 4 1 7</prosody>.
  <break time="500ms"/>
  Please keep it for your records.
</speak>

Interfaces. In browsers, the W3C Web Speech API’s speech synthesis interface lets a page speak text. On telephony platforms, MRCPv2 (RFC 6787) lets an IVR control a synthesizer running on a separate server.

Synthetic voices on calls. In February 2024 the US Federal Communications Commission ruled that calls made with AI-generated voices count as “artificial” voice calls under the Telephone Consumer Protection Act, so the Act’s rules for artificial or prerecorded voice calls apply to them.

Between agents. When a customer’s AI agent phones a business’s AI agent, text-to-speech on one side and speech-to-text on the other turn data into audio and back again. Each conversion adds delay and a chance of error that a structured request avoids.

Neighbouring terms. Speech-to-text is the reverse step. A voice agent combines both.

Sources

  1. W3C Speech Synthesis Markup Language (SSML) Version 1.1, Recommendation, 7 September 2010 (accessed )
  2. Web Speech API, W3C Draft Community Group Report (accessed )
  3. RFC 6787: Media Resource Control Protocol Version 2 (MRCPv2) (accessed )
  4. FCC Makes AI-Generated Voices in Robocalls Illegal (8 February 2024, Declaratory Ruling FCC 24-17) (accessed )