Prompt
How do I generate speech with SSML tags in an API?
Latest observation
Sep 30, 2026
OpenAI APIWeb search: off
To generate speech with SSML in an API, you generally:
- Send plain text or SSML-formatted XML to a text-to-speech endpoint.
- Set the request to indicate SSML, not plain text.
- The TTS engine parses the SSML tags and produces speech accordingly.
Typical flow
1) Write SSML
Example:
<speak>
Hello <break time="500ms"/> world.
<prosody rate="slow" pitch="+2st">
This part is slower and higher pitched.
</prosody>
</speak>
2) Call the TTS API
Most APIs expose a field like:
input.textinput.ssmltext_type: "ssml"input_format: "ssml"
Example JSON request
{
"voice": "en-US-JennyNeural",
"input": {
"ssml": "<speak>Hello <break time='500ms'/> world.</speak>"
},
"output_format": "audio/mp3"
}
Or sometimes:
{
"text": "<speak>Hello <break time='500ms'/> world.</speak>",
"textType": "ssml"
}
3) Receive audio
The API usually returns:
- a binary audio stream, or
- base64-encoded audio, or
- a URL to download the generated file
Common SSML tags
Useful tags include:
<break time="300ms"/>— pause<prosody rate="slow" pitch="+2st">...</prosody>— adjust speech<emphasis level="strong">...</emphasis>— emphasize words<say-as interpret-as="characters">API</say-as>— spell out text<voice name="...">...</voice>— choose a voice<lang xml:lang="es-ES">...</lang>— switch language
Example
<speak>
<voice name="en-US-JennyNeural">
Hello there. <break time="300ms"/>
Welcome to the demo.
</voice>
</speak>
Things to watch for
- SSML must be well-formed XML.
- Some APIs require the root
<speak>tag. - Not all SSML tags are supported by every vendor.
- You may need to escape special characters like
&,<, and>inside text.
If you want, I can show you how to do this for a specific provider like AWS Polly, Google Cloud TTS, Azure Speech, or OpenAI.