Prompt

How do I generate speech with SSML tags in an API?

Artificial Intelligence / Speech & Voice AI1 observationLast seen Sep 30, 2026

Latest observation

Sep 30, 2026OpenAI APIWeb search: off

To generate speech with SSML in an API, you generally:

  1. Send plain text or SSML-formatted XML to a text-to-speech endpoint.
  2. Set the request to indicate SSML, not plain text.
  3. The TTS engine parses the SSML tags and produces speech accordingly.

Typical flow

1) Write SSML

Example:

<speak>
  Hello <break time="500ms"/> world.
  <prosody rate="slow" pitch="+2st">
    This part is slower and higher pitched.
  </prosody>
</speak>

2) Call the TTS API

Most APIs expose a field like:

  • input.text
  • input.ssml
  • text_type: "ssml"
  • input_format: "ssml"

Example JSON request

{
  "voice": "en-US-JennyNeural",
  "input": {
    "ssml": "<speak>Hello <break time='500ms'/> world.</speak>"
  },
  "output_format": "audio/mp3"
}

Or sometimes:

{
  "text": "<speak>Hello <break time='500ms'/> world.</speak>",
  "textType": "ssml"
}

3) Receive audio

The API usually returns:

  • a binary audio stream, or
  • base64-encoded audio, or
  • a URL to download the generated file

Common SSML tags

Useful tags include:

  • <break time="300ms"/> — pause
  • <prosody rate="slow" pitch="+2st">...</prosody> — adjust speech
  • <emphasis level="strong">...</emphasis> — emphasize words
  • <say-as interpret-as="characters">API</say-as> — spell out text
  • <voice name="...">...</voice> — choose a voice
  • <lang xml:lang="es-ES">...</lang> — switch language

Example

<speak>
  <voice name="en-US-JennyNeural">
    Hello there. <break time="300ms"/>
    Welcome to the demo.
  </voice>
</speak>

Things to watch for

  • SSML must be well-formed XML.
  • Some APIs require the root <speak> tag.
  • Not all SSML tags are supported by every vendor.
  • You may need to escape special characters like &, <, and > inside text.

If you want, I can show you how to do this for a specific provider like AWS Polly, Google Cloud TTS, Azure Speech, or OpenAI.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.