Prompt
How do I add speech-to-text to my app using an API?
Latest observation
Oct 8, 2026
OpenAI APIWeb search: off
To add speech-to-text to your app with an API, the basic flow is:
- Capture audio from the user
- Send the audio to a speech-to-text API
- Receive transcribed text
- Display or process the text in your app
Common implementation options
Option 1: Use a hosted speech-to-text API
Popular providers include:
- OpenAI
- Google Cloud Speech-to-Text
- Azure Speech
- AWS Transcribe
- AssemblyAI
- Deepgram
This is usually the fastest way to integrate.
Typical integration steps
1) Record audio in your app
You can capture:
- Microphone input for live dictation
- Uploaded audio files like MP3, WAV, M4A, etc.
In web apps, you might use:
MediaRecorderAPIgetUserMedia()for microphone access
In mobile apps, you’d use the platform’s audio recording APIs.
2) Send audio to the speech-to-text endpoint
Most APIs accept:
- Multipart form uploads
- Base64 audio
- Streaming audio for real-time transcription
For non-real-time transcription, sending a file is simplest.
3) Parse the response
The API will usually return:
- The transcribed text
- Sometimes timestamps
- Sometimes confidence scores
- Sometimes speaker labels
Example response:
{
"text": "Hello, this is a test transcription."
}
Example: using an HTTP API from a backend
JavaScript/Node example
import fs from "fs";
import fetch from "node-fetch";
async function transcribeAudio() {
const audioBuffer = fs.readFileSync("audio.mp3");
const formData = new FormData();
formData.append("file", new Blob([audioBuffer]), "audio.mp3");
formData.append("model", "whisper-1");
const response = await fetch("https://api.example.com/v1/transcriptions", {
method: "POST",
headers: {
Authorization: `Bearer YOUR_API_KEY`,
},
body: formData,
});
const result = await response.json();
console.log(result.text);
}
If you want real-time speech-to-text
Use one of these approaches:
- WebSockets
- Streaming HTTP
- Realtime SDKs
This is better for:
- Live captions
- Voice assistants
- Dictation apps
Best practices
- Do not expose API keys in frontend code
- Send audio through your backend if the provider requires a secret key
- Validate file type and size
- Consider privacy and compliance if you store audio
- Add error handling for:
- No microphone permission
- Unsupported audio format
- Network failures
- Long audio files
If you’re using OpenAI specifically
You typically:
- Record or collect an audio file
- Upload it to the transcription endpoint
- Read the returned transcript
If you want, I can give you:
- A web app example
- A Python example
- A Node.js example
- A real-time streaming example
Just tell me your stack.