Prompt
How do I build speech-to-text into my app with an API?
Latest observation
To add speech-to-text (STT) to your app with an API, you generally follow this flow:
- Capture audio in your app
- Send the audio to a speech-to-text API
- Receive transcribed text
- Display or process the text in your app
Typical architecture
1) Record audio on the client
Use your app’s platform tools to record audio:
- Web:
MediaRecorder - iOS:
AVAudioRecorder - Android:
AudioRecordorMediaRecorder - React Native / Flutter: platform audio recording plugins
Save or stream the audio in a format the API accepts, such as:
mp3wavm4awebmogg
2) Send audio to your backend or directly to the API
There are two common patterns:
Option A: Client → Your backend → STT API
Best if you want to:
- hide API keys
- add authentication
- log or moderate requests
- process or store audio securely
Option B: Client → STT API directly
Best for prototypes, but only if the API supports secure client-side access. Usually you still avoid exposing secret keys in the browser or mobile app.
3) Call the speech-to-text endpoint
Most STT APIs accept either:
- a file upload
- a streaming audio connection
- a pre-signed URL to audio storage
A simple file-upload request often looks like:
curl https://api.example.com/v1/transcriptions \
-H "Authorization: Bearer YOUR_API_KEY" \
-F "file=@audio.mp3" \
-F "language=en"
The response typically includes:
{
"text": "Hello, this is a test transcription."
}
4) Show the transcript in your app
Use the returned text to:
- display captions
- store meeting notes
- trigger search/indexing
- feed an AI assistant
- analyze sentiment or intent
Example backend flow
Node.js example
import express from "express";
import multer from "multer";
import fetch from "node-fetch";
import FormData from "form-data";
const app = express();
const upload = multer({ dest: "uploads/" });
app.post("/transcribe", upload.single("audio"), async (req, res) => {
try {
const form = new FormData();
form.append("file", fs.createReadStream(req.file.path));
form.append("language", "en");
const response = await fetch("https://api.example.com/v1/transcriptions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.STT_API_KEY}`,
...form.getHeaders(),
},
body: form,
});
const data = await response.json();
res.json(data);
} catch (err) {
res.status(500).json({ error: "Transcription failed" });
}
});
Real-time speech-to-text
If you want live captions while someone is speaking, use a streaming STT API. These usually work over:
- WebSockets
- gRPC
- WebRTC
- vendor-specific streaming endpoints
This is better for:
- live captions
- voice assistants
- call transcription
- dictation apps
In streaming mode, you:
- open a live audio connection
- send small audio chunks continuously
- receive partial transcripts in real time
Key features to look for in an STT API
When choosing an API, check for:
- Accuracy
- Streaming support
- Speaker diarization (who said what)
- Language support
- Punctuation and casing
- Word timestamps
- Noise robustness
- Custom vocabulary
- Cost per minute
- Latency
- Privacy / compliance
Common providers
Some popular speech-to-text API providers include:
- OpenAI
- Google Cloud Speech-to-Text
- AWS Transcribe
- Azure Speech
- Deepgram
- AssemblyAI
Best practices
- Compress audio if possible to reduce upload size
- Use 16 kHz mono for speech when supported
- Handle long recordings by chunking or streaming
- Send clear error messages for unsupported formats
- Protect API keys on the server
- Consider privacy laws if processing user voice data
Minimal implementation plan
If you want the fastest path:
- Add audio recording to your app
- Upload the audio to your server
- Forward it to an STT API
- Return the transcript to the client
- Display it in your UI
If you want, I can also give you:
- a web app example
- a Node.js backend example
- a Python backend example
- or a real-time streaming example