Prompt

How do I build speech-to-text into my app with an API?

Artificial Intelligence / AI Platforms1 observationLast seen Oct 8, 2026

Latest observation

Oct 8, 2026OpenAI APIWeb search: off

To add speech-to-text (STT) to your app with an API, you generally follow this flow:

  1. Capture audio in your app
  2. Send the audio to a speech-to-text API
  3. Receive transcribed text
  4. Display or process the text in your app

Typical architecture

1) Record audio on the client

Use your app’s platform tools to record audio:

  • Web: MediaRecorder
  • iOS: AVAudioRecorder
  • Android: AudioRecord or MediaRecorder
  • React Native / Flutter: platform audio recording plugins

Save or stream the audio in a format the API accepts, such as:

  • mp3
  • wav
  • m4a
  • webm
  • ogg

2) Send audio to your backend or directly to the API

There are two common patterns:

Option A: Client → Your backend → STT API

Best if you want to:

  • hide API keys
  • add authentication
  • log or moderate requests
  • process or store audio securely

Option B: Client → STT API directly

Best for prototypes, but only if the API supports secure client-side access. Usually you still avoid exposing secret keys in the browser or mobile app.

3) Call the speech-to-text endpoint

Most STT APIs accept either:

  • a file upload
  • a streaming audio connection
  • a pre-signed URL to audio storage

A simple file-upload request often looks like:

curl https://api.example.com/v1/transcriptions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -F "file=@audio.mp3" \
  -F "language=en"

The response typically includes:

{
  "text": "Hello, this is a test transcription."
}

4) Show the transcript in your app

Use the returned text to:

  • display captions
  • store meeting notes
  • trigger search/indexing
  • feed an AI assistant
  • analyze sentiment or intent

Example backend flow

Node.js example

import express from "express";
import multer from "multer";
import fetch from "node-fetch";
import FormData from "form-data";

const app = express();
const upload = multer({ dest: "uploads/" });

app.post("/transcribe", upload.single("audio"), async (req, res) => {
  try {
    const form = new FormData();
    form.append("file", fs.createReadStream(req.file.path));
    form.append("language", "en");

    const response = await fetch("https://api.example.com/v1/transcriptions", {
      method: "POST",
      headers: {
        Authorization: `Bearer ${process.env.STT_API_KEY}`,
        ...form.getHeaders(),
      },
      body: form,
    });

    const data = await response.json();
    res.json(data);
  } catch (err) {
    res.status(500).json({ error: "Transcription failed" });
  }
});

Real-time speech-to-text

If you want live captions while someone is speaking, use a streaming STT API. These usually work over:

  • WebSockets
  • gRPC
  • WebRTC
  • vendor-specific streaming endpoints

This is better for:

  • live captions
  • voice assistants
  • call transcription
  • dictation apps

In streaming mode, you:

  1. open a live audio connection
  2. send small audio chunks continuously
  3. receive partial transcripts in real time

Key features to look for in an STT API

When choosing an API, check for:

  • Accuracy
  • Streaming support
  • Speaker diarization (who said what)
  • Language support
  • Punctuation and casing
  • Word timestamps
  • Noise robustness
  • Custom vocabulary
  • Cost per minute
  • Latency
  • Privacy / compliance

Common providers

Some popular speech-to-text API providers include:

  • OpenAI
  • Google Cloud Speech-to-Text
  • AWS Transcribe
  • Azure Speech
  • Deepgram
  • AssemblyAI

Best practices

  • Compress audio if possible to reduce upload size
  • Use 16 kHz mono for speech when supported
  • Handle long recordings by chunking or streaming
  • Send clear error messages for unsupported formats
  • Protect API keys on the server
  • Consider privacy laws if processing user voice data

Minimal implementation plan

If you want the fastest path:

  1. Add audio recording to your app
  2. Upload the audio to your server
  3. Forward it to an STT API
  4. Return the transcript to the client
  5. Display it in your UI

If you want, I can also give you:

  • a web app example
  • a Node.js backend example
  • a Python backend example
  • or a real-time streaming example

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.