Segment Transcription API

Synchronous client-VAD speech segment transcription (multipart).

POST /v1/audio/transcriptions

What it does

Transcribe a single speech clip synchronously. This endpoint is OpenAI Whisper-compatible (multipart/form-data + file).

Client responsibility: you must segment speech with your own VAD before calling. Do not send continuous silence — the service does not filter silence, empty clips, or low-energy audio. Whatever you POST is transcribed and billed.

Typical integrators: voice clients (e.g. Flow) that detect utterance boundaries locally and upload each clip as WAV.

Endpoint

POST /v1/audio/transcriptions

Authentication

Authorization: Bearer sk-...

Request

Content-Type: multipart/form-data

FieldTypeRequiredDefaultDescription
filefileyes—Speech clip (WAV, MP3, etc.)
modelstringnowhisper-1Model name passed to workers
languagestringno—ISO-639-1 hint (zh, en, …)
promptstringno—Context prompt for STT
response_formatstringnoverbose_jsonverbose_json, json, or text
worker_idstringno—Pin to a specific worker ID
Not for long files or raw recordings Use Batch Jobs (POST /v1/audio/transcriptions/jobs) for audio_url async processing. Use Realtime WS when the server should run VAD on a PCM stream.

Example

curl -X POST https://audio.lansonai.com/v1/audio/transcriptions \
  -H "Authorization: Bearer sk-..." \
  -F "file=@utterance.wav" \
  -F "language=zh" \
  -F "response_format=verbose_json"

Response — 200 OK

verbose_json (default):

{
  "text": "你好世界",
  "language": "zh",
  "duration": 1.2,
  "segments": [],
  "words": []
}

Response headers:

HeaderDescription
X-Worker-IdWorker that produced the result
X-Fallback-Used1 if a fallback worker was used, else 0

Latency and retries

RuleValue
Per-attempt timeout800 ms (Modal + ElevenLabs)
Max POST attempts per request2 (primary warm Modal → ElevenLabs when applicable)
Server HTTP retryNo — failed requests return an error; client may resend
Cold ModalNot POSTed on this request; health probe may run async; ElevenLabs used for cold path

Wall-clock budget is roughly ≤1.6 s on the warm Modal + fallback path.

Errors

StatusCodeDescription
400invalid_content_typeBody is not multipart/form-data
400missing_fileNo file field
401—Authentication failed
502transcription_failedAll worker attempts failed
{
  "error": {
    "message": "Realtime transcription timed out after 800ms",
    "type": "invalid_request_error",
    "code": "transcription_failed"
  }
}