Client-VAD speech segments

When and how to use synchronous segment transcription.

Client-VAD speech segments

Use Segment Transcription when your application already knows where speech starts and ends.

When to use Segment

Use SegmentUse something else
You run VAD locally and upload one clip per utteranceContinuous mic stream → Realtime WS
You need a synchronous transcript in ~1sLong file at a URL → Batch jobs
OpenAI Whisper file upload patternServer should detect silence → Realtime WS

Contract

  1. You segment — only POST clips that contain speech you want transcribed.
  2. We do not filter — silence, near-empty WAVs, and noise are transcribed as-is.
  3. You pay for what you send — wasting quota on silence is the caller's responsibility.
  4. No server HTTP retry — up to 2 worker attempts (800ms each) per request; on 502 the client may resend the same clip.

Typical flow

Client VAD detects utterance end
  → encode clip (e.g. WAV)
  → POST /v1/audio/transcriptions (multipart)
  → 200 + text (or 502 → client decides to retry)

Integrators like Flow follow this pattern: WebSocket voice session → local VAD → HTTP segment per utterance.

Example

curl -X POST https://audio.lansonai.com/v1/audio/transcriptions \
  -H "Authorization: Bearer sk-..." \
  -F "file=@utterance.wav" \
  -F "language=zh"

Anti-patterns

  • Uploading a full meeting recording to Segment — use batch audio_url instead.
  • Streaming continuous PCM to Segment — use Realtime WS.
  • Expecting the API to skip silence — it will not.