Client-VAD speech segments
When and how to use synchronous segment transcription.
Client-VAD speech segments
Use Segment Transcription when your application already knows where speech starts and ends.
When to use Segment
| Use Segment | Use something else |
|---|---|
| You run VAD locally and upload one clip per utterance | Continuous mic stream → Realtime WS |
| You need a synchronous transcript in ~1s | Long file at a URL → Batch jobs |
OpenAI Whisper file upload pattern | Server should detect silence → Realtime WS |
Contract
- You segment — only POST clips that contain speech you want transcribed.
- We do not filter — silence, near-empty WAVs, and noise are transcribed as-is.
- You pay for what you send — wasting quota on silence is the caller's responsibility.
- No server HTTP retry — up to 2 worker attempts (800ms each) per request; on 502 the client may resend the same clip.
Typical flow
Client VAD detects utterance end
→ encode clip (e.g. WAV)
→ POST /v1/audio/transcriptions (multipart)
→ 200 + text (or 502 → client decides to retry)
Integrators like Flow follow this pattern: WebSocket voice session → local VAD → HTTP segment per utterance.
Example
curl -X POST https://audio.lansonai.com/v1/audio/transcriptions \
-H "Authorization: Bearer sk-..." \
-F "file=@utterance.wav" \
-F "language=zh"
Anti-patterns
- Uploading a full meeting recording to Segment — use batch
audio_urlinstead. - Streaming continuous PCM to Segment — use Realtime WS.
- Expecting the API to skip silence — it will not.
Related
Reconnect a Live Session
Network disconnect, retry, and resume design.
API Reference Overview
Base URL, REST/WebSocket endpoints, versioning.
Resources
