# LansonAI Developer Platform LansonAI is the **Voice Context Layer for live speech**. It turns continuously changing speech into context that applications can read, display, translate, and act on while the conversation is still happening. Traditional speech recognition systems are primarily designed to answer: > What words were spoken? Live applications have a harder problem: > What can I safely show, understand, translate, or act on right now? LansonAI is built around that second problem. ## What you can build Use LansonAI to build experiences such as: - live captions - real-time multilingual experiences - voice interfaces - accessibility tools - live media experiences - conversational applications - systems that need structured context from ongoing speech ## Built for live speech Live speech is not a sequence of finished sentences. Recognition results evolve as more audio arrives. Words may be revised, sentence boundaries may move, and meaning may become clearer several seconds after the first tokens appear. LansonAI handles this changing state between raw audio and the application consuming it. Instead of treating every intermediate recognition result as equally reliable, the platform is designed around the lifecycle of spoken context. ## Core concepts Before integrating the API, we recommend understanding three concepts: **Voice Context Layer** The processing layer between raw speech recognition and the application consuming live speech. **StableStream** LansonAI's approach to producing continuously readable text while underlying recognition and context continue to evolve. **Trinity Engine** The shared processing foundation behind LansonAI's real-time voice capabilities. → Start with **[Voice Context Layer](https://docs.lansonai.com/concepts)** # Quickstart This page gets you from zero to your first transcription request. ## 0. Use an LLM to get started (recommended) The fastest way is to let an AI assistant read the docs for you. Paste this into ChatGPT, Claude, or your coding agent: ```text Read https://docs.lansonai.com/llms.txt and guide me to understand or integrate LansonAI's speech transcription based on my goal. ``` The docs are also available as full-page markdown for agents: [llms-full.txt](https://docs.lansonai.com/llms-full.txt){rel=""nofollow""} and [OpenAPI spec](https://audio.lansonai.com/openapi.json){rel=""nofollow""}. ## 1. Try it before you code - **See it live**: open this [replay link](https://live.lansonai.com/share/rethinking-live-captioning-accuracy-latency-and--9d7d72ec){rel=""nofollow""} — you will hear the audio and watch the realtime captions appear in sync, exactly what your users will experience. - **Need a sample file?** Use our [podcast episode](https://r2.lansonai.com/lanson/Readable%20Captions.m4a){rel=""nofollow""} (\~18 minutes, 34 MB) as `audio_url` for batch transcription, or [download it](https://r2.lansonai.com/lanson/Readable%20Captions.m4a){rel=""nofollow""} for realtime testing. ## 2. Get an API key Your API key starts with `sk-` and is provisioned by LansonAI. If you do not have one yet, contact the LansonAI team. ::callout{color="amber" icon="i-lucide-triangle-alert"} The plaintext key is returned only once at provisioning time. The server stores only a SHA-256 hash. Save it immediately. :: ## 3. Batch transcription from audio URL Submit a public URL, get a workflow id (`cf_…`), poll for the result: ```bash BASE="https://audio.lansonai.com" SK="sk-..." curl -X POST "$BASE/v1/audio/transcriptions/jobs" \ -H "Authorization: Bearer $SK" \ -H "Content-Type: application/json" \ -d '{ "audio_url": "https://cdn.example.com/audio/meeting.wav" }' ``` Response (202 Accepted): ```json { "request_id": "...", "workflow_id": "cf_...", "status": "queued", "poll_endpoint": "GET /v1/audio/transcriptions/jobs/cf_..." } ``` Poll: ```bash curl "$BASE/v1/audio/transcriptions/jobs/" \ -H "Authorization: Bearer $SK" ``` For **client-VAD speech clips** (multipart, synchronous 200) see [Segment API](https://docs.lansonai.com/api-reference/segment-transcription). ## 4. Realtime transcription (WebSocket) For live speech, connect over WebSocket: ```bash # 1) (browser only) get a 60-second session token RT=$(curl -s -X POST "$BASE/v1/audio/transcriptions/session-token" \ -H "Authorization: Bearer $SK" | jq -r .token) # 2) Connect # wss://audio.lansonai.com/v1/audio/transcriptions/stream?access_token=$RT # 3) Send PCM16LE / 16kHz / mono audio frames # Text frame: {"type":"input_audio_buffer.append","audio":""} # Or send raw binary PCM frames # 4) Flush when a speech segment ends # {"type":"input_audio_buffer.flush"} # 5) Receive transcription events # {"type":"conversation.item.input_audio_transcription.completed","text":"..."} ``` ## 5. Expected output ### Offline (complete status) ```json { "workflow_id": "...", "status": "complete", "output": { "result": { "segments": [ { "id": 0, "start_time": 0.0, "end_time": 3.2, "text": "The weather is nice today", "confidence": 0.95 } ], "summary": { "total_duration": 120.5, "num_segments": 45 }, "metadata": { "language": "zh", "model": "whisper-large-v3-turbo", "chunk_count": 3 } } } } ``` ### Realtime event ```json { "type": "conversation.item.input_audio_transcription.completed", "utterance_index": 0, "text": "The weather is nice today", "language": "zh", "audio_duration_ms": 3200, "latency_ms": 480 } ``` ## Next steps - [Realtime Quickstart](https://docs.lansonai.com/realtime/quickstart) — full WebSocket examples - [Transcribe Audio](https://docs.lansonai.com/recorded/transcribe-audio) — offline transcription details - [Authentication](https://docs.lansonai.com/start/authentication) — auth methods and browser security # Authentication All API requests require an API key. This page explains the key format, usage patterns, and browser-side security model. ## API key format - Prefix: `sk-` - The plaintext key is returned only once at provisioning. The server stores only a SHA-256 hash. - Minimum length: 16 characters (at least 13 characters after `sk-`) ## Server-side usage All endpoints accept the key via the `Authorization: Bearer` header: ```bash curl -H "Authorization: Bearer sk-..." https://audio.lansonai.com/v1/audio/transcriptions ``` This works for both offline HTTP requests and WebSocket upgrade handshakes. ## Browser-side security ::callout{color="red" icon="i-lucide-circle-alert"} Never expose your long-lived API key in frontend code or WebSocket URLs. Use a short-lived session token for browser connections. :: ### Session token flow Browsers cannot safely store a long-lived key. LansonAI provides a short-lived token mechanism: ```text 1. Frontend → your backend: request a session token 2. Your backend → LansonAI: POST /v1/audio/transcriptions/session-token Authorization: Bearer sk-... (your long-lived key, server-side only) 3. LansonAI → your backend: returns rt_... token (valid 60 seconds) 4. Your backend → frontend: pass the rt_... token 5. Frontend → LansonAI: open WebSocket ?access_token=rt_... ``` ### Session token details | Property | Value | | -------------- | --------------------------------------------------- | | Prefix | `rt_` | | Validity | 60 seconds | | Signature | HMAC-SHA256 | | Signing secret | `GATEWAY_SESSION_TOKEN_SECRET` (server-side config) | | Transport | URL query `?access_token=rt_...` or `?token=rt_...` | ### Token response ```json { "token": "rt_eyJ...", "expires_in": 60, "endpoint": "/v1/audio/transcriptions/stream" } ``` If `GATEWAY_SESSION_TOKEN_SECRET` is not configured, the endpoint returns `503 token_issuer_unavailable`. ## WebSocket authentication methods The realtime WebSocket accepts three authentication methods: | Method | Use case | Usage | | ------------------------------ | ------------------- | --------------------------- | | `Authorization: Bearer sk-...` | Server-side clients | HTTP upgrade header | | `?access_token=rt_...` | Browser | URL query parameter | | `?token=rt_...` | Browser | URL query parameter (alias) | ::callout{color="amber" icon="i-lucide-triangle-alert"} Never put `sk-...` in a URL. Browser connections must always use `rt_...` session tokens. :: ## Admin API Key provisioning, revocation, and plan-catalog endpoints exist under `/v1/external-transcription`, but they are for internal operators and dashboards, not regular API consumers. See [Admin API](https://docs.lansonai.com/api-reference/admin) for details. ## Error handling | Error code | HTTP | Cause | | -------------------------- | ---- | -------------------------------------- | | `missing_api_key` | 401 | No Authorization header or query token | | `invalid_api_key` | 401 | Key format invalid or unknown key | | `revoked_api_key` | 401 | Key has been revoked | | `expired_api_key` | 401 | Key has expired | | `auth_unavailable` | 503 | Key store unavailable | | `invalid_session_token` | 401 | Session token invalid or expired | | `token_issuer_unavailable` | 503 | Token issuer not configured | See [Errors](https://docs.lansonai.com/api-reference/errors) for the full error reference. ## Next steps - [Quickstart](https://docs.lansonai.com/start/quickstart) — first request - [Errors](https://docs.lansonai.com/api-reference/errors) — full error reference - [Realtime API](https://docs.lansonai.com/api-reference/realtime-api) — WebSocket endpoint details # Choose an API LansonAI provides three live transcription APIs today, plus Voice Agent which is coming soon. ## Decision table | Scenario | API | Endpoint | Who does VAD | | ------------------------------------------ | ------------ | ------------------------------------ | ----------------------- | | Live captions / dialogue | Realtime WS | `WS /v1/audio/transcriptions/stream` | **Server** | | Pre-segmented speech clips (e.g. voice UI) | Segment HTTP | `POST /v1/audio/transcriptions` | **Client** | | Meeting / podcast file from URL | Batch jobs | `POST /v1/audio/transcriptions/jobs` | **Server** (time slice) | | Voice agent | Voice Agent | Coming soon | — | ::callout{color="primary" icon="i-lucide-info"} Segment vs Batch **Segment** = you already cut speech with client VAD; send multipart `file`, get **200** immediately. **Batch** = you have a long `audio_url`; get **202** + poll. Do not send silence to Segment — we transcribe whatever you POST. :: ## Realtime Voice Context **When to use**: text while the person is still speaking. - Live captions, real-time display - Continuous PCM stream; server-side VAD and utterance boundaries **Characteristics**: - WebSocket full-duplex - PCM16LE / 16kHz / mono - Millisecond-level latency → [Realtime Overview](https://docs.lansonai.com/realtime) ## Segment transcription (client VAD) **When to use**: your app already detected speech boundaries (VAD) and has a short clip per utterance. - Voice clients that upload one WAV per utterance - OpenAI Whisper-compatible `multipart/form-data` - Synchronous **200** response **Characteristics**: - **No server-side silence filtering** — silent clips are transcribed and billed - \~800ms × up to 2 worker attempts per request; **no HTTP auto-retry** - On failure (**502**), client decides whether to resend → [Segment API](https://docs.lansonai.com/api-reference/segment-transcription) · [Client-VAD guide](https://docs.lansonai.com/guides/client-vad-segments) ## Batch jobs (recorded files) **When to use**: complete audio file at a public URL; processing can take minutes. - Post-meeting / podcast processing - Server slices by time, concurrent chunk transcription, aggregation - Optional LLM review and webhook **Characteristics**: - HTTP async (**202** + poll) - `workflow_id` is `cf_` + 64 hex (not UUID) → [Recorded Overview](https://docs.lansonai.com/recorded) · [Batch Jobs API](https://docs.lansonai.com/api-reference/batch-jobs) ## Voice Agent **Status**: Coming soon ## Comparison | Dimension | Realtime WS | Segment HTTP | Batch jobs | | ------------ | -------------- | ---------------- | -------------------- | | Transport | WebSocket | HTTP POST | HTTP POST + GET poll | | Latency | Milliseconds | \~≤1.6s sync | Minutes (async) | | Input | PCM16LE stream | Multipart `file` | JSON `audio_url` | | Output | Event stream | Transcript JSON | Full job JSON | | VAD | Server | **Client** | Server (slice) | | Success code | WS events | **200** | **202** | ## Next steps - [Realtime Quickstart](https://docs.lansonai.com/realtime/quickstart) - [Client-VAD segments](https://docs.lansonai.com/guides/client-vad-segments) - [Transcribe Audio](https://docs.lansonai.com/recorded/transcribe-audio) - [Authentication](https://docs.lansonai.com/start/authentication) # Realtime Overview Realtime is a **Voice Context Layer** mode for live speech. [Recorded](https://docs.lansonai.com/recorded) is the other mode, for full-file processing. See [Voice Context Layer](https://docs.lansonai.com/concepts) for the umbrella model. The Realtime API provides streaming speech transcription over WebSocket. You send PCM16LE audio frames, the server returns transcription events. ## What it does - **Live captions**: text appears while the person is still speaking - **Real-time translation**: source and translated text together *(coming soon for external sessions)* - **Multi-language**: Whisper-compatible language support - **Server-side VAD**: automatic speech segment detection, no client-side splitting needed ## Input - Audio format: **PCM16LE / 16kHz / mono** - Transport: JSON text frames (base64) or binary frames (raw PCM) - Frame size: recommended 100ms frames (\~3200 bytes), max 1 MiB ## Output Server pushes JSON events. The primary output is: - `conversation.item.input_audio_transcription.completed` — a speech segment has been transcribed Supporting events: - `input_audio_buffer.speech_started` / `speech_stopped` — speech segment boundaries - `session.created` — connection established - `error` — error events ## Architecture ```text Your app LansonAI Gateway Upstream STT │ │ │ ├── WS connect (Bearer) ───→ session.created │ │ │ │ ├── Send PCM16 frames ──────→ gateway relay (binary PCM) ─────→ VAD + STT │ │ │ │←── input_audio_buffer.speech_started ──────│←── input_audio_buffer.speech_started │ │ │ │←── conversation.item.input_audio_transcription.completed ─│←── conversation.item.input_audio_transcription.completed │ │ │ ├── Close ─────────────────→ meter flush → R2 ledger │ ``` ## vs. Segment and Batch | Dimension | Realtime WS | Segment HTTP | Batch jobs | | --------- | -------------- | ------------------ | ------------------ | | Transport | WebSocket | HTTP sync | HTTP async | | Latency | Milliseconds | \~≤1.6s | Minutes | | Input | PCM16LE stream | Multipart `file` | Audio URL | | Output | Event stream | Transcript JSON | Complete job JSON | | VAD | **Server** | **Client** | **Server** (slice) | | Best for | Live stream | Pre-cut utterances | Long files | ## Next steps - [Realtime Quickstart](https://docs.lansonai.com/realtime/quickstart) — runnable example - [Connection Lifecycle](https://docs.lansonai.com/realtime/connection-lifecycle) — connect, timeout, close - [Audio Input](https://docs.lansonai.com/realtime/audio-input) — audio format reference # Realtime Quickstart Minimal runnable examples for real-time streaming transcription. ## Node.js / TypeScript ```typescript import WebSocket from "ws"; const BASE = "wss://audio.lansonai.com"; const SK = "sk-..."; const ws = new WebSocket(`${BASE}/v1/audio/transcriptions/stream`, { headers: { Authorization: `Bearer ${SK}` }, }); ws.on("open", () => { // Optional: configure session ws.send(JSON.stringify({ type: "session.update", language: "zh" })); // Send PCM16LE base64 audio frames (100ms recommended) ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: "", })); // Flush when a speech segment ends ws.send(JSON.stringify({ type: "input_audio_buffer.flush" })); }); ws.on("message", (data) => { const event = JSON.parse(data.toString()); switch (event.type) { case "session.created": console.log("Session:", event.session_id, "Plan:", event.plan); break; case "conversation.item.input_audio_transcription.completed": console.log(`[${event.utterance_index}] ${event.text}`); break; case "error": console.error("Error:", event.code, event.message); break; } }); ws.on("close", (code) => console.log(`Closed: ${code}`)); ``` ## Browser JavaScript Browser connections require a session token obtained from your backend: ```javascript // 1. Get session token from your backend const { token } = await fetch("/your-backend/session-token", { method: "POST", }).then(r => r.json()); // 2. Open WebSocket const ws = new WebSocket( `wss://audio.lansonai.com/v1/audio/transcriptions/stream?access_token=${token}` ); ws.onopen = () => { // 3. Capture microphone audio at 16kHz mono const ctx = new AudioContext({ sampleRate: 16000 }); navigator.mediaDevices.getUserMedia({ audio: true }).then((stream) => { const source = ctx.createMediaStreamSource(stream); // Use AudioWorklet to convert to PCM16LE and send frames }); }; ws.onmessage = (event) => { const data = JSON.parse(event.data); if (data.type === "conversation.item.input_audio_transcription.completed") { console.log(data.text); } }; ``` ## Python ```python import json, websocket ws = websocket.create_connection( "wss://audio.lansonai.com/v1/audio/transcriptions/stream", header=["Authorization: Bearer sk-..."] ) print(json.loads(ws.recv())) # session.created ws.send(json.dumps({ "type": "input_audio_buffer.append", "audio": "" })) ws.send(json.dumps({"type": "input_audio_buffer.flush"})) while True: event = json.loads(ws.recv()) if event["type"] == "conversation.item.input_audio_transcription.completed": print(f"[{event['utterance_index']}] {event['text']}") elif event["type"] == "error": print(f"Error: {event['code']}: {event['message']}") break ``` ## Reference script The repository's `scripts/realtime-inspect.ts` is a complete end-to-end validation script: ```bash bun run inspect:realtime -- --audio ./sample.wav ``` It handles session-token flow, PCM16LE encoding, 100ms frame simulation, and completed-event verification. Requires `ffmpeg`. ## Next steps - [Connection Lifecycle](https://docs.lansonai.com/realtime/connection-lifecycle) — timeouts and close codes - [Audio Input](https://docs.lansonai.com/realtime/audio-input) — format and frame size - [Client Messages](https://docs.lansonai.com/api-reference/client-messages) — all message types # Connection Lifecycle The full lifecycle of a realtime WebSocket connection. ## Stages ```text 1. Connect → 2. session.created → 3. Configure (optional) → 4. Stream audio ↓ 6. Close/Reconnect ← 5. Receive events ← ``` ## 1. Connect WebSocket handshake: - Server-side: `Authorization: Bearer sk-...` header - Browser: `?access_token=rt_...` query parameter Conditions checked on connect: - WebSocket upgrade required (else 426) - Gateway enabled (else 503 `gateway_disabled`) - Circuit breaker closed (else 503 `upstream_unavailable` + `Retry-After`) - Connection quota passes (concurrent + rate) ## 2. session.created Received immediately on connect: ```json { "type": "session.created", "session_id": "sess_...", "limits": { "max_concurrent_utterances": 2, "max_session_seconds": 900, "idle_timeout_seconds": 60, "remaining_audio_seconds": 3600 } } ``` ## 3. Configure (optional) Send `session.update` to set parameters: ```json { "type": "session.update", "language": "zh" } ``` Receive `session.updated` confirmation. ## 4. Stream audio Continuously send PCM16LE audio frames: - Text frames: `input_audio_buffer.append` (base64) - Binary frames: raw PCM data - End of speech segment: `input_audio_buffer.flush` ## 5. Receive events ```text input_audio_buffer.speech_started → input_audio_buffer.speech_stopped → conversation.item.input_audio_transcription.completed ``` This cycle repeats for each speech segment. ## 6. Close / Reconnect ### Normal close Client calls `ws.close()`. Server flushes final meter data. ### Timeout close | Close code | Cause / close reason | Trigger | | ---------- | ------------------------------: | ------------------------------------------------- | | 4408 | `idle_timeout` | No audio received within `idle_timeout_seconds` | | 1008 | `session_duration_limit` | Exceeded `max_session_seconds` | | 1008 | `session_audio_limit` | Exceeded `max_audio_seconds_per_session` | | 1008 | `audio_quota_exhausted` | Billing-window audio budget exhausted | | 1009 | `frame_too_large` | Single frame exceeds 1 MiB | | 1011 | `client_error` / internal error | Client socket error or unexpected gateway failure | | 1013 | `upstream_unavailable` | Circuit breaker open | | 1013 | `upstream_backpressure` | Upstream is not draining fast enough | | 1013 | `client_backpressure` | Client is not reading events fast enough | ## Idle timeout To avoid idle timeout, either: - Send silent frames to keep the connection alive - Or close explicitly and reconnect when needed ## Recommended reconnect design 1. Listen for `close` events, decide reconnect based on close code 2. 4408 (idle) and 1008 (duration limit): reconnect directly 3. 1013 (upstream): exponential backoff before reconnect 4. 1000 (normal): do not reconnect 5. After reconnect, re-send `session.update` to restore configuration See [Reconnect a Live Session](https://docs.lansonai.com/guides/reconnection) for a detailed guide. ## Related - [Realtime API](https://docs.lansonai.com/api-reference/realtime-api) — endpoint reference - [Reconnection & Retries](https://docs.lansonai.com/production/reconnection-retries) — retry strategies - [Rate Limits](https://docs.lansonai.com/api-reference/rate-limits) — plan timeout parameters # Audio Input The realtime API accepts PCM16LE audio. This page covers format requirements, frame size, and latency tradeoffs. ## Format requirements | Property | Requirement | | ----------- | ------------------------------------- | | Encoding | PCM16LE (16-bit signed little-endian) | | Sample rate | 16000 Hz | | Channels | 1 (mono) | ::callout{color="amber" icon="i-lucide-triangle-alert"} Audio not matching this format will be rejected or produce incorrect results. Convert with ffmpeg first. :: ## Converting audio ```bash ffmpeg -i input.mp3 -ar 16000 -ac 1 -c:a pcm_s16le output.wav ``` ## Transport methods ### Text frames (base64) ```json { "type": "input_audio_buffer.append", "audio": "" } ``` Has \~33% base64 overhead. Good for debugging and quick integration. ### Binary frames (raw PCM) Send raw `ArrayBuffer` directly. No encoding overhead. Recommended for production. ## Frame size and latency | Frame size | Interval | Data size | Latency impact | | ---------- | --------- | -------------- | ----------------------------------- | | 50ms | 50ms | 1600 bytes | Lowest latency, higher CPU overhead | | **100ms** | **100ms** | **3200 bytes** | **Recommended balance** | | 200ms | 200ms | 6400 bytes | Increased latency, fewer frames | | 500ms | 500ms | 16000 bytes | Noticeable latency | **100ms frames** (3200 bytes) is the recommended sweet spot. ## Max frame size 1 MiB (1,048,576 bytes). Exceeding this returns `audio_frame_too_large` and closes the connection (1009). ## Browser capture ```javascript const ctx = new AudioContext({ sampleRate: 16000 }); const stream = await navigator.mediaDevices.getUserMedia({ audio: true }); const source = ctx.createMediaStreamSource(stream); // Use AudioWorklet or ScriptProcessorNode to convert float samples to PCM16LE // Ensure 16000 Hz sample rate and mono ``` ## Server-side capture - Use `Authorization: Bearer sk-...` header - Send binary PCM frames directly (no base64 needed) - Keep sending to avoid idle timeout ## Backpressure When the upstream cannot keep up: | Condition | Behavior | | -------------------------------- | ------------------------------------ | | Soft threshold (1 MiB in-flight) | Drop frames, send `lanson.throttled` | | Hard threshold (8 MiB in-flight) | Close connection (1013) | Monitor `lanson.throttled` events and reduce send rate accordingly. ## Related - [Client Messages](https://docs.lansonai.com/api-reference/client-messages) — audio message format - [Connection Lifecycle](https://docs.lansonai.com/realtime/connection-lifecycle) — idle timeout - [Realtime Quickstart](https://docs.lansonai.com/realtime/quickstart) — complete example # Transcript Lifecycle State transitions for transcribed text in a realtime session. Understanding these states is essential for building caption UIs. ## State flow ```text audio input ↓ speech_started (VAD detects speech) ↓ [upstream processing — segment in progress] ↓ speech_stopped (VAD detects silence / flush) ↓ transcription.completed (stable text) ``` ## Event semantics ### speech\_started VAD detected the start of a speech segment. - **At this point**: no transcribed text yet - **Client action**: optionally show a "listening" indicator - **Mutable**: no, `utterance_index` is fixed ### speech\_stopped VAD detected the end of a speech segment (silence, flush, or max speech duration). - **At this point**: upstream begins final transcription - **Client action**: optionally show "processing" state - **Mutable**: no ### conversation.item.input\_audio\_transcription.completed A speech segment has been transcribed. This is the **primary output**. - **At this point**: `text` is stable and will not change - **Client action**: display the text, safe to write to database - **Mutable**: no, text is final - **Can trigger downstream**: yes ::callout{color="primary" icon="i-lucide-info"} **When is text stable?** The `text` in a `conversation.item.input_audio_transcription.completed` event is final. Once received, it will never be modified for that utterance. :: ## utterance\_index Each speech segment has a unique, incrementing `utterance_index`. Use it to: - Order transcribed results - Detect missing events (index gaps) - Merge into a complete transcript ## Concurrent utterances The upstream may process multiple speech segments simultaneously. `limits.max_concurrent_utterances` limits concurrency. When exceeded, frames are dropped and `lanson.throttled` is sent. ## No partial text In the current version: - **No partial transcription is pushed**: no intermediate text before `transcription.completed` - **One result per utterance**: each utterance produces one `completed` event - **Text is immediately stable**: no partial → final transition needed on the client This means clients do not need to handle caption jitter — each segment's text arrives in its final form. ## Related - [StableStream](https://docs.lansonai.com/concepts/stablestream) — the stability contract - [Build a Stable Caption UI](https://docs.lansonai.com/guides/stable-caption-ui) — UI implementation - [Server Events](https://docs.lansonai.com/api-reference/server-events) — event field reference # Realtime Translation Real-time translation: get translated text alongside transcription while speech is happening. ::callout{color="amber" icon="i-lucide-triangle-alert"} Real-time `translate_to` is **not currently enabled** for external sessions. The public realtime gateway forwards `language` but does not forward `translate_to`, so a `session.update` with that field has no effect. :: ## Usage ### Configure translation Set the source language and translation target via `session.update`: ```json { "type": "session.update", "language": "zh", "translate_to": "en" } ``` - `language`: source language of the audio - `translate_to`: target language for translation ### Translation output When enabled, translation results are returned alongside transcription in the `conversation.item.input_audio_transcription.completed` event. ## Source and translated text correspondence Source and translation for each utterance arrive in the same `conversation.item.input_audio_transcription.completed` event, ensuring correspondence. No client-side matching needed. ## Switching languages Change language at any time during the session: ```json { "type": "session.update", "language": "en" } ``` New utterances after the switch use the new language. In-flight utterances complete with the previous language. ## Latency Translation adds additional latency: | Component | Typical p50 | | ------------------- | ----------- | | STT (transcription) | \~310ms | | Translation | \~200ms | | End-to-end | \~510ms | Translation latency depends on audio length, context, and service load. Measured data on [Service Status](https://control.lansonai.com/status/lanson-audio){rel=""nofollow""}. ## Multilingual / code-switching - **Auto-detect**: set `language: "auto"` - **Mixed language**: upstream supports code-switching, accuracy depends on model and audio quality - **Chinese aliases**: `zh-cn`, `zh-tw`, `mandarin` etc. are normalized to `zh` ## UI strategy 1. Receive `conversation.item.input_audio_transcription.completed` event 2. Display source text on first line 3. Display translation on second line 4. Each utterance is independent, no replacement needed See [Build Live Translation](https://docs.lansonai.com/guides/live-translation) for a detailed guide. ## Related - [Session Configuration](https://docs.lansonai.com/realtime/session-configuration) — `session.update` details - [Live Translation](https://docs.lansonai.com/concepts/live-translation) — translation concepts - [Build Live Translation](https://docs.lansonai.com/guides/live-translation) — UI guide # Session Configuration Configure session parameters via `session.update` messages. ## When to configure - **After connect**: send `session.update` right after receiving `session.created` - **During session**: update parameters at any time - **Confirmation**: server returns `session.updated` to confirm ## All options ### Language and context | Field | Type | Description | | ---------- | ------ | ----------------------------------------------- | | `language` | string | Audio source language (e.g. `zh`, `en`, `auto`) | | `prompt` | string | Context hint, max 2000 characters | ### Model selection | Field | Snake-case alias | Description | | ----------------- | ------------------ | -------------------------- | | `backend` | — | Upstream backend selection | | `sttModel` | `stt_model` | STT model | | `normalizerModel` | `normalizer_model` | Text normalizer model | ### VAD configuration | Field | Snake-case alias | Type | Description | | ----------------------- | --------------------------- | ------- | ----------------------------------------- | | `vad` | — | boolean | Enable VAD | | `vadThreshold` | `vad_threshold` | number | Sensitivity threshold | | `vadSilenceMs` | `vad_silence_ms` | number | Silence duration to trigger end-of-speech | | `vadPrefixMs` | `vad_prefix_ms` | number | Prefix padding | | `vadMinSpeechMs` | `vad_min_speech_ms` | number | Minimum speech segment | | `vadTargetSpeechMs` | `vad_target_speech_ms` | number | Target speech segment | | `vadMaxSpeechMs` | `vad_max_speech_ms` | number | Maximum speech segment | | `vadSmartSplitWindowMs` | `vad_smart_split_window_ms` | number | Smart split window | ### Audio configuration Fixed by the gateway, client cannot change: ```json { "format": "pcm_s16le", "sample_rate": 16000, "channels": 1 } ``` ## Examples ### Set language ```json { "type": "session.update", "language": "zh" } ``` ### Set context hint ```json { "type": "session.update", "prompt": "Cardiology conference, terms: atrial fibrillation, anticoagulation, warfarin" } ``` ### Adjust VAD ```json { "type": "session.update", "vad": true, "vadThreshold": 0.5, "vadSilenceMs": 800 } ``` ### Switch language at runtime ```json { "type": "session.update", "language": "en" } ``` ## Tenant isolation All non-whitelisted fields are stripped before forwarding upstream. These fields are injected by the gateway and cannot be overridden: - `request_id` - `session_id` - `user_id` ## Related - [Client Messages](https://docs.lansonai.com/api-reference/client-messages) — message format - [Server Events](https://docs.lansonai.com/api-reference/server-events) — `session.updated` event - [Live Translation](https://docs.lansonai.com/concepts/live-translation) — translation concepts # Handling Interruptions & Silence How silence, pauses, and speech interruptions are handled in realtime transcription. ## VAD and speech segmentation Server-side VAD automatically detects speech segments. The WebSocket events pushed to your client are: 1. Speech detected → `input_audio_buffer.speech_started` (abbreviated as `speech_started`) 2. Speech continues → audio accumulated 3. Silence detected → `input_audio_buffer.speech_stopped` (abbreviated as `speech_stopped`) 4. Transcription complete → `conversation.item.input_audio_transcription.completed` (abbreviated as `completed`) Use the full `type` string in your event dispatcher. See [Server Events](https://docs.lansonai.com/api-reference/server-events) for the complete payload schemas. Clients do not need to implement VAD — just consume events. ## Silence behavior ### Short pauses Normal speaking pauses (commas, sentence ends) do not trigger segment splits. VAD has a silence threshold (`vadSilenceMs`); only silence exceeding this duration triggers end-of-speech. ### Long silence - Silence exceeding `idle_timeout_seconds` closes the connection with code `4408`. - The actual value is returned in `session.created.limits.idle_timeout_seconds`; do not hard-code it. - To keep the connection alive, continue sending PCM16LE audio frames, even when they contain silence (all-zero samples). For example, a 100 ms silent binary frame at 16 kHz mono is `new Int16Array(1600).fill(0)`. See [Audio Input](https://docs.lansonai.com/realtime/audio-input) for frame-size guidance and [Connection Lifecycle](https://docs.lansonai.com/realtime/connection-lifecycle) for close-code details. ## Manual flush Send `input_audio_buffer.flush` (or its alias `flush`) to force-end the current speech segment: ```json { "type": "input_audio_buffer.flush" } ``` Use cases: - You know a speech segment has ended (e.g. the user released a push-to-talk button or tapped an end-of-utterance button). - Force the upstream to process already-sent audio instead of waiting for VAD. - Reduce latency by not waiting for `vadSilenceMs`. ## Max speech segment duration `vadMaxSpeechMs` limits the maximum duration of a single speech segment. When exceeded: - The current segment is finalized: `input_audio_buffer.speech_stopped` → `conversation.item.input_audio_transcription.completed`. - A new segment starts immediately with `input_audio_buffer.speech_started`. - The `reason` field of `speech_stopped` explains why the segment ended (e.g. `end_of_speech` or `max_duration`). See [Server Events](https://docs.lansonai.com/api-reference/server-events) for the full list of `reason` values. ## Interruption handling ### Speaker interrupted In live scenarios, if a speaker is interrupted: - Upstream detects the speech boundary. - Current segment: `input_audio_buffer.speech_stopped` → `conversation.item.input_audio_transcription.completed`. - A new segment starts with `input_audio_buffer.speech_started`. ### Client strategy 1. Each `conversation.item.input_audio_transcription.completed` event is independent and final. 2. No need to cancel or roll back already-displayed text. 3. New segments are ordered by `utterance_index`. If two segments overlap in time, use `audio_duration_ms` / `speech_duration_ms` to position them on a timeline rather than relying on arrival order. ## Recommended client handling ```typescript ws.onmessage = (event) => { if (typeof event.data !== "string") return; const data = JSON.parse(event.data); if (data.type === "input_audio_buffer.speech_started") { showListeningIndicator(data.utterance_index); } else if (data.type === "input_audio_buffer.speech_stopped") { showProcessingIndicator(data.utterance_index); } else if (data.type === "conversation.item.input_audio_transcription.completed") { appendTranscript(data.utterance_index, data.text); hideIndicators(data.utterance_index); } }; ``` ## Related - [Transcript Lifecycle](https://docs.lansonai.com/realtime/transcript-lifecycle) — state transitions - [Session Configuration](https://docs.lansonai.com/realtime/session-configuration) — VAD parameters - [Connection Lifecycle](https://docs.lansonai.com/realtime/connection-lifecycle) — idle timeout and reconnect - [Client Messages](https://docs.lansonai.com/api-reference/client-messages) — `input_audio_buffer.flush` - [Server Events](https://docs.lansonai.com/api-reference/server-events) — event schemas and `reason` values # Recorded Overview Recorded is **batch transcription**: submit a public `audio_url`, process asynchronously, receive complete results. For **synchronous client-VAD clips** see [Segment Transcription](https://docs.lansonai.com/api-reference/segment-transcription). For **live streams** see [Realtime](https://docs.lansonai.com/realtime). ## How it works ```text Submit audio URL (POST /v1/audio/transcriptions/jobs) → Probe audio format and duration → Slice into time-based chunks → Concurrent transcription → Aggregate (merge into global timestamps) → Optional LLM review → Webhook callback (optional) ``` ## When to use | Scenario | Suitable | | -------------------------------- | ------------------------------------------------------------------------------ | | Post-meeting processing | ✅ | | Batch audio file transcription | ✅ | | Need precise timestamps | ✅ | | Need structured results | ✅ | | Need subtitle files | ✅ | | Short utterance after client VAD | ❌ use [Segment](https://docs.lansonai.com/api-reference/segment-transcription) | | Live microphone stream | ❌ use [Realtime](https://docs.lansonai.com/realtime) | ## vs. Realtime and Segment | Dimension | Batch jobs | Segment HTTP | Realtime | | ------------ | ----------------------- | ---------------- | ---------------- | | Transport | HTTP async | HTTP sync | WebSocket | | Input | Audio URL | Multipart `file` | PCM16LE stream | | Latency | Minutes | \~≤1.6s | Milliseconds | | Output | Complete JSON | Transcript JSON | Event stream | | Segmentation | **Server** (time slice) | **Client** (VAD) | **Server** (VAD) | | Success | **202** + poll | **200** | WS events | ## Current capabilities | Capability | Status | | ----------------------- | ----------------------------- | | Full-file transcription | ✅ | | Segment timestamps | ✅ | | Confidence scores | ✅ | | Translation | Coming soon | | Webhook callback | ✅ | | LLM review correction | ✅ (optional) | | Speaker diarization | Coming soon | | SRT/VTT subtitles | ✅ (generated from timestamps) | ## Next steps - [Transcribe Audio](https://docs.lansonai.com/recorded/transcribe-audio) — submit a job - [Batch Jobs API](https://docs.lansonai.com/api-reference/batch-jobs) - [Segment API](https://docs.lansonai.com/api-reference/segment-transcription) — client-VAD clips - [Timestamps & Speakers](https://docs.lansonai.com/recorded/timestamps-speakers) # Transcribe Audio Submit a **batch transcription job** from a public audio URL. For client-VAD speech clips use [Segment Transcription](https://docs.lansonai.com/api-reference/segment-transcription) instead. ## Submit a job ### Request ```bash curl -X POST https://audio.lansonai.com/v1/audio/transcriptions/jobs \ -H "Authorization: Bearer sk-..." \ -H "Content-Type: application/json" \ -d '{ "audio_url": "https://cdn.example.com/meeting.wav", "language": "zh", "prompt": "Medical cardiology conference", "webhook_url": "https://your-server.com/webhook" }' ``` ### Request body | Field | Type | Required | Default | Description | | ----------------- | ------- | -------- | -------------- | --------------------------------------------------- | | `audio_url` | string | **yes** | — | Public audio URL, must be `http(s)://` | | `request_id` | string | no | random UUID | Idempotency key; same ID resumes from R2 checkpoint | | `language` | string | no | — | Language hint (e.g. `zh`, `en`) | | `model` | string | no | env default | STT model override | | `prompt` | string | no | — | Transcription context hint | | `segment_seconds` | number | no | 300 | Slice length in seconds, must be > 0 | | `response_format` | string | no | `verbose_json` | Response format | | `concurrency` | number | no | 6 | Concurrent chunk transcription, must be > 0 | | `webhook_url` | string | no | — | POST final result to this URL on completion | | `review` | boolean | no | `false` | Enable two-stage LLM review | | `metadata` | object | no | — | Context forwarded to review stages | ::callout{icon="i-lucide-info"} `response_format` defaults to `verbose_json` in the OpenAPI schema, but the workflow currently hardcodes `verbose_json` for every chunk transcription. Setting this field has no effect today. :: ### Response — 202 Accepted ```json { "request_id": "...", "workflow_id": "cf_55190d1a608984daf77cbfca6b7b5438891436f8522f2be7be12fc93fa239ad4", "status": "queued", "endpoint": "GET /cf_...", "poll_endpoint": "GET /v1/audio/transcriptions/jobs/cf_..." } ``` `workflow_id` is a Cloudflare Workflow id (`cf_` + 64 hex), not a UUID. ## Poll for results ```bash curl https://audio.lansonai.com/v1/audio/transcriptions/jobs/ \ -H "Authorization: Bearer sk-..." ``` Legacy aliases: `GET /`, `GET /v1/audio/transcriptions/` ### Status values | Status | Meaning | | ------------ | ------------------------- | | `queued` | Waiting to start | | `running` | Processing | | `complete` | Done (result in `output`) | | `errored` | Failed (error in `error`) | | `terminated` | Terminated | ### Complete output ```json { "status": "complete", "output": { "result": { "segments": [ { "id": 0, "start_time": 0.0, "end_time": 3.2, "duration": 3.2, "text": "The weather is nice today", "confidence": 0.95 } ], "summary": { "total_duration": 120.5, "total_speech_duration": 95.3, "overall_speech_ratio": 0.79, "num_segments": 45 }, "metadata": { "language": "zh", "chunk_count": 3, "audio_duration_seconds": 120.5 } } } } ``` ## Webhook Set `webhook_url` to receive the final result via HTTP POST when the job completes. ## LLM review Set `review: true` to enable two-stage LLM correction. ## R2 artifacts Results are stored in R2 under `transcription/{request_id}/`: | Artifact | Key | | ---------------------------- | --------------------------------------------- | | Per-chunk transcription | `chunk_{i}.json` | | Raw aggregated transcript | `result.json` | | Per-chunk review annotations | `review_chunk_{i}.json` (when `review: true`) | | Final corrected transcript | `reviewed_result.json` (when `review: true`) | ## Idempotency Provide `request_id` for idempotency — already-transcribed chunks are skipped on retry. ## Related - [Batch Jobs API](https://docs.lansonai.com/api-reference/batch-jobs) - [Segment Transcription](https://docs.lansonai.com/api-reference/segment-transcription) - [Timestamps & Speakers](https://docs.lansonai.com/recorded/timestamps-speakers) - [Translation](https://docs.lansonai.com/recorded/translation) # Timestamps & Speakers Timestamp and speaker information in transcription results. ## Segment timestamps Each segment includes timestamps: ```json { "id": 0, "start_time": 0.0, "end_time": 3.2, "duration": 3.2, "text": "The weather is nice today", "confidence": 0.95 } ``` | Field | Type | Description | | ------------ | ------ | ---------------------------------------- | | `id` | number | Segment index (from 0) | | `start_time` | number | Start time in seconds | | `end_time` | number | End time in seconds | | `duration` | number | Duration in seconds | | `text` | string | Transcribed text | | `confidence` | number | Confidence score (0–1, 3 decimal places) | ### Confidence - Range: 0 to 1 - 3 decimal places - Calculated from `1 - no_speech_prob` when upstream provides it ### Global timestamps After slicing, timestamps are aggregated to global time. `start_time` and `end_time` are relative to the original audio, not chunk-internal time. ## Summary statistics ```json { "summary": { "total_duration": 120.5, "total_speech_duration": 95.3, "overall_speech_ratio": 0.79, "num_segments": 45 } } ``` | Field | Description | | ----------------------- | -------------------------------- | | `total_duration` | Total audio duration (seconds) | | `total_speech_duration` | Actual speech duration (seconds) | | `overall_speech_ratio` | Speech ratio | | `num_segments` | Number of segments | ## Metadata ```json { "metadata": { "language": "zh", "model": "whisper-large-v3-turbo", "chunk_count": 3, "audio_duration_seconds": 120.5 } } ``` ## Speaker diarization Speaker diarization is **not available** in the current version and is on the roadmap. ::callout{color="amber" icon="i-lucide-triangle-alert"} Current segments do not include a `speaker` field. Do not rely on speaker information. :: ## Precision notes - Timestamp precision depends on the upstream model and slicing granularity - Timestamps at slice boundaries are corrected during aggregation - Short speech segments (<1s) may have less precise timestamps ## Related - [Transcribe Audio](https://docs.lansonai.com/recorded/transcribe-audio) — submit a job - [Subtitles](https://docs.lansonai.com/recorded/subtitles) — generate subtitles from timestamps - [Batch Jobs API](https://docs.lansonai.com/api-reference/batch-jobs) — API reference # Subtitles Generate subtitle files from transcription result timestamps. ## From transcription results to subtitles Segment data from a completed transcription can be directly converted to SRT or VTT: ```json { "segments": [ { "id": 0, "start_time": 0.0, "end_time": 3.2, "text": "The weather is nice today" }, { "id": 1, "start_time": 3.5, "end_time": 8.1, "text": "It might rain tomorrow" } ] } ``` ### SRT format ```text 1 00:00:00,000 --> 00:00:03,200 The weather is nice today 2 00:00:03,500 --> 00:00:08,100 It might rain tomorrow ``` ### VTT format ```text WEBVTT 00:00:00.000 --> 00:00:03.200 The weather is nice today 00:00:03.500 --> 00:00:08.100 It might rain tomorrow ``` ## Time format conversion | Format | Time notation | | ------ | --------------------------------------- | | SRT | `HH:MM:SS,mmm` (comma for milliseconds) | | VTT | `HH:MM:SS.mmm` (dot for milliseconds) | | API | Seconds (float) | ## Conversion function ```python def seconds_to_srt(seconds): h = int(seconds // 3600) m = int((seconds % 3600) // 60) s = int(seconds % 60) ms = int((seconds % 1) * 1000) return f"{h:02d}:{m:02d}:{s:02d},{ms:03d}" def generate_srt(segments): lines = [] for i, seg in enumerate(segments, 1): lines.append(str(i)) lines.append(f"{seconds_to_srt(seg['start_time'])} --> {seconds_to_srt(seg['end_time'])}") lines.append(seg['text']) lines.append("") return "\n".join(lines) ``` ## Segmentation rules - Each API segment maps to one subtitle entry - Gaps between segments naturally become subtitle breaks - No additional subtitle length limits are imposed by the API - For traditional subtitle formatting (character-per-line limits), split long segments client-side ## Related - [Timestamps & Speakers](https://docs.lansonai.com/recorded/timestamps-speakers) — timestamp reference - [Transcribe Audio](https://docs.lansonai.com/recorded/transcribe-audio) — get transcription results - [Save a Transcript](https://docs.lansonai.com/guides/save-transcript) — persistence guide # Recorded Translation File-level translation for offline transcription. ::callout{color="amber" icon="i-lucide-triangle-alert"} `translate_to` is **not currently accepted** by the offline transcription endpoint. Sending it in the request body has no effect. :: ## Usage Specify `translate_to` when submitting a transcription job: ```bash curl -X POST https://audio.lansonai.com/v1/audio/transcriptions/jobs \ -H "Authorization: Bearer sk-..." \ -H "Content-Type: application/json" \ -d '{ "audio_url": "https://cdn.example.com/audio.wav", "language": "zh", "translate_to": "en" }' ``` | Parameter | Description | | -------------- | ------------------------------- | | `language` | Source language of the audio | | `translate_to` | Target language for translation | Omit `translate_to` to get transcription only, without translation. ## Source and translation correspondence Translation results correspond to source segments. Each segment's translation follows the source text. ## Latency Translation adds processing time. Measured translation latency is available on the [Service Status](https://control.lansonai.com/status/lanson-audio){rel=""nofollow""} page. ## Related - [Transcribe Audio](https://docs.lansonai.com/recorded/transcribe-audio) — submit a job - [Live Translation](https://docs.lansonai.com/concepts/live-translation) — translation concepts - [Realtime Translation](https://docs.lansonai.com/realtime/realtime-translation) — real-time translation # Structured Results ## Segment list Each segment's text, timestamps, and confidence. The core structured output. ```json [Response] { "id": 0, "start_time": 0.0, "end_time": 3.2, "text": "The weather is nice today", "confidence": 0.95 } ``` ## Summary statistics Audio-level statistics: ```json [Response] { "summary": { "total_duration": 120.5, "total_speech_duration": 95.3, "overall_speech_ratio": 0.79, "num_segments": 45 } } ``` ## Metadata Processing metadata: ```json [Response] { "metadata": { "language": "zh", "model": "whisper-large-v3-turbo", "chunk_count": 3, "audio_duration_seconds": 120.5 } } ``` ## LLM review results When `review: true` is enabled, the reviewed transcript replaces the raw result: 1. **Per-chunk review**: each chunk independently corrected 2. **Global review**: final correction after aggregation Reviewed results are stored in `reviewed_result.json` in R2. ## Not yet stable | Capability | Status | | ------------------------ | ------- | | Automatic summary (text) | Roadmap | | Chapter segmentation | Roadmap | | Action item extraction | Roadmap | | Speaker diarization | Roadmap | ::callout{color="amber" icon="i-lucide-triangle-alert"} Do not treat experimental or unmarked capabilities as stable. If a capability is marked as "Roadmap", it is not yet implemented or may change. :: ## Stability markers | Marker | Meaning | | ------------ | ------------------------------------------ | | Stable | Shipped. Safe to depend on in production. | | Preview | Available, but contracts may still change. | | Experimental | May change or be removed without notice. | | Roadmap | Planned, not yet implemented. | ## Related - [Transcribe Audio](https://docs.lansonai.com/recorded/transcribe-audio) — get results - [Timestamps & Speakers](https://docs.lansonai.com/recorded/timestamps-speakers) — timestamp details - [Batch Jobs API](https://docs.lansonai.com/api-reference/batch-jobs) — API reference # Voice Context Layer A Voice Context Layer sits between live speech recognition and the application that consumes it. **Realtime**, **Segment**, and **Batch jobs** are three HTTP/WS planes of this layer: | Plane | Who segments speech | | ------------ | ------------------------------------------ | | Realtime WS | Server (VAD on PCM stream) | | Segment HTTP | **Client** (you POST each utterance clip) | | Batch jobs | Server (time-based slicing of `audio_url`) | Its job is not simply to produce words. Its job is to turn continuously changing speech into **usable context**. ## Speech recognition is only the first layer A speech recognition system may produce something like: ```text I think we should meet on I think we should meet on Thursday I think we should meet on Thursday afternoon I think we should meet Thursday afternoon instead ``` Each output may be reasonable at the moment it was generated. But an application has different questions: - Which text should already be visible? - Which text may still change? - When is a thought complete enough to translate? - When should downstream logic act on it? - How should corrections appear without disrupting the reader? Those questions exist above recognition itself. That is the problem addressed by the Voice Context Layer. ## From audio to usable context Conceptually, a live speech system can be viewed as: ```text Audio ↓ Speech recognition ↓ Voice Context Layer ↓ Application ``` The recognition layer determines what was likely spoken. The Voice Context Layer determines how that evolving information should become usable by the application. This may include: - stabilization - contextual correction - segmentation - translation readiness - presentation continuity - lifecycle state ## Why this matters For offline transcription, the system can wait until the audio has finished before producing the final result. Live applications cannot. They must continuously balance two competing goals: **Responsiveness** Show useful information as soon as possible. **Stability** Avoid repeatedly changing information the user has already read. LansonAI is designed around that tradeoff. The objective is not simply to make text appear faster. It is to make live speech **ready to use while it is still happening**. # Stable vs. Partial Text ::callout{icon="i-lucide-info"} Current API behavior This page explains the general concept of partial vs. stable text. In the **current realtime API**, only final/stable transcription events are pushed to the client (`conversation.item.input_audio_transcription.completed`). There is no partial-text event stream at this time, so you do not need to implement rollback or replacement logic. :: Live transcription is inherently provisional. When a speaker is still talking, the system does not yet have all of the information required to interpret the utterance. That means some text should be treated as **working text**, while other text can become increasingly stable. ## Partial text Partial text represents the system's current interpretation of ongoing speech. It is optimized for responsiveness. Partial text may change as: - additional audio arrives - sentence boundaries become clearer - ambiguous words are resolved - surrounding context changes the interpretation Applications should assume partial text is mutable. ## Stable text Stable text represents content that has reached a stronger level of contextual confidence. It is intended to provide a more reliable boundary for applications that need to: - display persistent text - translate a segment - save conversation history - trigger downstream processing - build conversational context The exact stabilization signals exposed by each API are documented in the corresponding API reference. ## Stability is a lifecycle It is useful to think of live text as moving through states rather than suddenly becoming correct: ```text speech ↓ working interpretation ↓ context develops ↓ stabilization ↓ stable context ``` The important distinction is therefore not simply: ```text wrong → correct ``` but: ```text mutable → increasingly reliable ``` This distinction becomes especially important when building user interfaces for live captions and translation. # StableStream StableStream is LansonAI's approach to maintaining readable continuity while live speech continues to evolve. Real-time speech recognition produces changing information. A naive interface can expose those changes directly: ```text we should probably we should probably ship we should probably ship the we should probably ship this we should probably ship this next we should probably ship this next week ``` If earlier words are repeatedly replaced or rearranged, the interface may technically be updating quickly while becoming difficult to read. StableStream addresses this boundary between **model updates** and **human-readable output**. ## Ready to read, not racing to display The fastest possible token is not always the most useful token. A live interface must balance: - response latency - linguistic uncertainty - correction quality - visual stability - reading continuity StableStream is designed around that balance. Its goal is to allow new information to arrive continuously without making previously presented information unnecessarily unstable. ## Recognition stability and visual stability These are related but different concepts. **Recognition stability** describes how confident the system is that speech has been interpreted correctly. **Visual stability** describes how much already-presented content moves or changes on screen. A good live speech experience needs both. StableStream operates at this boundary. ## Application behavior Applications consuming a live stream should distinguish between content that is still evolving and content that has become stable. The API documentation describes the exact events and state transitions available to clients. This page describes the underlying principle: > Live text should evolve without forcing the reader to repeatedly reconstruct what they have already understood. # Context-Aware Processing Speech is ambiguous when interpreted in isolation. A short audio fragment may contain multiple plausible interpretations. Additional speech often makes the intended meaning clearer. Consider: ```text Let's send it to Alex... Let's send it to Alex Chen... Let's send it to Alex Chen after the review. ``` The later context changes how earlier information should be interpreted and structured. ## Context is part of recognition LansonAI treats live speech as an evolving context rather than a sequence of independent audio fragments. This allows the system to use surrounding information when resolving ambiguity. Context can help with: - ambiguous words - names and terminology - sentence boundaries - corrections - semantic continuity - translation ## Context does not mean waiting for completion A system could obtain maximum context simply by waiting until the speaker finishes. That would defeat the purpose of real-time processing. The challenge is therefore: > Use enough context to improve interpretation without turning live speech into offline transcription. This tradeoff is central to LansonAI's real-time architecture. ## Context accumulates over time Conceptually: ```text audio₁ → interpretation₁ audio₂ + previous context → interpretation₂ audio₃ + accumulated context → interpretation₃ ``` The system continuously updates its understanding as new evidence arrives. Applications therefore receive speech as an evolving stream of context rather than a collection of isolated recognition requests. # Understanding Latency There is no single latency number for a live speech system. What users experience as "latency" is the result of several different stages. ## The latency pipeline A simplified real-time pipeline looks like: ```text speaker ↓ audio capture ↓ network transport ↓ speech recognition ↓ context processing ↓ optional translation ↓ application rendering ``` Each stage contributes to the final experience. ## First-result latency The time between incoming speech and the first usable recognition result. Lower first-result latency generally makes an interface feel more responsive. However, extremely early results may contain greater uncertainty. ## Stabilization latency The time required before evolving speech becomes sufficiently stable for a particular use. This is different from first-result latency. For example: ```text 300 ms → first interpretation appears 900 ms → surrounding context resolves ambiguity 1.2 s → segment becomes stable ``` These numbers are illustrative only. The important point is that **responsiveness and stability are different measurements**. ## Translation latency Live translation introduces another dependency. Translation quality improves when more linguistic context is available, while live experiences require output before the full conversation is known. This creates another latency-quality tradeoff. LansonAI therefore treats live translation as a streaming context problem rather than simply translating a completed transcript. ## Measure the user experience For live applications, useful latency measurements should reflect what the user actually experiences. Depending on the application, this may include: - time to first readable text - time to stable text - time to translated text - correction frequency - visible reflow - end-to-end interaction latency Optimizing only one number can make another part of the experience worse. # Live Translation Live translation is different from translating a finished transcript. With a finished transcript, the system already knows: - the complete sentence - punctuation - sentence boundaries - surrounding context - what comes next During live speech, none of those conditions are guaranteed. ## Translation while meaning is still forming Consider: ```text I don't think we should... I don't think we should launch... I don't think we should launch tomorrow. ``` A translation system receiving the first fragment must decide whether to produce output immediately or wait for additional context. Producing too early can cause repeated corrections. Waiting too long creates noticeable latency. Live translation therefore requires balancing: - responsiveness - linguistic context - translation quality - stability ## Source and translated context In live applications, source text and translated text are related streams. A change in the interpretation of the source may also affect the translation. Applications should therefore avoid assuming that translation is simply a second independent transcript. Instead: ```text live speech ↓ source context ↓ translation context ``` Both continue evolving while the conversation progresses. ## Designed for delivery The objective of live translation is not merely to eventually produce a correct translated transcript. It is to allow someone to **follow speech in another language while it is happening**. That distinction affects how the entire system should be designed. # Trinity Engine Trinity Engine is the shared processing foundation behind LansonAI's real-time voice capabilities. It sits below product experiences and developer-facing APIs. Conceptually: ```text Lanson Live Lanson Reception Developer APIs ↓ Trinity Engine ↓ Speech and context infrastructure ``` ## What Trinity Engine does Trinity Engine coordinates the processing required to turn incoming speech into application-ready context. This includes multiple stages that may operate at different timescales, such as: - speech processing - segmentation - contextual interpretation - correction - stabilization - multilingual processing The exact internal implementation may evolve independently of the developer-facing API. ## Why the separation matters Applications should not need to understand the internal processing architecture in order to use LansonAI. The developer contract exists at the API layer. Trinity Engine exists below that contract. This separation allows LansonAI to improve its internal speech and context processing while keeping application integrations stable. ## Relationship to StableStream Trinity Engine and StableStream describe different layers. **Trinity Engine** The processing foundation. **StableStream** The behavior and continuity of evolving context exposed to live applications. This distinction is important because the developer experience should be defined by observable behavior rather than internal implementation details. # Browser Live Captions Build a browser-based live caption interface using the LansonAI realtime API. ## Architecture ```text Browser Your backend LansonAI │ │ │ ├── request session token ─────→ POST session-token ───────→ rt_... token │←── rt_... token ──────────────│ │ │ │ │ ├── open WebSocket ─────────────────────────────────────────→ session.created │ ?access_token=rt_... │ │ │ │ ├── capture mic (16kHz mono) ───│ │ ├── send PCM16 frames ──────────────────────────────────────→ VAD + STT │←── conversation.item.input_audio_transcription.completed ───────────────────────────│ │ │ │ └── display captions │ │ ``` ## Step 1: Get a session token Your backend exchanges the long-lived API key for a 60-second session token: ```javascript // backend route app.post('/session-token', async (req, res) => { const resp = await fetch('https://audio.lansonai.com/v1/audio/transcriptions/session-token', { method: 'POST', headers: { Authorization: `Bearer ${process.env.LANSON_AUDIO_API_KEY}` }, }); res.json(await resp.json()); }); ``` ::callout{color="amber" icon="i-lucide-triangle-alert"} Never expose `sk-...` in frontend code. Always proxy through your backend. :: ## Step 2: Open WebSocket ```javascript const { token } = await fetch('/session-token', { method: 'POST' }).then(r => r.json()); const ws = new WebSocket( `wss://audio.lansonai.com/v1/audio/transcriptions/stream?access_token=${token}` ); ``` ## Step 3: Capture and send audio ```javascript const audioContext = new AudioContext({ sampleRate: 16000 }); const stream = await navigator.mediaDevices.getUserMedia({ audio: true }); const source = audioContext.createMediaStreamSource(stream); // Use AudioWorklet to get PCM16LE frames await audioContext.audioWorklet.addModule('pcm-processor.js'); const node = new AudioWorkletNode(audioContext, 'pcm-processor'); source.connect(node); // pcm-processor.js posts PCM16LE buffers to main thread node.port.onmessage = (e) => { // Send as binary frame ws.send(e.data); // ArrayBuffer of PCM16LE data }; ``` Minimal `pcm-processor.js`: ```javascript class PcmProcessor extends AudioWorkletProcessor { process(inputs) { const input = inputs[0][0]; // mono, 16kHz if (input) { const pcm16 = new Int16Array(input.length); for (let i = 0; i < input.length; i++) { const s = Math.max(-1, Math.min(1, input[i])); pcm16[i] = s < 0 ? s * 0x8000 : s * 0x7fff; } this.port.postMessage(pcm16.buffer); } return true; } } registerProcessor('pcm-processor', PcmProcessor); ``` ## Step 4: Display captions ```javascript ws.onmessage = (event) => { const data = JSON.parse(event.data); if (data.type === 'conversation.item.input_audio_transcription.completed') { const div = document.getElementById('captions'); const p = document.createElement('p'); p.textContent = data.text; div.appendChild(p); } }; ``` ## Browser security - **Never** put `sk-...` in frontend code - Always obtain `rt_...` tokens from your backend - Tokens expire in 60 seconds — get a fresh one for each connection - WebSocket `?access_token` keeps the long-lived key off the client ## Related - [Build a Stable Caption UI](https://docs.lansonai.com/guides/stable-caption-ui) — advanced UI patterns - [Authentication](https://docs.lansonai.com/start/authentication) — session token details - [Realtime Quickstart](https://docs.lansonai.com/realtime/quickstart) — more examples # Server-side Streaming Connect your backend to LansonAI to stream audio from server-side sources (files, pipelines, telephony). ## Architecture ```text Your backend LansonAI Upstream STT │ │ │ ├── WS connect (Bearer sk-...) ─→ session.created │ │ │ │ ├── read audio file ───────────│ │ ├── convert to PCM16LE/16k ─────│ │ ├── send binary frames ─────────→ gateway relay ────────────→ VAD + STT │←── conversation.item.input_audio_transcription.completed ─│←──────────│ │ │ │ ├── close ──────────────────────→ meter flush │ ``` ## Authentication Server-side connections use the API key directly: ```typescript const ws = new WebSocket( "wss://audio.lansonai.com/v1/audio/transcriptions/stream", { headers: { Authorization: "Bearer sk-..." } } ); ``` No session token needed — the key never leaves your server. ## Streaming a file ```typescript import { readFileSync } from "fs"; import WebSocket from "ws"; const ws = new WebSocket( "wss://audio.lansonai.com/v1/audio/transcriptions/stream", { headers: { Authorization: "Bearer sk-..." } } ); ws.on("open", () => { // Read pre-converted PCM16LE 16kHz mono audio const audio = readFileSync("audio.pcm"); const FRAME_BYTES = 3200; // 100ms at 16kHz 16-bit mono let offset = 0; const sendFrame = () => { if (offset >= audio.length) { ws.send(JSON.stringify({ type: "input_audio_buffer.flush" })); return; } const frame = audio.slice(offset, offset + FRAME_BYTES); ws.send(frame); // binary frame offset += FRAME_BYTES; setTimeout(sendFrame, 100); // simulate real-time }; sendFrame(); }); ws.on("message", (data) => { const event = JSON.parse(data.toString()); if (event.type === "conversation.item.input_audio_transcription.completed") { console.log(`[${event.utterance_index}] ${event.text}`); } }); ``` ## Converting audio server-side ```bash ffmpeg -i input.mp3 -ar 16000 -ac 1 -c:a pcm_s16le output.pcm ``` ## Connection management - Keep the connection alive by sending frames at a steady rate - If audio source pauses, send silent frames to avoid idle timeout - Close explicitly when done to flush meter data - For long-running streams, monitor session duration limits ## Token management Server-side connections do not need session tokens. Use `Authorization: Bearer sk-...` in the WebSocket upgrade headers. ## Related - [Realtime Quickstart](https://docs.lansonai.com/realtime/quickstart) — basic examples - [Audio Input](https://docs.lansonai.com/realtime/audio-input) — format requirements - [Connection Lifecycle](https://docs.lansonai.com/realtime/connection-lifecycle) — timeouts # Build a Stable Caption UI How to consume realtime transcription events and build a caption UI that doesn't jitter. ## The problem with naive approaches A naive live caption UI might try to update text on every event. With partial-text systems, this causes visible jitter as words are repeatedly replaced. LansonAI's StableStream eliminates this: each `conversation.item.input_audio_transcription.completed` event contains final, stable text. ## Recommended render strategy ### Append-only display ```typescript const captions = document.getElementById('captions'); ws.onmessage = (event) => { const data = JSON.parse(event.data); if (data.type === 'conversation.item.input_audio_transcription.completed') { // Append directly — text is final, no replacement needed const p = document.createElement('p'); p.dataset.utterance = data.utterance_index; p.textContent = data.text; captions.appendChild(p); captions.scrollTop = captions.scrollHeight; } }; ``` ### With optional status indicators ```typescript ws.onmessage = (event) => { const data = JSON.parse(event.data); if (data.type === 'input_audio_buffer.speech_started') { showIndicator(data.utterance_index, 'listening'); } if (data.type === 'input_audio_buffer.speech_stopped') { showIndicator(data.utterance_index, 'processing'); } if (data.type === 'conversation.item.input_audio_transcription.completed') { hideIndicator(data.utterance_index); appendCaption(data.utterance_index, data.text); } }; ``` ## What you do NOT need to do - ❌ No partial → final text replacement mapping - ❌ No waiting for "stabilization" (received text is already stable) - ❌ No text rollback or undo - ❌ No debounce or jitter elimination - ❌ No re-rendering of previously shown text ## Scrolling behavior For long sessions, manage scroll behavior: ```typescript function appendCaption(index, text) { const p = document.createElement('p'); p.textContent = text; // Auto-scroll if user is near bottom const isNearBottom = captions.scrollHeight - captions.scrollTop - captions.clientHeight < 100; captions.appendChild(p); if (isNearBottom) { captions.scrollTop = captions.scrollHeight; } } ``` ## Styling for readability ```css #captions p { margin: 0.25em 0; padding: 0.25em 0.5em; border-radius: 4px; transition: opacity 0.2s; } ``` ## Related - [StableStream](https://docs.lansonai.com/concepts/stablestream) — the stability contract - [Stable vs. Partial Text](https://docs.lansonai.com/concepts/stable-vs-partial-text) — concept - [Browser Live Captions](https://docs.lansonai.com/guides) — full browser setup # Build Live Translation Build a bilingual caption UI that shows source text and translation side by side. ::callout{color="amber" icon="i-lucide-triangle-alert"} Real-time `translate_to` is **not currently enabled** for external sessions. The public realtime gateway forwards `language` but does not forward `translate_to`, so a `session.update` with that field has no effect. :: ## Layout ```text ┌──────────────────────────────────────────┐ │ Source │ │ The weather is nice today │ │ ─────────────────────────────── │ │ Translation │ │ The weather is nice today │ └──────────────────────────────────────────┘ ``` ## Implementation ```javascript ws.onmessage = (event) => { const data = JSON.parse(event.data); if (data.type === 'conversation.item.input_audio_transcription.completed') { const entry = document.createElement('div'); entry.className = 'translation-entry'; // Source text const source = document.createElement('div'); source.className = 'source-text'; source.textContent = data.text; entry.appendChild(source); // Translation (if available in the event) if (data.translation) { const translated = document.createElement('div'); translated.className = 'translated-text'; translated.textContent = data.translation; entry.appendChild(translated); } container.appendChild(entry); container.scrollTop = container.scrollHeight; } }; ``` ## CSS ```css .translation-entry { margin: 0.5em 0; padding: 0.5em; border-left: 3px solid #4f46e5; } .source-text { font-weight: 500; } .translated-text { margin-top: 0.25em; opacity: 0.85; font-style: italic; } ``` ## Language switching When the user switches the source language, update the session: ```javascript function switchLanguage(lang) { ws.send(JSON.stringify({ type: 'session.update', language: lang, })); } ``` New utterances after the switch use the new source language. ## Related - [Live Translation](https://docs.lansonai.com/concepts/live-translation) — translation concepts - [Realtime Translation](https://docs.lansonai.com/realtime/realtime-translation) — API details - [Build a Stable Caption UI](https://docs.lansonai.com/guides/stable-caption-ui) — caption UI patterns # Save a Transcript How to collect and persist a complete transcript after a realtime session ends. ## Collecting segments During the session, accumulate `conversation.item.input_audio_transcription.completed` events: ```typescript const segments = []; ws.onmessage = (event) => { const data = JSON.parse(event.data); if (data.type === 'conversation.item.input_audio_transcription.completed') { segments.push({ utterance_index: data.utterance_index, text: data.text, language: data.language, audio_duration_ms: data.audio_duration_ms, latency_ms: data.latency_ms, timestamp: Date.now(), }); } }; ws.onclose = () => { // Session ended — save the complete transcript saveTranscript(segments); }; ``` ## Merging into a full transcript ```typescript function mergeTranscript(segments) { // Sort by utterance_index to ensure order segments.sort((a, b) => a.utterance_index - b.utterance_index); return segments.map(s => s.text).join('\n'); } ``` ## Storage format Recommended JSON structure: ```json { "session_id": "sess_...", "started_at": "2026-08-15T10:00:00Z", "ended_at": "2026-08-15T10:30:00Z", "language": "zh", "segments": [ { "utterance_index": 0, "text": "The weather is nice today", "audio_duration_ms": 3200, "latency_ms": 480 } ], "full_text": "The weather is nice today\nIt might rain tomorrow" } ``` ## Detecting missing segments Check for gaps in `utterance_index`: ```typescript function findMissing(segments) { const indices = segments.map(s => s.utterance_index); const max = Math.max(...indices); const missing = []; for (let i = 0; i <= max; i++) { if (!indices.includes(i)) missing.push(i); } return missing; } ``` Missing segments may occur if frames were dropped due to backpressure or concurrent utterance limits. ## Offline alternative For cases where you need a guaranteed complete transcript, consider using the [Recorded API](https://docs.lansonai.com/recorded/transcribe-audio) instead — submit the recorded audio file and get the full structured result. ## Related - [Transcript Lifecycle](https://docs.lansonai.com/realtime/transcript-lifecycle) — event states - [Transcribe Audio](https://docs.lansonai.com/recorded/transcribe-audio) — offline alternative - [Subtitles](https://docs.lansonai.com/recorded/subtitles) — subtitle generation # Reconnect a Live Session Recommended design for handling WebSocket disconnects in realtime sessions. ## Reconnect decision matrix | Close code | Meaning | Action | | ---------: | ---------------------- | ----------------------------------- | | 1000 | Normal close | Do not reconnect | | 4408 | Idle timeout | Reconnect immediately | | 1008 | Session duration limit | Reconnect immediately (new session) | | 1009 | Frame too large | Fix frame size, then reconnect | | 1011 | Client socket error | Reconnect with backoff | | 1013 | Upstream unavailable | Exponential backoff | ## Exponential backoff ```typescript function reconnectWithBackoff(url, maxRetries = 5) { let attempt = 0; let lastLanguage = 'zh'; function connect() { const ws = new WebSocket(url); ws.onopen = () => { console.log('Connected'); attempt = 0; // reset backoff on success // Re-send session configuration ws.send(JSON.stringify({ type: 'session.update', language: lastLanguage })); }; ws.onclose = (event) => { if (event.code === 1000) return; // normal close if (attempt < maxRetries) { const delay = Math.min(1000 * Math.pow(2, attempt), 30000); attempt++; console.log(`Reconnecting in ${delay}ms (attempt ${attempt})`); setTimeout(connect, delay); } }; ws.onmessage = (event) => { // Handle events as normal }; } connect(); } ``` ## Session continuation After reconnecting: - You get a **new** `session_id` — sessions do not resume - Re-send `session.update` to restore language, VAD, and other settings - Previous utterances are not re-delivered - If you need the complete transcript, accumulate segments client-side ## Keep-alive strategies To avoid idle timeout (4408): ```typescript // Send silent frames to keep connection alive function keepAlive(ws) { setInterval(() => { if (ws.readyState === WebSocket.OPEN) { const silence = new ArrayBuffer(3200); // 100ms of silence ws.send(silence); } }, 5000); // every 5 seconds } ``` ## Session token refresh Browser connections: session tokens expire in 60 seconds. If reconnection happens after token expiry: ```typescript async function connectWithFreshToken() { const { token } = await fetch('/session-token', { method: 'POST' }).then(r => r.json()); return new WebSocket( `wss://audio.lansonai.com/v1/audio/transcriptions/stream?access_token=${token}` ); } ``` ## Related - [Connection Lifecycle](https://docs.lansonai.com/realtime/connection-lifecycle) — close codes and timeouts - [Reconnection & Retries](https://docs.lansonai.com/production/reconnection-retries) — production retry strategies - [Authentication](https://docs.lansonai.com/start/authentication) — session token refresh # Client-VAD speech segments Use [Segment Transcription](https://docs.lansonai.com/api-reference/segment-transcription) when **your application** already knows where speech starts and ends. ## When to use Segment | Use Segment | Use something else | | ----------------------------------------------------- | -------------------------------------------------------------------------------------- | | You run VAD locally and upload one clip per utterance | Continuous mic stream → [Realtime WS](https://docs.lansonai.com/realtime) | | You need a synchronous transcript in \~1s | Long file at a URL → [Batch jobs](https://docs.lansonai.com/recorded/transcribe-audio) | | OpenAI Whisper `file` upload pattern | Server should detect silence → Realtime WS | ## Contract 1. **You segment** — only POST clips that contain speech you want transcribed. 2. **We do not filter** — silence, near-empty WAVs, and noise are transcribed as-is. 3. **You pay for what you send** — wasting quota on silence is the caller's responsibility. 4. **No server HTTP retry** — up to 2 worker attempts (800ms each) per request; on **502** the client may resend the same clip. ## Typical flow ```text Client VAD detects utterance end → encode clip (e.g. WAV) → POST /v1/audio/transcriptions (multipart) → 200 + text (or 502 → client decides to retry) ``` Integrators like Flow follow this pattern: WebSocket voice session → local VAD → HTTP segment per utterance. ## Example ```bash curl -X POST https://audio.lansonai.com/v1/audio/transcriptions \ -H "Authorization: Bearer sk-..." \ -F "file=@utterance.wav" \ -F "language=zh" ``` ## Anti-patterns - Uploading a **full meeting recording** to Segment — use batch `audio_url` instead. - Streaming **continuous PCM** to Segment — use Realtime WS. - Expecting the API to **skip silence** — it will not. ## Related - [Segment API reference](https://docs.lansonai.com/api-reference/segment-transcription) - [Choose an API](https://docs.lansonai.com/start/choose-an-api) - [Errors](https://docs.lansonai.com/api-reference/errors) # API Reference — Overview ## Base URL ```text https://audio.lansonai.com ``` All relative paths in this reference are relative to this domain. ## Three transcription planes | Plane | Endpoint | Body | Success | Who segments speech | | -------------- | ------------------------------------ | --------------------- | ------------- | ------------------------- | | **Realtime** | `WS /v1/audio/transcriptions/stream` | PCM16LE stream | WS events | **Server** (VAD) | | **Segment** | `POST /v1/audio/transcriptions` | `multipart/form-data` | **200** sync | **Client** (VAD) | | **Batch jobs** | `POST /v1/audio/transcriptions/jobs` | JSON `audio_url` | **202** async | **Server** (time slicing) | ## Endpoints | Endpoint | Method | Purpose | Auth | | -------------------------------------------- | ------ | ----------------------------------------- | --------------------------- | | `/v1/audio/transcriptions/stream` | WS | Realtime streaming transcription | `Bearer sk-...` or `rt_...` | | `/v1/audio/transcriptions` | POST | Segment transcription (multipart, sync) | `Bearer sk-...` | | `/v1/audio/transcriptions/jobs` | POST | Submit batch transcription job | `Bearer sk-...` | | `/v1/audio/transcriptions/jobs/{workflowId}` | GET | Poll batch job (preferred) | `Bearer sk-...` | | `/v1/audio/transcriptions/{workflowId}` | GET | Poll batch job (alias) | `Bearer sk-...` | | `/{workflowId}` | GET | Poll batch job (legacy alias) | `Bearer sk-...` | | `/v1/audio/transcriptions/session-token` | POST | Issue short-lived browser WebSocket token | `Bearer sk-...` | | `/health` | GET | Liveness check | Public | | `/readyz` | GET | Realtime readiness check | Public | | `/openapi.json` | GET | OpenAPI specification | Public | | `/docs` | GET | Swagger UI | Public | ::callout{color="primary" icon="i-lucide-info"} Legacy JSON on `POST /v1/audio/transcriptions` Still accepted temporarily with `Deprecation: true` and a `Link` to `/jobs`. Prefer `POST /v1/audio/transcriptions/jobs` for batch jobs. :: Service health board (latency / probe SLI) is **not** on this host. Use the public ops page: - Production: {rel=""nofollow""} ## Versioning The current version uses the `/v1/...` path prefix. Future breaking changes will use `/v2/...`; `/v1/` will remain supported during a reasonable transition period. ## Authentication All transcription endpoints require an API key via `Authorization: Bearer sk-...`. See [Authentication](https://docs.lansonai.com/start/authentication) for browser-side session token flow. ## Page layout Each endpoint page follows a consistent template: - **What it does** — functional description - **Endpoint** — URL and method - **Authentication** — auth requirements - **Request** — request format - **Parameters** — parameter reference - **Example** — code example - **Response** — response format - **Errors** — relevant error codes - **Related** — links to guides and concepts ::callout{color="primary" icon="i-lucide-info"} Reference pages are for field facts This section does not contain product narratives or tutorials. For concepts, see [Concepts](https://docs.lansonai.com/concepts). For integration guides, see [Guides](https://docs.lansonai.com/guides). :: # Admin API These endpoints are for **internal operators and dashboards**, not for regular API consumers. All key-management routes require the `GATEWAY_ADMIN_TOKEN` via `Authorization: Bearer `. The single exception is `GET /v1/external-transcription/plans`, which is public and intended for SDK clients and dashboards that need to display plan limits. ## Base URL ```text https://audio.lansonai.com ``` ## Authentication | Header | Value | | --------------- | ------------------------------ | | `Authorization` | `Bearer ` | A missing or incorrect admin token returns: ```json { "ok": false, "error": { "code": "admin_unauthorized", "message": "Admin token required.", "retryable": false } } ``` ## Response envelope All admin endpoints use the same envelope: ```json { "ok": true, "data": { ... } } ``` Errors: ```json { "ok": false, "error": { "code": "...", "message": "...", "retryable": false } } ``` ## Endpoints ### `GET /v1/external-transcription/plans` Public. Returns every known plan and its resolved limits. ```bash curl https://audio.lansonai.com/v1/external-transcription/plans ``` Response: ```json { "ok": true, "data": [ { "plan": "free", "limits": { "maxConcurrentConnections": 1, "maxConcurrentUtterances": 2, "maxConnectionsPerMinute": 6, "maxRequestsPerMinute": 12, "maxAudioSecondsPerRequest": 900, "maxSessionSeconds": 900, "maxAudioSecondsPerSession": 900, "maxAudioSecondsPerWindow": 3600, "idleTimeoutSeconds": 60 } } ] } ``` ### `POST /v1/external-transcription/keys` Issue a new API key. The **plaintext key is returned exactly once** in the response; after that only metadata and a hash are stored. ```bash curl -X POST https://audio.lansonai.com/v1/external-transcription/keys \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{ "owner_id": "tenant_a", "plan": "pro", "label": "production", "expires_at": "2026-12-31T23:59:59Z" }' ``` Request fields: | Field | Type | Required | Description | | ------------ | ----------------- | -------- | -------------------------------------- | | `owner_id` | string | no | Tenant identifier (default: `default`) | | `plan` | string | no | Plan name (default: `free`) | | `label` | string | no | Human-readable label | | `expires_at` | string (ISO 8601) | no | Expiration timestamp | Response: ```json { "ok": true, "data": { "key": "sk-...", "id": "...", "ownerId": "tenant_a", "keyPrefix": "sk-xxxxxxxxx", "lastFour": "abcd", "plan": "pro", "status": "active", "label": "production", "expiresAt": "2026-12-31T23:59:59Z", "revokedAt": null, "lastUsedAt": null, "createdAt": "2026-08-16T...", "limits": { ... } } } ``` ### `GET /v1/external-transcription/keys` List key metadata. Optionally filter by `?owner_id=`. ```bash curl https://audio.lansonai.com/v1/external-transcription/keys \ -H "Authorization: Bearer " ``` Response: ```json { "ok": true, "data": [ { "id": "...", "ownerId": "tenant_a", "plan": "pro", "status": "active", ... } ] } ``` ### `GET /v1/external-transcription/keys/{id}` Inspect a single key. ```bash curl https://audio.lansonai.com/v1/external-transcription/keys/ \ -H "Authorization: Bearer " ``` ### `DELETE /v1/external-transcription/keys/{id}` Revoke a key. ```bash curl -X DELETE https://audio.lansonai.com/v1/external-transcription/keys/ \ -H "Authorization: Bearer " ``` ## Error codes | Code | HTTP | Description | | -------------------- | ---- | -------------------------------- | | `admin_unauthorized` | 401 | Missing or incorrect admin token | | `invalid_plan` | 400 | Unknown plan name | | `key_not_found` | 404 | Key ID does not exist | | `issue_failed` | 500 | Unable to issue a key | | `store_unavailable` | 500 | R2 key store unavailable | | `not_found` | 404 | Unknown admin route | ## Related - [Authentication](https://docs.lansonai.com/start/authentication) — regular API key and session-token flow - [Rate Limits](https://docs.lansonai.com/api-reference/rate-limits) — plan limits returned by `/plans` # Realtime API ## What it does Real-time streaming speech transcription. Open a WebSocket, send PCM16LE audio frames, receive transcription events. ## Endpoint ```text WSS /v1/audio/transcriptions/stream ``` ## Authentication Three authentication methods: | Method | Usage | Use case | | ------------------- | ------------------------------ | ----------------------- | | Bearer header | `Authorization: Bearer sk-...` | Server-side clients | | Query token | `?access_token=rt_...` | Browser (session token) | | Query token (alias) | `?token=rt_...` | Browser (session token) | Session tokens are obtained via `POST /v1/audio/transcriptions/session-token` and are valid for 60 seconds. See [Authentication](https://docs.lansonai.com/start/authentication). ## Connection ### Upgrade conditions - Request must be a WebSocket upgrade, otherwise `426 websocket_required` - Gateway must be enabled, otherwise `503 gateway_disabled` - Upstream circuit breaker must be closed, otherwise `503 upstream_unavailable` with `Retry-After` header - Connection quota must pass (concurrent connections + handshake rate) ### On success Receives a `session.created` event: ```json { "type": "session.created", "request_id": "...", "session_id": "sess_...", "audio": { "format": "pcm_s16le", "sample_rate": 16000, "channels": 1 }, "turn_detection": { "type": "server_vad" }, "plan": "free", "limits": { "max_concurrent_utterances": 2, "max_session_seconds": 900, "idle_timeout_seconds": 60, "remaining_audio_seconds": 3600 } } ``` ## Message format All messages are JSON text frames or binary PCM audio frames. - **Text frames**: JSON objects, must contain a `type` field - **Binary frames**: raw PCM audio data (ArrayBuffer) ## WebSocket close codes | Code | Close reason | Meaning | | ---: | ------------------------ | ---------------------------------------------------- | | 1000 | — | Normal close | | 1008 | `session_audio_limit` | Realtime session audio budget exceeded | | 1008 | `audio_quota_exhausted` | Billing-window audio budget exhausted | | 1008 | `session_duration_limit` | Session wall-clock duration exceeded | | 1009 | `audio_frame_too_large` | Audio frame exceeds 1 MiB | | 1011 | `client_error` | Client socket error or unexpected gateway error | | 1013 | `upstream_unavailable` | Realtime upstream unavailable (circuit breaker open) | | 1013 | `upstream_backpressure` | Upstream is not draining fast enough | | 1013 | `client_backpressure` | Client is not reading events fast enough | | 4408 | `idle_timeout` | No audio received within idle timeout window | ## Related - [Client Messages](https://docs.lansonai.com/api-reference/client-messages) — all client-sent messages - [Server Events](https://docs.lansonai.com/api-reference/server-events) — all server-pushed events - [Realtime Quickstart](https://docs.lansonai.com/realtime/quickstart) — complete example - [Authentication](https://docs.lansonai.com/start/authentication) — auth details # Client Messages Messages the client sends over the WebSocket connection. ## Overview | type | Transport | Purpose | | --------------------------- | ----------------- | ------------------------------------ | | `input_audio_buffer.append` | Text frame (JSON) | Send base64-encoded PCM audio | | *(binary frame)* | Binary frame | Send raw PCM audio data | | `input_audio_buffer.flush` | Text frame (JSON) | Trigger manual flush | | `flush` | Text frame (JSON) | Alias for `input_audio_buffer.flush` | | `session.update` | Text frame (JSON) | Update session parameters | ## input\_audio\_buffer.append Send base64-encoded PCM16LE audio data. ```json { "type": "input_audio_buffer.append", "audio": "" } ``` | Field | Type | Required | Description | | ------- | ------ | -------- | ------------------------------------- | | `type` | string | yes | Must be `"input_audio_buffer.append"` | | `audio` | string | yes | Base64-encoded PCM16LE audio data | **Errors**: - `invalid_audio`: `audio` is empty or not valid base64 ## Binary audio frames Send raw PCM binary data directly as an `ArrayBuffer`. Same effect as `input_audio_buffer.append` but without base64 encoding overhead. **Limits**: - Max frame size: 1 MiB (1,048,576 bytes). Exceeding this returns `audio_frame_too_large` and closes the connection (1009). - Frames are dropped when concurrent utterance limit is exceeded (`concurrent_utterance_limit`) - Frames are dropped or connection closed (1013) on upstream backpressure ## input\_audio\_buffer.flush Trigger a manual flush of the upstream buffer, ending the current speech segment. ```json { "type": "input_audio_buffer.flush" } ``` `"flush"` is an alias with the same effect. ## session.update Update session parameters. Can be sent at any time after connection. ```json { "type": "session.update", "language": "zh", "prompt": "Optional context hint" } ``` ### Supported fields | Field | Snake-case alias | Type | Description | | ----------------------- | --------------------------- | ------- | ----------------------------------------- | | `language` | — | string | Audio language (e.g. `zh`, `en`, `auto`) | | `prompt` | — | string | Context hint, max 2000 characters | | `backend` | — | string | Upstream backend selection | | `sttModel` | `stt_model` | string | STT model | | `normalizerModel` | `normalizer_model` | string | Text normalizer model | | `vad` | — | boolean | Enable VAD | | `vadThreshold` | `vad_threshold` | number | VAD sensitivity threshold | | `vadSilenceMs` | `vad_silence_ms` | number | Silence duration to trigger end-of-speech | | `vadPrefixMs` | `vad_prefix_ms` | number | Prefix padding duration | | `vadMinSpeechMs` | `vad_min_speech_ms` | number | Minimum speech segment duration | | `vadTargetSpeechMs` | `vad_target_speech_ms` | number | Target speech segment duration | | `vadMaxSpeechMs` | `vad_max_speech_ms` | number | Maximum speech segment duration | | `vadSmartSplitWindowMs` | `vad_smart_split_window_ms` | number | Smart split window | ::callout{color="amber" icon="i-lucide-triangle-alert"} Tenant isolation All fields not listed above are stripped before forwarding to the upstream. Identity and credential fields (`request_id`, `session_id`, `user_id`) are injected by the gateway and cannot be overridden by the client. :: ### Errors - `bad_json`: JSON parse failure - `bad_message`: Not a JSON object or unsupported `type` - `invalid_audio`: Invalid audio data ## Related - [Server Events](https://docs.lansonai.com/api-reference/server-events) — server-pushed events - [Session Configuration](https://docs.lansonai.com/realtime/session-configuration) — session config details - [Audio Input](https://docs.lansonai.com/realtime/audio-input) — audio format reference # Server Events Events the server pushes over the WebSocket connection. ## Overview | type | Source | Description | | ------------------------------------------------------- | ------------------ | ---------------------------------------------------------- | | `session.created` | Gateway | Sent on connection establishment | | `session.updated` | Upstream | Confirms session parameter update | | `input_audio_buffer.speech_started` | Upstream | Speech segment detected | | `input_audio_buffer.speech_stopped` | Upstream | Speech segment ended | | `conversation.item.input_audio_transcription.completed` | Upstream | A speech segment has been transcribed (**primary output**) | | `lanson.throttled` | Gateway | Throttle notification (max once per second) | | `error` | Gateway / upstream | Error event | ## session.created Sent by the gateway immediately after connection. ```json { "type": "session.created", "request_id": "...", "session_id": "sess_...", "audio": { "format": "pcm_s16le", "sample_rate": 16000, "channels": 1 }, "turn_detection": { "type": "server_vad" }, "plan": "free", "limits": { "max_concurrent_utterances": 2, "max_session_seconds": 900, "idle_timeout_seconds": 60, "remaining_audio_seconds": 3600 } } ``` | Field | Type | Description | | ---------------------------------- | -------------- | -------------------------------------------- | | `session_id` | string | Session identifier | | `audio.format` | string | Fixed: `pcm_s16le` | | `audio.sample_rate` | number | Fixed: 16000 | | `audio.channels` | number | Fixed: 1 | | `turn_detection.type` | string | Fixed: `server_vad` | | `plan` | string | Current plan | | `limits.max_concurrent_utterances` | number | Max concurrent utterances | | `limits.max_session_seconds` | number | Max session duration (seconds) | | `limits.idle_timeout_seconds` | number | Idle timeout (seconds) | | `limits.remaining_audio_seconds` | number \| null | Remaining audio budget (null when unlimited) | ## session.updated Confirms that a `session.update` message has taken effect. ```json { "type": "session.updated", "request_id": "...", "language": "zh", "turn_detection": { "type": "server_vad" } } ``` ## input\_audio\_buffer.speech\_started VAD detected the start of a speech segment. ```json { "type": "input_audio_buffer.speech_started", "request_id": "...", "utterance_index": 0 } ``` ## input\_audio\_buffer.speech\_stopped VAD detected the end of a speech segment. ```json { "type": "input_audio_buffer.speech_stopped", "request_id": "...", "utterance_index": 0, "reason": "end_of_speech", "audio_duration_ms": 3200, "speech_duration_ms": 2800 } ``` | Field | Type | Description | | -------------------- | ------- | ---------------------- | | `utterance_index` | number | Speech segment number | | `reason` | string | Stop reason | | `audio_duration_ms` | number | Total audio duration | | `speech_duration_ms` | number? | Actual speech duration | | `split_offset_ms` | number? | Split offset | | `split_rms` | number? | Split RMS value | ## conversation.item.input\_audio\_transcription.completed **Primary output event.** A speech segment has been transcribed. ```json { "type": "conversation.item.input_audio_transcription.completed", "request_id": "...", "utterance_index": 0, "text": "The weather is nice today", "language": "zh", "audio_duration_ms": 3200, "latency_ms": 480, "segments": [], "verbose": {} } ``` | Field | Type | Description | | ------------------- | ------ | ------------------------ | | `utterance_index` | number | Speech segment number | | `text` | string | Transcribed text | | `language` | string | Detected language | | `audio_duration_ms` | number | Audio duration | | `latency_ms` | number | Processing latency | | `segments` | array | Segment details | | `verbose` | object | Upstream additional info | ## lanson.throttled Gateway throttle notification, sent at most once per second. ```json { "type": "lanson.throttled", "reason": "concurrent_utterance_limit", "dropped_audio_frames": 3, "in_flight_utterances": 2 } ``` | `reason` value | Meaning | | ---------------------------- | ----------------------------------- | | `concurrent_utterance_limit` | Concurrent utterance limit exceeded | | `upstream_backpressure` | Upstream backpressure | | `upstream_connecting` | Upstream still connecting | ## error Error event. Can be sent at any time during the WebSocket connection. ```json { "type": "error", "code": "audio_frame_too_large", "message": "Audio frame exceeds maximum size.", "request_id": "..." } ``` Error codes are listed in [Errors](https://docs.lansonai.com/api-reference/errors). ## Event ordering ```text session.created → [audio frames...] → input_audio_buffer.speech_started → [more audio...] → input_audio_buffer.speech_stopped → conversation.item.input_audio_transcription.completed → [repeat...] → session close ``` ## Related - [Client Messages](https://docs.lansonai.com/api-reference/client-messages) — client-sent messages - [Errors](https://docs.lansonai.com/api-reference/errors) — full error reference - [Transcript Lifecycle](https://docs.lansonai.com/realtime/transcript-lifecycle) — state transitions # Segment Transcription API ## POST /v1/audio/transcriptions ### What it does Transcribe a **single speech clip** synchronously. This endpoint is OpenAI Whisper-compatible (`multipart/form-data` + `file`). **Client responsibility:** you must segment speech with your own VAD before calling. Do not send continuous silence — the service does not filter silence, empty clips, or low-energy audio. Whatever you POST is transcribed and billed. Typical integrators: voice clients (e.g. Flow) that detect utterance boundaries locally and upload each clip as WAV. ### Endpoint ```text POST /v1/audio/transcriptions ``` ### Authentication ```text Authorization: Bearer sk-... ``` ### Request `Content-Type: multipart/form-data` | Field | Type | Required | Default | Description | | ----------------- | ------ | -------- | -------------- | --------------------------------- | | `file` | file | **yes** | — | Speech clip (WAV, MP3, etc.) | | `model` | string | no | `whisper-1` | Model name passed to workers | | `language` | string | no | — | ISO-639-1 hint (`zh`, `en`, …) | | `prompt` | string | no | — | Context prompt for STT | | `response_format` | string | no | `verbose_json` | `verbose_json`, `json`, or `text` | | `worker_id` | string | no | — | Pin to a specific worker ID | ::callout{color="amber" icon="i-lucide-triangle-alert"} Not for long files or raw recordings Use [Batch Jobs](https://docs.lansonai.com/api-reference/batch-jobs) (`POST /v1/audio/transcriptions/jobs`) for `audio_url` async processing. Use [Realtime WS](https://docs.lansonai.com/api-reference/realtime-api) when the server should run VAD on a PCM stream. :: ### Example ```bash curl -X POST https://audio.lansonai.com/v1/audio/transcriptions \ -H "Authorization: Bearer sk-..." \ -F "file=@utterance.wav" \ -F "language=zh" \ -F "response_format=verbose_json" ``` ### Response — 200 OK `verbose_json` (default): ```json { "text": "你好世界", "language": "zh", "duration": 1.2, "segments": [], "words": [] } ``` Response headers: | Header | Description | | ----------------- | ------------------------------------------- | | `X-Worker-Id` | Worker that produced the result | | `X-Fallback-Used` | `1` if a fallback worker was used, else `0` | ### Latency and retries | Rule | Value | | ----------------------------- | ------------------------------------------------------------------------------------- | | Per-attempt timeout | 800 ms (Modal + ElevenLabs) | | Max POST attempts per request | 2 (primary warm Modal → ElevenLabs when applicable) | | Server HTTP retry | **No** — failed requests return an error; client may resend | | Cold Modal | Not POSTed on this request; health probe may run async; ElevenLabs used for cold path | Wall-clock budget is roughly **≤1.6 s** on the warm Modal + fallback path. ### Errors | Status | Code | Description | | ------ | ---------------------- | --------------------------------- | | 400 | `invalid_content_type` | Body is not `multipart/form-data` | | 400 | `missing_file` | No `file` field | | 401 | — | Authentication failed | | 502 | `transcription_failed` | All worker attempts failed | ```json { "error": { "message": "Realtime transcription timed out after 800ms", "type": "invalid_request_error", "code": "transcription_failed" } } ``` ## Related - [Client-VAD segments guide](https://docs.lansonai.com/guides/client-vad-segments) - [Batch Jobs](https://docs.lansonai.com/api-reference/batch-jobs) - [Choose an API](https://docs.lansonai.com/start/choose-an-api) - [Errors](https://docs.lansonai.com/api-reference/errors) # Batch Transcription Jobs ## POST /v1/audio/transcriptions/jobs ### What it does Submit an **audio URL** for asynchronous batch transcription. The service fetches the file, slices it by time, transcribes chunks concurrently, aggregates results, and optionally posts to a webhook. For synchronous client-VAD clips use [Segment Transcription](https://docs.lansonai.com/api-reference/segment-transcription) (`POST /v1/audio/transcriptions` with multipart). ### Endpoint ```text POST /v1/audio/transcriptions/jobs ``` `POST /` and `POST /v1/audio/transcriptions` with JSON still work but are **deprecated** (return `Deprecation: true`). ### Authentication ```text Authorization: Bearer sk-... ``` ### Request body ```json { "audio_url": "https://cdn.example.com/audio/meeting.wav", "language": "zh", "prompt": "Medical cardiology conference", "review": true, "metadata": { "medical_specialty": "cardiology" }, "webhook_url": "https://your-server.com/webhook" } ``` | Field | Type | Required | Default | Description | | ----------------- | ------- | -------- | -------------- | ---------------------------------------------------------------------------- | | `audio_url` | string | **yes** | — | Public URL of the audio file. Must be `http(s)://`. | | `request_id` | string | no | random UUID | Idempotency key. Re-submitting with the same ID resumes from R2 checkpoints. | | `language` | string | no | — | Audio language hint (e.g. `zh`, `en`). | | `model` | string | no | env default | STT model override. | | `prompt` | string | no | — | Transcription context hint. | | `segment_seconds` | number | no | 300 | Target slice length in seconds. Must be > 0. | | `response_format` | string | no | `verbose_json` | Response format (workflow hardcodes verbose\_json per chunk today). | | `concurrency` | number | no | 6 | Concurrent transcription of chunks. Must be > 0. | | `webhook_url` | string | no | — | POSTed the full result (raw + reviewed) when the job completes. | | `review` | boolean | no | `false` | Enable two-stage LLM review pipeline. | | `metadata` | object | no | — | Contextual metadata forwarded to review stages. | ### Response — 202 Accepted ```json { "request_id": "...", "workflow_id": "cf_55190d1a608984daf77cbfca6b7b5438891436f8522f2be7be12fc93fa239ad4", "status": "queued", "endpoint": "GET /cf_...", "poll_endpoint": "GET /v1/audio/transcriptions/jobs/cf_..." } ``` `workflow_id` is a Cloudflare Workflow instance id (`cf_` + 64 hex), **not** a UUID. ### Errors | Status | Description | | ------ | -------------------------------------------- | | 400 | Non-JSON body or invalid/missing `audio_url` | | 401 | Authentication failed | --- ## GET /v1/audio/transcriptions/jobs/{workflowId} ### What it does Poll batch job status and result (**preferred** path). ### Endpoint ```text GET /v1/audio/transcriptions/jobs/{workflowId} ``` Aliases: `GET /v1/audio/transcriptions/{workflowId}`, `GET /{workflowId}` ### Response ```json { "workflow_id": "cf_...", "status": "complete", "steps": [ ... ], "output": { ... }, "error": null } ``` ### Status values | Status | Meaning | | ------------ | ------------------------- | | `queued` | Waiting to start | | `running` | Processing | | `complete` | Done (result in `output`) | | `errored` | Failed (error in `error`) | | `terminated` | Terminated | | `paused` | Paused | | `waiting` | Waiting | ### Complete output structure ```json { "result": { "segments": [ { "id": 0, "start_time": 0.0, "end_time": 3.2, "duration": 3.2, "text": "The weather is nice today", "confidence": 0.95 } ], "summary": { "total_duration": 120.5, "total_speech_duration": 95.3, "overall_speech_ratio": 0.79, "num_segments": 45 }, "metadata": { "language": "zh", "model": "whisper-large-v3-turbo", "chunk_count": 3, "audio_duration_seconds": 120.5 } } } ``` ## R2 artifacts Result artifacts are stored in R2 under `transcription/{request_id}/`: | Artifact | R2 key | | ---------------------------- | --------------------------------------------- | | Per-chunk transcription | `chunk_{i}.json` | | Raw aggregated transcript | `result.json` | | Per-chunk review annotations | `review_chunk_{i}.json` (when `review: true`) | | Final corrected transcript | `reviewed_result.json` (when `review: true`) | ## Webhook Set `webhook_url` to receive the final result via HTTP POST when the job completes. ## OpenAPI Full OpenAPI specification at `GET /openapi.json`. Swagger UI at `GET /docs`. ## Related - [Transcribe Audio](https://docs.lansonai.com/recorded/transcribe-audio) — usage guide - [Segment Transcription](https://docs.lansonai.com/api-reference/segment-transcription) — sync client-VAD clips - [Timestamps & Speakers](https://docs.lansonai.com/recorded/timestamps-speakers) — timestamp details - [Errors](https://docs.lansonai.com/api-reference/errors) — error codes # Languages & Models ## Languages The service passes the `language` value straight through to the upstream STT endpoint. Use the language code expected by the upstream model. Common examples: - `zh` — Chinese - `en` — English - `auto` — let the upstream auto-detect the language Real-time, [segment](https://docs.lansonai.com/api-reference/segment-transcription), and [batch jobs](https://docs.lansonai.com/api-reference/batch-jobs) all accept a `language` field. In the realtime gateway, update the language at any time with a `session.update` message. ## Model and capability identifiers - `model` — offline STT model override (default `whisper-large-v3-turbo`) - `sttModel` / `stt_model` — realtime STT model selection - `normalizerModel` / `normalizer_model` — text normalizer model selection - `backend` — upstream backend selection Not all model fields are forwarded by the gateway; see [Session Configuration](https://docs.lansonai.com/realtime/session-configuration) for the current whitelist. Specific identifier values depend on the upstream provider and are not contractually fixed by this gateway. ## See also - [Live Translation](https://docs.lansonai.com/concepts/live-translation) — source language, target language, auto-detection, and code-switching - [Segment Transcription](https://docs.lansonai.com/api-reference/segment-transcription) — sync multipart parameters - [Batch Jobs](https://docs.lansonai.com/api-reference/batch-jobs) — async `audio_url` parameters - [Session Configuration](https://docs.lansonai.com/realtime/session-configuration) — realtime `session.update` fields # Errors ## Error response formats ### OpenAI format (`/v1/audio/*` endpoints) Used for realtime WebSocket, offline HTTP, and session-token endpoints: ```json { "error": { "message": "API key is missing.", "type": "invalid_request_error", "code": "missing_api_key" } } ``` `type` is `rate_limit_error` for 429 responses, otherwise `invalid_request_error`. ### Envelope format (Key management API) ```json { "ok": false, "error": { "code": "admin_unauthorized", "message": "Admin token required.", "retryable": false } } ``` `retryable` is `true` for 429 and ≥500 responses. ## HTTP error codes | Code | HTTP | Cause | Client action | | ----------------------------- | ---- | ----------------------------------------------------------------- | --------------------------- | | `missing_api_key` | 401 | No Authorization header / query token | Add API key | | `invalid_api_key` | 401 | Key format invalid or unknown | Check key | | `revoked_api_key` | 401 | Key has been revoked | Contact provider | | `expired_api_key` | 401 | Key has expired | Re-provision | | `auth_unavailable` | 503 | Key store unavailable | Retry later | | `invalid_session_token` | 401 | Session token invalid or expired | Obtain new token | | `token_issuer_unavailable` | 503 | Token issuer not configured | Retry later | | `websocket_required` | 426 | Non-WS upgrade request | Use WebSocket client | | `gateway_disabled` | 503 | Realtime gateway disabled | Contact provider | | `upstream_unavailable` | 503 | Upstream unavailable (circuit breaker) | Back off and retry | | `connection_rate_limited` | 429 | Connection rate exceeded | Honor `Retry-After` | | `concurrent_connection_limit` | 429 | Max concurrent connections exceeded | Wait for release | | `request_rate_limited` | 429 | Request rate exceeded | Honor `Retry-After` | | `audio_quota_exhausted` | 429 | Audio window budget exhausted | Retry later | | `audio_too_long` | 413 | Single request audio exceeds plan limit | Reduce audio duration | | `internal_error` | 500 | Uncaught exception | Contact provider | | `bad_request` | 400 | Invalid request body | Fix and retry | | `invalid_content_type` | 400 | Segment endpoint expects multipart; batch expects JSON on `/jobs` | Fix Content-Type and path | | `transcription_failed` | 502 | Segment transcription failed after worker attempts | Client may resend utterance | | `missing_file` | 400 | Segment request missing `file` field | Add multipart file | ## WebSocket error events Within an active WebSocket connection, the server may send `error` events: ```json { "type": "error", "code": "...", "message": "...", "request_id": "..." } ``` ### WebSocket error codes | Code | Description | Close code | | ---------------------------- | ----------------------------------- | ---------- | | `audio_frame_too_large` | Audio frame exceeds 1 MiB | 1009 | | `concurrent_utterance_limit` | Concurrent utterance limit exceeded | — | | `upstream_backpressure` | Upstream backpressure | 1013 | | `client_backpressure` | Client backpressure | — | | `upstream_unavailable` | Upstream unavailable | 1013 | | `idle_timeout` | Idle timeout reached | 4408 | | `session_duration_limit` | Session duration exceeded | 1008 | | `session_audio_limit` | Session audio budget exceeded | 1008 | | `audio_quota_exhausted` | Audio quota exhausted | 1008 | | `bad_json` | JSON parse failure | — | | `bad_message` | Invalid message type | — | | `invalid_audio` | Invalid audio data | — | ## Key management API errors | Code | HTTP | Description | | -------------------- | ---- | --------------------------- | | `admin_unauthorized` | 401 | Missing admin token | | `invalid_plan` | 400 | Unknown plan | | `invalid_expiry` | 400 | Invalid or past expiry date | | `store_unavailable` | 500 | R2 storage unavailable | | `key_not_found` | 404 | Key ID not found | | `not_found` | 404 | Unknown management route | ## Troubleshooting | Problem | Check | | -------------------------- | ------------------------------------------------------ | | 401 `missing_api_key` | Ensure request includes `Authorization: Bearer sk-...` | | 401 `invalid_api_key` | Verify key starts with `sk-`, length ≥ 16 | | 429 `*_rate_limited` | Honor `Retry-After` header, exponential backoff | | 503 `upstream_unavailable` | Check `/readyz` for upstream status | | WS `idle_timeout` | Keep sending audio frames, or close explicitly | | WS `audio_frame_too_large` | Reduce frame size to ≤ 1 MiB | ## Related - [Authentication](https://docs.lansonai.com/start/authentication) — auth methods - [Reconnection & Retries](https://docs.lansonai.com/production/reconnection-retries) — retry strategies - [Rate Limits](https://docs.lansonai.com/api-reference/rate-limits) — plan limits # Rate Limits Limits are enforced per API key and derived from `BASE_PLAN_LIMITS` in the gateway configuration. A value of `0` means unlimited for audio-second budgets. ## Plan limits | Limit | free | starter | pro | enterprise | | -------------------------------------- | -----------: | -------------: | ---------------: | ------------: | | Max concurrent WebSocket connections | 1 | 3 | 10 | 50 | | Max concurrent utterances | 2 | 4 | 8 | 16 | | Max connections per minute | 6 | 20 | 60 | 240 | | Max offline requests per minute | 12 | 60 | 300 | 1,200 | | Max audio seconds per offline request | 900 (15 min) | 7,200 (2 hr) | 14,400 (4 hr) | unlimited | | Max session wall-clock seconds | 900 (15 min) | 7,200 (2 hr) | 28,800 (8 hr) | 28,800 (8 hr) | | Max audio seconds per realtime session | 900 (15 min) | 7,200 (2 hr) | 28,800 (8 hr) | 28,800 (8 hr) | | Max audio seconds per billing window | 3,600 (1 hr) | 72,000 (20 hr) | 720,000 (200 hr) | unlimited | | Idle timeout (no audio) | 60 s | 120 s | 180 s | 300 s | ## What happens when a limit is hit | Scenario | Response | | -------------------------------- | ---------------------------------------------------------- | | Exceed concurrent connections | HTTP 429 `concurrent_connection_limit` | | Exceed connection rate | HTTP 429 `connection_rate_limited` with `Retry-After` | | Exceed offline request rate | HTTP 429 `request_rate_limited` with `Retry-After` | | Exceed per-request audio length | HTTP 413 `audio_too_long` | | Exceed realtime session duration | WebSocket close `1008` `session_duration_limit` | | Exceed realtime session audio | WebSocket close `1008` `session_audio_limit` | | Exhaust billing-window audio | WebSocket close `1008` `audio_quota_exhausted` or HTTP 429 | | No audio within idle timeout | WebSocket close `4408` `idle_timeout` | | Upstream backpressure | WebSocket close `1013` `upstream_backpressure` | | Client not reading fast enough | WebSocket close `1013` `client_backpressure` | | Frame larger than 1 MiB | WebSocket close `1009` `audio_frame_too_large` | ## Raising limits Contact the LansonAI team to move to a higher plan or negotiate custom `GATEWAY_PLAN_LIMITS_JSON` overrides. ## Related - [Realtime API](https://docs.lansonai.com/api-reference/realtime-api) — WebSocket endpoint details - [Client Messages](https://docs.lansonai.com/api-reference/client-messages) — audio frame and `session.update` messages - [Errors](https://docs.lansonai.com/api-reference/errors) — error and close codes # Production Checklist > 🚧 **Pending API contract** — This page will be written once the API endpoint, event schema, parameters, and limits are finalized. ## Planned coverage - What to verify before going live - Checklist items for authentication, reconnection, rate limits, monitoring, and data retention # Reconnection & Retries > 🚧 **Pending API contract** — This page will be written once the API endpoint, event schema, parameters, and limits are finalized. ## Planned coverage - WebSocket disconnects, 429, 5xx, backoff - Recommended retry strategies and boundaries # Latency Best Practices > 🚧 **Pending API contract** — This page will be written once the API endpoint, event schema, parameters, and limits are finalized. ## Planned coverage - Practical recommendations for reducing end-to-end latency # Security & Privacy > 🚧 **Pending API contract** — This page will be written once the API endpoint, event schema, parameters, and limits are finalized. ## Planned coverage - Data processing flow - API key security - Retention policies # Data Retention > 🚧 **Pending API contract** — This page will be written once the API endpoint, event schema, parameters, and limits are finalized. ## Planned coverage - Whether audio and transcripts are retained - Retention duration - How to request deletion # Changelog ## 2026-08-22 ### API paths - **Batch jobs** — preferred enqueue: `POST /v1/audio/transcriptions/jobs` (JSON `audio_url`, **202**). - **Poll** — preferred: `GET /v1/audio/transcriptions/jobs/{workflow_id}` (`cf_` + 64 hex). - **Segment** — `POST /v1/audio/transcriptions` is now documented as **multipart-only** synchronous client-VAD transcription (**200**). - **Deprecation** — `POST /v1/audio/transcriptions` with JSON still works but returns `Deprecation: true` and `Link: `. ### Documentation - New [Segment Transcription API](https://docs.lansonai.com/api-reference/segment-transcription) reference. - [Batch Jobs](https://docs.lansonai.com/api-reference/batch-jobs) split from the old combined transcription page. - [Client-VAD segments guide](https://docs.lansonai.com/guides/client-vad-segments).