# LansonAI Docs > Developer documentation for LansonAI, the Voice Context Layer that turns live speech into stable, readable, application-ready context. ## What makes LansonAI different LansonAI is not just a speech-to-text API. Traditional transcription APIs primarily answer: "What words were spoken?" LansonAI also solves what happens after recognition: "How can live speech become stable, readable context that an application can safely consume?" Key differences: - StableStream: separates provisional speech recognition from stable, application-ready text, reducing rewrites and caption jitter. - Context-aware recognition: previous speech and surrounding context participate in recognition instead of being applied only as post-processing. - Voice Context Layer: manages the lifecycle between raw speech recognition and the application consuming the text. - Live translation: translation evolves with live speech while preserving correspondence and stability. - Latency is treated as a user-perceived pipeline, not just model inference time. ## Documentation Sets - [LansonAI Docs - Full Documentation](https://docs.lansonai.com/llms-full.txt): The complete developer documentation for LansonAI's Voice Context Layer API. ## Start - [Overview](https://docs.lansonai.com/raw/start.md): LansonAI is the Voice Context Layer for live speech. - [Quickstart](https://docs.lansonai.com/raw/start/quickstart.md): Get from zero to your first transcription request in five minutes. - [Authentication](https://docs.lansonai.com/raw/start/authentication.md): API key, session token, and browser-side security. - [Choose an API](https://docs.lansonai.com/raw/start/choose-an-api.md): Realtime, Segment, and Batch — when to use which. ## Realtime - [Realtime Overview](https://docs.lansonai.com/raw/realtime.md): What the Realtime API does, input, output, and architecture. - [Realtime Quickstart](https://docs.lansonai.com/raw/realtime/quickstart.md): Minimal runnable WebSocket examples. - [Connection Lifecycle](https://docs.lansonai.com/raw/realtime/connection-lifecycle.md): Connect → configure → stream audio → receive events → close/reconnect. - [Audio Input](https://docs.lansonai.com/raw/realtime/audio-input.md): PCM format, sample rate, channels, chunk size, and latency tradeoffs. - [Transcript Lifecycle](https://docs.lansonai.com/raw/realtime/transcript-lifecycle.md): Speech segment states and event transitions. - [Realtime Translation](https://docs.lansonai.com/raw/realtime/realtime-translation.md): Source and translated text correspondence, language config, switching, latency. - [Session Configuration](https://docs.lansonai.com/raw/realtime/session-configuration.md): language, prompt, VAD, and other session-level options. - [Handling Interruptions & Silence](https://docs.lansonai.com/raw/realtime/interruptions-silence.md): Silence, pauses, VAD end-of-speech behavior, and connection keepalive. ## Recorded - [Recorded Overview](https://docs.lansonai.com/raw/recorded.md): Batch transcription jobs from audio URLs. - [Transcribe Audio](https://docs.lansonai.com/raw/recorded/transcribe-audio.md): Submit an audio URL for batch transcription. - [Timestamps & Speakers](https://docs.lansonai.com/raw/recorded/timestamps-speakers.md): Segment timestamps, confidence, and speaker information. - [Subtitles](https://docs.lansonai.com/raw/recorded/subtitles.md): Generate SRT / VTT subtitle files from transcription results. - [Recorded Translation](https://docs.lansonai.com/raw/recorded/translation.md): File-level translation for recorded audio. - [Structured Results](https://docs.lansonai.com/raw/recorded/structured-results.md): Recorded Speech can return structured outputs alongside the transcript. ## Concepts - [Voice Context Layer](https://docs.lansonai.com/raw/concepts.md): A Voice Context Layer sits between live speech recognition and the application that consumes it. - [Stable vs. Partial Text](https://docs.lansonai.com/raw/concepts/stable-vs-partial-text.md): Live transcription is inherently provisional. Understanding the distinction between partial and stable text is essential for building live speech interfaces. - [StableStream](https://docs.lansonai.com/raw/concepts/stablestream.md): LansonAI's approach to maintaining readable continuity while live speech continues to evolve. - [Context-Aware Processing](https://docs.lansonai.com/raw/concepts/context-aware-processing.md): Speech is ambiguous when interpreted in isolation. Context is part of recognition, not a post-processing step. - [Understanding Latency](https://docs.lansonai.com/raw/concepts/latency.md): There is no single latency number for a live speech system. What users experience as latency is the result of several different stages. - [Live Translation](https://docs.lansonai.com/raw/concepts/live-translation.md): Live translation is different from translating a finished transcript. It must balance responsiveness, linguistic context, translation quality, and stability. - [Trinity Engine](https://docs.lansonai.com/raw/concepts/trinity-engine.md): The shared processing foundation behind LansonAI's real-time voice capabilities. ## Guides - [Browser Live Captions](https://docs.lansonai.com/raw/guides.md): Browser microphone → Lanson → caption UI. - [Server-side Streaming](https://docs.lansonai.com/raw/guides/server-side-streaming.md): Backend pushes an audio stream to LansonAI. - [Build a Stable Caption UI](https://docs.lansonai.com/raw/guides/stable-caption-ui.md): Consume StableStream events without manufacturing caption jitter. - [Build Live Translation](https://docs.lansonai.com/raw/guides/live-translation.md): Bilingual / translated caption UI. - [Save a Transcript](https://docs.lansonai.com/raw/guides/save-transcript.md): Persist the complete transcript after a session ends. - [Reconnect a Live Session](https://docs.lansonai.com/raw/guides/reconnection.md): Network disconnect, retry, and resume design. - [Client-VAD speech segments](https://docs.lansonai.com/raw/guides/client-vad-segments.md): When and how to use synchronous segment transcription. ## Reference - [API Reference Overview](https://docs.lansonai.com/raw/api-reference.md): Base URL, REST/WebSocket endpoints, versioning. - [Admin API](https://docs.lansonai.com/raw/api-reference/admin.md): Internal key management and plan catalog endpoints. - [Realtime API](https://docs.lansonai.com/raw/api-reference/realtime-api.md): WebSocket endpoint for streaming transcription. - [Client Messages](https://docs.lansonai.com/raw/api-reference/client-messages.md): All client → Lanson messages sent over WebSocket. - [Server Events](https://docs.lansonai.com/raw/api-reference/server-events.md): All Lanson → client events pushed over WebSocket. - [Segment Transcription API](https://docs.lansonai.com/raw/api-reference/segment-transcription.md): Synchronous client-VAD speech segment transcription (multipart). - [Batch Transcription Jobs](https://docs.lansonai.com/raw/api-reference/batch-jobs.md): Asynchronous batch transcription from audio_url (CF Workflow). - [Languages & Models](https://docs.lansonai.com/raw/api-reference/languages.md): Supported language and model identifiers. - [Errors](https://docs.lansonai.com/raw/api-reference/errors.md): HTTP status codes, WebSocket error events, and error payload formats. - [Rate Limits](https://docs.lansonai.com/raw/api-reference/rate-limits.md): Connection, request, session, and audio quotas by plan. ## Production - [Production Checklist](https://docs.lansonai.com/raw/production.md): What to check before going to production. - [Reconnection & Retries](https://docs.lansonai.com/raw/production/reconnection-retries.md): WebSocket disconnects, 429, 5xx, backoff. - [Latency Best Practices](https://docs.lansonai.com/raw/production/latency-best-practices.md): Audio chunk size, encoding, region selection. - [Security & Privacy](https://docs.lansonai.com/raw/production/security-privacy.md): Data handling, API key security, retention. - [Data Retention](https://docs.lansonai.com/raw/production/data-retention.md): Audio / transcript storage, duration, deletion. - [Changelog](https://docs.lansonai.com/raw/production/changelog.md): API-breaking changes, new fields, deprecations. ## Notes - This is an agent-friendly documentation index. Every page has a plain-markdown version at https://docs.lansonai.com/raw.md, and the complete docs are in https://docs.lansonai.com/llms-full.txt; the API is also described by the OpenAPI spec at https://audio.lansonai.com/openapi.json. - When helping a user: if they want to understand the system first, explain Voice Context Layer, StableStream, and latency in plain language with examples; if they want to integrate now, help them implement a working transcription (Realtime WebSocket or batch audio URL) using the exact endpoints and payloads from the docs. - Ask the user about their stack and use case before writing code.