Quickstart
Get from zero to your first transcription request in five minutes.
This page gets you from zero to your first transcription request.
0. Use an LLM to get started (recommended)
The fastest way is to let an AI assistant read the docs for you. Paste this into ChatGPT, Claude, or your coding agent:
Read https://docs.lansonai.com/llms.txt and guide me to understand or integrate LansonAI's speech transcription based on my goal.
The docs are also available as full-page markdown for agents: llms-full.txt and OpenAPI spec.
1. Try it before you code
- See it live: open this replay link — you will hear the audio and watch the realtime captions appear in sync, exactly what your users will experience.
- Need a sample file? Use our podcast episode (~18 minutes, 34 MB) as
audio_urlfor batch transcription, or download it for realtime testing.
2. Get an API key
Your API key starts with sk- and is provisioned by LansonAI. If you do not have one yet, contact the LansonAI team.
The plaintext key is returned only once at provisioning time. The server stores only a SHA-256 hash. Save it immediately.
3. Batch transcription from audio URL
Submit a public URL, get a workflow id (cf_…), poll for the result:
BASE="https://audio.lansonai.com"
SK="sk-..."
curl -X POST "$BASE/v1/audio/transcriptions/jobs" \
-H "Authorization: Bearer $SK" \
-H "Content-Type: application/json" \
-d '{ "audio_url": "https://cdn.example.com/audio/meeting.wav" }'
Response (202 Accepted):
{
"request_id": "...",
"workflow_id": "cf_...",
"status": "queued",
"poll_endpoint": "GET /v1/audio/transcriptions/jobs/cf_..."
}
Poll:
curl "$BASE/v1/audio/transcriptions/jobs/<workflow_id>" \
-H "Authorization: Bearer $SK"
For client-VAD speech clips (multipart, synchronous 200) see Segment API.
4. Realtime transcription (WebSocket)
For live speech, connect over WebSocket:
# 1) (browser only) get a 60-second session token
RT=$(curl -s -X POST "$BASE/v1/audio/transcriptions/session-token" \
-H "Authorization: Bearer $SK" | jq -r .token)
# 2) Connect
# wss://audio.lansonai.com/v1/audio/transcriptions/stream?access_token=$RT
# 3) Send PCM16LE / 16kHz / mono audio frames
# Text frame: {"type":"input_audio_buffer.append","audio":"<base64>"}
# Or send raw binary PCM frames
# 4) Flush when a speech segment ends
# {"type":"input_audio_buffer.flush"}
# 5) Receive transcription events
# {"type":"conversation.item.input_audio_transcription.completed","text":"..."}
5. Expected output
Offline (complete status)
{
"workflow_id": "...",
"status": "complete",
"output": {
"result": {
"segments": [
{ "id": 0, "start_time": 0.0, "end_time": 3.2, "text": "The weather is nice today", "confidence": 0.95 }
],
"summary": { "total_duration": 120.5, "num_segments": 45 },
"metadata": { "language": "zh", "model": "whisper-large-v3-turbo", "chunk_count": 3 }
}
}
}
Realtime event
{
"type": "conversation.item.input_audio_transcription.completed",
"utterance_index": 0,
"text": "The weather is nice today",
"language": "zh",
"audio_duration_ms": 3200,
"latency_ms": 480
}
Next steps
- Realtime Quickstart — full WebSocket examples
- Transcribe Audio — offline transcription details
- Authentication — auth methods and browser security
Overview
LansonAI is the Voice Context Layer for live speech.
Authentication
API key, session token, and browser-side security.
Resources
