Transcripts and commit strategies
How-to guide · Assumes you have completed the client-side or server-side streaming guide.
Overview
When transcribing audio, you will receive partial and committed transcripts.
- Partial transcripts - the interim results of the transcription
- Committed transcripts - the final results of the transcription segment that are sent when a “commit” message is received. A session can have multiple committed transcripts.
The commit transcript can optionally contain word-level timestamps. This is only received when the “include timestamps” option is set to true.
Commit strategies
When sending audio chunks via the WebSocket, transcript segments can be committed in two ways: Manual Commit or Voice Activity Detection (VAD).
Manual commit
With the manual commit strategy, you control when to commit transcript segments. This is the strategy that is used by default. Committing a segment will clear the processed accumulated transcript and start a new segment without losing context. Committing every 20-30 seconds is good practice to improve latency. Even if you do not commit manually, the model automatically commits after approximately 36 seconds of accumulated audio.
For best results, commit during silence periods or another logical point like a turn model.
Committing manually several times in a short sequence can degrade model performance.
Sending previous text context
When sending audio for transcription, you can send previous text context alongside the first audio chunk to help the model understand the context of the speech. This is useful in a few scenarios:
- Agent text for conversational AI use cases - Allows the model to more easily understand the context of the conversation and produce better transcriptions.
- Reconnection after a network error - This allows the model to continue transcribing, using the previous text as guidance.
- General contextual information - A short description of what the transcription will be about helps the model understand the context.
Sending previous_text context is only possible when sending the first audio chunk via
connection.send(). Sending it in subsequent chunks will result in an error. Previous text works
best when it’s under 50 characters long.
Voice Activity Detection (VAD)
With the VAD strategy, the transcription engine automatically detects speech and silence segments. When a silence threshold is reached, the transcription engine will commit the transcript segment automatically.
When transcribing audio from the microphone in the client-side integration, it is recommended to use the VAD strategy.
Keeping the connection alive during silence
The official SDKs don’t drop the connection when no messages arrive, so most integrations don’t need this. If your own WebSocket client, or a proxy or load balancer in between, closes the connection when no frames arrive for a while, pass the optional keepalive_interval_ms query parameter when connecting. This matters during long stretches of silence, for example a phone call with 10-15 second pauses. About once per interval, the server sends a keepalive partial_transcript: empty (text: "") if the current segment has no uncommitted text, or a repeat of the latest partial text if it does.
Keepalives are not a free-running ping: you must keep streaming audio (silence frames are fine). If you stop sending audio, no keepalives are sent, and the server closes the connection after 15 seconds without any client messages. This server limit is not configurable.
- Accepts an integer between
500and10000(milliseconds). It is disabled by default; omit the parameter to keep the existing behavior. - Out-of-range or non-integer values cause the server to send an
invalid_requesterror and close the connection. - Keepalives are driven by the model actually processing the silent audio you send, not a free-running timer, so they also confirm the transcription path is alive.
- Audio is processed in roughly 1-second chunks, so keepalives arrive about once per interval rounded to that cadence (for example,
1000fires about once a second,3000about once every 3 seconds). The first keepalive of a session arrives about 2 seconds after it starts, since the server buffers the first ~2 seconds of audio before transcribing. Set your interval to at most about a third of your own read timeout to leave margin. - A keepalive’s
textis only empty when the current segment has no uncommitted text yet. If a pause happens after speech but before a commit — most visible withfilter_background_audio=trueor in manual commit mode without committing — the keepalive repeats the latest partial text instead, so it never blanks interim text. After a commit, keepalives go back to empty until new speech arrives. - Works with both
commit_strategy=manualandcommit_strategy=vad, and withfilter_background_audio=true. Has no billing impact beyond the audio you’re already streaming. - The
session_startedmessage’sconfigechoes backkeepalive_interval_ms(nullwhen disabled).
An empty partial_transcript (text: "") means “no speech in the current segment.” A repeated,
identical partial_transcript during a pause is also a keepalive — clients should simply render
partials as they already do rather than special-casing the repeat.
Add the parameter to the WebSocket URL:
Supported audio formats
Best practices
Audio quality
- For best results, use a 16kHz sample rate for an optimum balance of quality and bandwidth.
- Ensure clean audio input with minimal background noise.
- Use an appropriate microphone gain to avoid clipping.
- Only mono audio is supported at this time.
Chunk size
- Send audio chunks of 0.1 - 1 second in length for smooth streaming.
- Smaller chunks result in lower latency but more overhead.
- Larger chunks are more efficient but can introduce latency.