Skip to content

Streaming transcription

Push audio over a WebSocket and read transcripts as the words are still being spoken.

What you can build

Streaming transcription gives you text while the audio is still happening, not after it ends. Results arrive as partial transcripts that refine in place and settle into complete ones.

  • Live captions on a call or meeting, updating as words are recognized.
  • Real-time agent assist. Detect an intent or keyword mid-sentence and surface a suggestion before the caller finishes talking.
  • Voice agents. Feed final segments into your own logic and respond while the caller is still on the line.
  • Compliance monitoring. Watch for required or prohibited phrases as they are said, rather than in a nightly pass.

How it works

One WebSocket carries everything. You send a JSON config frame, then binary audio frames. The server sends JSON transcript frames back on the same connection until you signal end of audio.

The protocol

1. Connect

2. Send the config frame

The first message, as JSON text. Nothing is transcribed until it arrives.

{
  "config": {
    "sample_rate": 8000,
    "transaction_id": "a-uuid-you-generate",
    "model": "hi-banking-v2-8khz",
    "aux": true
  }
}
Field Type Description
sample_rate number Sample rate of the PCM frames you will send. Must match your audio. 8000 is what telephone channels carry, and every model is optimised for it.
transaction_id string An id you generate to correlate this session with your own logs. Must be a valid UUID.
model string Which model to transcribe with. See Models and languages.
aux boolean true to receive segment metadata alongside text.

The config frame also accepts hotwords for context biasing, parse_number for converting spoken numbers to digits, exclude_partial to suppress partial frames, and endpoint_silence_duration to tune when a pause ends a segment. See Advanced features, or the API reference for the full field list.

3. Stream audio frames

Binary WebSocket frames carrying 16-bit signed PCM, mono, at the sample rate you declared.

There is no required frame size. Smaller frames lower latency, larger ones reduce overhead, and around 100 ms is a reasonable default. At 8 kHz mono that is 1600 bytes per frame, since one second of audio is sample_rate × 2 bytes.

When you are streaming live audio, frames arrive at the rate the caller speaks and there is nothing to pace. The sleep in the example below exists only because it replays a file, and imitating real-time capture keeps the latency behaviour representative.

4. Read transcripts

{
  "call_id": "0f8c1b2e-4a55-4c7e-9d31-1b6a0e2f7c44",
  "segment_id": 0,
  "eos": false,
  "type": "partial",
  "text": "the words recognised so far",
  "segment_meta": {
    "start_time": 0,
    "confidence": 0.91,
    "tokens": ["the", "words", "recognised", "so", "far"],
    "timestamps": [0.2, 0.5, 0.8, 1.1, 1.3]
  }
}

type is "partial" while a segment is still being refined and "complete" once it is final. Replace on segment_id, do not append. Successive partials for the same segment are revisions of each other, not additions.

segments: dict[int, str] = {}

async for raw in socket:
    frame = json.loads(raw)
    if "error" in frame:
        raise RuntimeError(frame)
    segments[frame["segment_id"]] = frame["text"]
    if frame["type"] == "complete":
        handle_final_segment(frame)
    if frame.get("eos"):
        break

transcript = " ".join(segments[k] for k in sorted(segments))

segment_meta.timestamps gives per-token offsets in seconds, pairing positionally with segment_meta.tokens. A "complete" frame also carries words, a per-word confidence breakdown.

5. End the call

Send the EOF frame, then stop sending audio. Remaining transcripts flush, and the last frame carries eos: true.

{ "eof": 1 }

A complete client

stream.pypython
import asyncio, json, os, uuid, wave
from websockets.asyncio.client import connect

URL = "wss://stt.navana.ai"
CHUNK_MS = 100

async def transcribe(path: str) -> str:
    wf = wave.open(path, "rb")
    chunk_frames = int(wf.getframerate() * CHUNK_MS / 1000)

    async with connect(
        URL,
        additional_headers={"X-Api-Key": os.environ["BODHI_API_KEY"]},
    ) as socket:
        await socket.send(json.dumps({
            "config": {
                "sample_rate": wf.getframerate(),
                "transaction_id": str(uuid.uuid4()),
                "model": "hi-banking-v2-8khz",
                "aux": True,
            }
        }))

        segments: dict[int, str] = {}

        async def send_audio():
            while chunk := wf.readframes(chunk_frames):
                await socket.send(chunk)
                await asyncio.sleep(CHUNK_MS / 1000)
            await socket.send(json.dumps({"eof": 1}))

        async def receive():
            async for raw in socket:
                frame = json.loads(raw)
                if "error" in frame:
                    raise RuntimeError(frame)
                segments[frame["segment_id"]] = frame["text"]
                if frame.get("eos"):
                    return

        await asyncio.gather(send_audio(), receive())
        return " ".join(segments[k] for k in sorted(segments))

Using the Python SDK

If you are writing Python, the SDK handles the framing, the config message, and reconnection for you. You supply audio and read events.

pip install bodhi-api-sdk
import asyncio, os, wave
from bodhi import BodhiClient, TranscriptionConfig, LiveTranscriptionEvents

async def main():
    client = BodhiClient(api_key=os.environ["BODHI_API_KEY"])

    async def on_transcript(response):
        kind = "final " if response.type == "complete" else "partial"
        print(kind, response.text)

    client.on(LiveTranscriptionEvents.Transcript, on_transcript)

    with wave.open("recording.wav", "rb") as wf:
        await client.start_connection(config=TranscriptionConfig(
            model="hi-banking-v2-8khz",
            sample_rate=wf.getframerate(),
        ))

        chunk_bytes = int(wf.getframerate() * wf.getsampwidth() * 0.1)  # 100 ms
        audio = wf.readframes(wf.getnframes())
        for i in range(0, len(audio), chunk_bytes):
            await client.send_audio_stream(audio[i:i + chunk_bytes])
            await asyncio.sleep(0.1)   # pace like live capture

    await client.close_connection()

asyncio.run(main())

Other events you can subscribe to are SpeechStarted, UtteranceEnd, Error and Close. Building a voice agent? See Pipecat, which wraps this same client as a pipeline service.

Important considerations

You are billed for connected time, not speech time. Silence costs the same as speech, so close the socket when you are done.

Keep credit headroom. If the balance runs out while a connection is open, the stream is cut off mid-call. Watch for an error frame and treat the close that follows it as part of the same failure rather than a second one.

Wait for eos, not a fixed delay. Sending eof does not mean transcripts stop immediately, since how much audio is still in flight varies. Settle on whichever comes first, the eos frame or the socket closing.

The connection drops after 15 seconds of silence. The server resets a 15 second read deadline every time it receives a message from you. If your audio source can go quiet for longer than that, expect to reconnect rather than holding the socket open.

Sample rate must match reality. The config frame declares it and nothing verifies it. Declaring 16 kHz while sending 8 kHz audio produces a poor transcript rather than an error.

Measuring latency

Streaming latency is the gap between how much audio you have sent and how much of it has come back as text. Measuring it means tracking two positions.

audio_cursor is how much audio you have sent, in seconds. Advance it as you send each chunk.

transcript_cursor is how far the transcript has got, taken from the last frame you received:

current_offset = 0
timestamps = message["segment_meta"]["timestamps"]
if timestamps:
    current_offset = timestamps[-1]

transcript_cursor = message["segment_meta"]["start_time"] + current_offset

The difference between the two gives you three useful numbers: the gap after the latest segment is processed is your minimum latency, the gap before processing starts on the current segment is your maximum, and weighting each segment’s midpoint latency by its duration gives you a meaningful average.

Errors

Connection-time failures come back as a normal HTTP response on the upgrade, so you get a readable status rather than a silent close.

Status Cause
400 Malformed config frame, invalid transaction_id, or unknown model.
401 Missing or incorrect API key.
402 Credit balance exhausted.
403 Account inactive, or the key lacks the required scope.
503 Service unavailable or temporarily overloaded.
Navigation

Type to search…

↑↓ navigate↵ selectEsc close