---
title: "Streaming transcription"
description: "Push audio over a WebSocket and read transcripts as the words are still being spoken."
---

> Documentation Index
> Fetch the complete documentation index at: https://docs.navana.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Streaming transcription

## What you can build

Streaming transcription gives you text while the audio is still happening, not
after it ends. Results arrive as `partial` transcripts that refine in place and
settle into `complete` ones.

- **Live captions** on a call or meeting, updating as words are recognized.
- **Real-time agent assist.** Detect an intent or keyword mid-sentence and
  surface a suggestion before the caller finishes talking.
- **Voice agents.** Feed final segments into your own logic and respond while
  the caller is still on the line.
- **Compliance monitoring.** Watch for required or prohibited phrases as they
  are said, rather than in a nightly pass.

## How it works

One WebSocket carries everything. You send a JSON config frame, then binary
audio frames. The server sends JSON transcript frames back on the same
connection until you signal end of audio.

<Diagram
  title="A streaming session: config frame, audio frames, transcript frames, then eof and eos."
  code={`sequenceDiagram
  participant C as Your backend
  participant B as Bodhi
  C->>B: WebSocket upgrade (X-Api-Key)
  C->>B: {"config": {...}}
  loop while audio arrives
    C->>B: binary PCM16 frame
    B-->>C: {"type":"partial", "text": "..."}
    B-->>C: {"type":"complete", "text": "..."}
  end
  C->>B: {"eof": 1}
  B-->>C: final frames, last one has eos: true
  B->>C: close`}
/>

## The protocol

### 1. Connect

```python
import os
from websockets.asyncio.client import connect

async def transcribe():
async with connect(
    "wss://stt.navana.ai",
    additional_headers={"X-Api-Key": os.environ["BODHI_API_KEY"]},
) as socket:
    ...
```

```ts
import WebSocket from 'ws'

const socket = new WebSocket('wss://stt.navana.ai', {
  headers: { 'X-Api-Key': process.env.BODHI_API_KEY! },
})
```
```go
header := http.Header{}
header.Set("X-Api-Key", os.Getenv("BODHI_API_KEY"))

conn, _, err := websocket.DefaultDialer.Dial("wss://stt.navana.ai", header)
```

### 2. Send the config frame

The first message, as JSON text. Nothing is transcribed until it arrives.

```json
{
  "config": {
"sample_rate": 8000,
"transaction_id": "a-uuid-you-generate",
"model": "hi-banking-v2-8khz",
"aux": true
  }
}
```

| Field | Type | Description |
| --- | --- | --- |
| `sample_rate` | number | Sample rate of the PCM frames you will send. Must match your audio. `8000` is what telephone channels carry, and every model is optimised for it. |
| `transaction_id` | string | An id you generate to correlate this session with your own logs. Must be a valid UUID. |
| `model` | string | Which model to transcribe with. See [Models and languages](/speech-to-text/models). |
| `aux` | boolean | `true` to receive segment metadata alongside text. |

The config frame also accepts `hotwords` for context biasing, `parse_number`
for converting spoken numbers to digits, `exclude_partial` to suppress partial
frames, and `endpoint_silence_duration` to tune when a pause ends a segment.
See [Advanced features](/speech-to-text/advanced-features), or the
[API reference](/api-reference/transcribe-streaming) for the full field list.

### 3. Stream audio frames

Binary WebSocket frames carrying **16-bit signed PCM, mono**, at
the sample rate you declared.

There is no required frame size. Smaller frames lower latency, larger ones
reduce overhead, and around 100 ms is a reasonable default. At 8 kHz mono that
is 1600 bytes per frame, since one second of audio is `sample_rate × 2` bytes.

When you are streaming live audio, frames arrive at the rate the caller speaks
and there is nothing to pace. The `sleep` in the example below exists only
because it replays a file, and imitating real-time capture keeps the latency
behaviour representative.

### 4. Read transcripts

```json
{
  "call_id": "0f8c1b2e-4a55-4c7e-9d31-1b6a0e2f7c44",
  "segment_id": 0,
  "eos": false,
  "type": "partial",
  "text": "the words recognised so far",
  "segment_meta": {
"start_time": 0,
"confidence": 0.91,
"tokens": ["the", "words", "recognised", "so", "far"],
"timestamps": [0.2, 0.5, 0.8, 1.1, 1.3]
  }
}
```

`type` is `"partial"` while a segment is still being refined and `"complete"`
once it is final. **Replace on `segment_id`, do not append.** Successive
partials for the same segment are revisions of each other, not additions.

```python
segments: dict[int, str] = {}

async for raw in socket:
frame = json.loads(raw)
if "error" in frame:
    raise RuntimeError(frame)
segments[frame["segment_id"]] = frame["text"]
if frame["type"] == "complete":
    handle_final_segment(frame)
if frame.get("eos"):
    break

transcript = " ".join(segments[k] for k in sorted(segments))
```

`segment_meta.timestamps` gives per-token offsets in seconds, pairing
positionally with `segment_meta.tokens`. A `"complete"` frame also carries
`words`, a per-word confidence breakdown.

### 5. End the call

Send the EOF frame, then stop sending audio. Remaining transcripts flush, and
the last frame carries `eos: true`.

```json
{ "eof": 1 }
```

## A complete client

```python title="stream.py"
import asyncio, json, os, uuid, wave
from websockets.asyncio.client import connect

URL = "wss://stt.navana.ai"
CHUNK_MS = 100

async def transcribe(path: str) -> str:
wf = wave.open(path, "rb")
chunk_frames = int(wf.getframerate() * CHUNK_MS / 1000)

async with connect(
    URL,
    additional_headers={"X-Api-Key": os.environ["BODHI_API_KEY"]},
) as socket:
    await socket.send(json.dumps({
        "config": {
            "sample_rate": wf.getframerate(),
            "transaction_id": str(uuid.uuid4()),
            "model": "hi-banking-v2-8khz",
            "aux": True,
        }
    }))

    segments: dict[int, str] = {}

    async def send_audio():
        while chunk := wf.readframes(chunk_frames):
            await socket.send(chunk)
            await asyncio.sleep(CHUNK_MS / 1000)
        await socket.send(json.dumps({"eof": 1}))

    async def receive():
        async for raw in socket:
            frame = json.loads(raw)
            if "error" in frame:
                raise RuntimeError(frame)
            segments[frame["segment_id"]] = frame["text"]
            if frame.get("eos"):
                return

    await asyncio.gather(send_audio(), receive())
    return " ".join(segments[k] for k in sorted(segments))
```

## Using the Python SDK

If you are writing Python, the SDK handles the framing, the config message, and
reconnection for you. You supply audio and read events.

```bash
pip install bodhi-api-sdk
```

```python
import asyncio, os, wave
from bodhi import BodhiClient, TranscriptionConfig, LiveTranscriptionEvents

async def main():
client = BodhiClient(api_key=os.environ["BODHI_API_KEY"])

async def on_transcript(response):
    kind = "final " if response.type == "complete" else "partial"
    print(kind, response.text)

client.on(LiveTranscriptionEvents.Transcript, on_transcript)

with wave.open("recording.wav", "rb") as wf:
    await client.start_connection(config=TranscriptionConfig(
        model="hi-banking-v2-8khz",
        sample_rate=wf.getframerate(),
    ))

    chunk_bytes = int(wf.getframerate() * wf.getsampwidth() * 0.1)  # 100 ms
    audio = wf.readframes(wf.getnframes())
    for i in range(0, len(audio), chunk_bytes):
        await client.send_audio_stream(audio[i:i + chunk_bytes])
        await asyncio.sleep(0.1)   # pace like live capture

await client.close_connection()

asyncio.run(main())
```

> **Results arrive through listeners, not return values**
>
> `start_connection` and `close_connection` return nothing. Every transcript
> reaches you through the `Transcript` listener, so register one before you
> start sending audio. The SDK logs a warning if you forget.

Other events you can subscribe to are `SpeechStarted`, `UtteranceEnd`, `Error`
and `Close`. Building a voice agent? See [Pipecat](/speech-to-text/pipecat),
which wraps this same client as a pipeline service.

## Important considerations

**You are billed for connected time, not speech time.** Silence costs the same
as speech, so close the socket when you are done.

**Keep credit headroom.** If the balance runs out while a connection is open,
the stream is cut off mid-call. Watch for an error frame and treat the close
that follows it as part of the same failure rather than a second one.

**Wait for `eos`, not a fixed delay.** Sending `eof` does not mean transcripts
stop immediately, since how much audio is still in flight varies. Settle on
whichever comes first, the `eos` frame or the socket closing.

**The connection drops after 15 seconds of silence.** The server resets a 15
second read deadline every time it receives a message from you. If your audio
source can go quiet for longer than that, expect to reconnect rather than
holding the socket open.

**Sample rate must match reality.** The config frame declares it and nothing
verifies it. Declaring 16 kHz while sending 8 kHz audio produces a poor
transcript rather than an error.

## Measuring latency

Streaming latency is the gap between how much audio you have sent and how much
of it has come back as text. Measuring it means tracking two positions.

**`audio_cursor`** is how much audio you have sent, in seconds. Advance it as
you send each chunk.

**`transcript_cursor`** is how far the transcript has got, taken from the last
frame you received:

```python
current_offset = 0
timestamps = message["segment_meta"]["timestamps"]
if timestamps:
current_offset = timestamps[-1]

transcript_cursor = message["segment_meta"]["start_time"] + current_offset
```

The difference between the two gives you three useful numbers: the gap after
the latest segment is processed is your minimum latency, the gap before
processing starts on the current segment is your maximum, and weighting each
segment's midpoint latency by its duration gives you a meaningful average.

> **Measure from somewhere close**
>
> Network round-trip is part of anything you measure this way, and it can
> easily dominate. Run the client from a server in India for a representative
> number. If you need lower latency than a public endpoint can give you, ask
> about on-premise deployment at
> [support@navanatech.in](mailto:support@navanatech.in).

## Errors

Connection-time failures come back as a normal HTTP response on the upgrade, so
you get a readable status rather than a silent close.

| Status | Cause |
| --- | --- |
| `400` | Malformed config frame, invalid `transaction_id`, or unknown model. |
| `401` | Missing or incorrect API key. |
| `402` | Credit balance exhausted. |
| `403` | Account inactive, or the key lacks the required scope. |
| `503` | Service unavailable or temporarily overloaded. |

## Related

- [API reference](/api-reference/transcribe-streaming) — Exact frame shapes and every config field.
- [Non-streaming transcription](/speech-to-text/non-streaming) — When you already have the whole file.
- [Models and languages](/speech-to-text/models) — Pick the right model.

Source: https://docs.navana.ai/speech-to-text/streaming/index.mdx
