What you can build
Streaming transcription gives you text while the audio is still happening, not
after it ends. Results arrive as partial transcripts that refine in place and
settle into complete ones.
- Live captions on a call or meeting, updating as words are recognized.
- Real-time agent assist. Detect an intent or keyword mid-sentence and surface a suggestion before the caller finishes talking.
- Voice agents. Feed final segments into your own logic and respond while the caller is still on the line.
- Compliance monitoring. Watch for required or prohibited phrases as they are said, rather than in a nightly pass.
How it works
One WebSocket carries everything. You send a JSON config frame, then binary audio frames. The server sends JSON transcript frames back on the same connection until you signal end of audio.
sequenceDiagram
participant C as Your backend
participant B as Bodhi
C->>B: WebSocket upgrade (X-Api-Key)
C->>B: {"config": {...}}
loop while audio arrives
C->>B: binary PCM16 frame
B-->>C: {"type":"partial", "text": "..."}
B-->>C: {"type":"complete", "text": "..."}
end
C->>B: {"eof": 1}
B-->>C: final frames, last one has eos: true
B->>C: close
The protocol
1. Connect
import os
from websockets.asyncio.client import connect
async def transcribe():
async with connect(
"wss://stt.navana.ai",
additional_headers={"X-Api-Key": os.environ["BODHI_API_KEY"]},
) as socket:
...import WebSocket from 'ws'
const socket = new WebSocket('wss://stt.navana.ai', {
headers: { 'X-Api-Key': process.env.BODHI_API_KEY! },
})header := http.Header{}
header.Set("X-Api-Key", os.Getenv("BODHI_API_KEY"))
conn, _, err := websocket.DefaultDialer.Dial("wss://stt.navana.ai", header)2. Send the config frame
The first message, as JSON text. Nothing is transcribed until it arrives.
{
"config": {
"sample_rate": 8000,
"transaction_id": "a-uuid-you-generate",
"model": "hi-banking-v2-8khz",
"aux": true
}
}| Field | Type | Description |
|---|---|---|
sample_rate |
number | Sample rate of the PCM frames you will send. Must match your audio. 8000 is what telephone channels carry, and every model is optimised for it. |
transaction_id |
string | An id you generate to correlate this session with your own logs. Must be a valid UUID. |
model |
string | Which model to transcribe with. See Models and languages. |
aux |
boolean | true to receive segment metadata alongside text. |
The config frame also accepts hotwords for context biasing, parse_number
for converting spoken numbers to digits, exclude_partial to suppress partial
frames, and endpoint_silence_duration to tune when a pause ends a segment.
See Advanced features, or the
API reference for the full field list.
3. Stream audio frames
Binary WebSocket frames carrying 16-bit signed PCM, mono, at the sample rate you declared.
There is no required frame size. Smaller frames lower latency, larger ones
reduce overhead, and around 100 ms is a reasonable default. At 8 kHz mono that
is 1600 bytes per frame, since one second of audio is sample_rate × 2 bytes.
When you are streaming live audio, frames arrive at the rate the caller speaks
and there is nothing to pace. The sleep in the example below exists only
because it replays a file, and imitating real-time capture keeps the latency
behaviour representative.
4. Read transcripts
{
"call_id": "0f8c1b2e-4a55-4c7e-9d31-1b6a0e2f7c44",
"segment_id": 0,
"eos": false,
"type": "partial",
"text": "the words recognised so far",
"segment_meta": {
"start_time": 0,
"confidence": 0.91,
"tokens": ["the", "words", "recognised", "so", "far"],
"timestamps": [0.2, 0.5, 0.8, 1.1, 1.3]
}
}type is "partial" while a segment is still being refined and "complete"
once it is final. Replace on segment_id, do not append. Successive
partials for the same segment are revisions of each other, not additions.
segments: dict[int, str] = {}
async for raw in socket:
frame = json.loads(raw)
if "error" in frame:
raise RuntimeError(frame)
segments[frame["segment_id"]] = frame["text"]
if frame["type"] == "complete":
handle_final_segment(frame)
if frame.get("eos"):
break
transcript = " ".join(segments[k] for k in sorted(segments))segment_meta.timestamps gives per-token offsets in seconds, pairing
positionally with segment_meta.tokens. A "complete" frame also carries
words, a per-word confidence breakdown.
5. End the call
Send the EOF frame, then stop sending audio. Remaining transcripts flush, and
the last frame carries eos: true.
{ "eof": 1 }A complete client
import asyncio, json, os, uuid, wave
from websockets.asyncio.client import connect
URL = "wss://stt.navana.ai"
CHUNK_MS = 100
async def transcribe(path: str) -> str:
wf = wave.open(path, "rb")
chunk_frames = int(wf.getframerate() * CHUNK_MS / 1000)
async with connect(
URL,
additional_headers={"X-Api-Key": os.environ["BODHI_API_KEY"]},
) as socket:
await socket.send(json.dumps({
"config": {
"sample_rate": wf.getframerate(),
"transaction_id": str(uuid.uuid4()),
"model": "hi-banking-v2-8khz",
"aux": True,
}
}))
segments: dict[int, str] = {}
async def send_audio():
while chunk := wf.readframes(chunk_frames):
await socket.send(chunk)
await asyncio.sleep(CHUNK_MS / 1000)
await socket.send(json.dumps({"eof": 1}))
async def receive():
async for raw in socket:
frame = json.loads(raw)
if "error" in frame:
raise RuntimeError(frame)
segments[frame["segment_id"]] = frame["text"]
if frame.get("eos"):
return
await asyncio.gather(send_audio(), receive())
return " ".join(segments[k] for k in sorted(segments))Using the Python SDK
If you are writing Python, the SDK handles the framing, the config message, and reconnection for you. You supply audio and read events.
pip install bodhi-api-sdkimport asyncio, os, wave
from bodhi import BodhiClient, TranscriptionConfig, LiveTranscriptionEvents
async def main():
client = BodhiClient(api_key=os.environ["BODHI_API_KEY"])
async def on_transcript(response):
kind = "final " if response.type == "complete" else "partial"
print(kind, response.text)
client.on(LiveTranscriptionEvents.Transcript, on_transcript)
with wave.open("recording.wav", "rb") as wf:
await client.start_connection(config=TranscriptionConfig(
model="hi-banking-v2-8khz",
sample_rate=wf.getframerate(),
))
chunk_bytes = int(wf.getframerate() * wf.getsampwidth() * 0.1) # 100 ms
audio = wf.readframes(wf.getnframes())
for i in range(0, len(audio), chunk_bytes):
await client.send_audio_stream(audio[i:i + chunk_bytes])
await asyncio.sleep(0.1) # pace like live capture
await client.close_connection()
asyncio.run(main())Other events you can subscribe to are SpeechStarted, UtteranceEnd, Error
and Close. Building a voice agent? See Pipecat,
which wraps this same client as a pipeline service.
Important considerations
You are billed for connected time, not speech time. Silence costs the same as speech, so close the socket when you are done.
Keep credit headroom. If the balance runs out while a connection is open, the stream is cut off mid-call. Watch for an error frame and treat the close that follows it as part of the same failure rather than a second one.
Wait for eos, not a fixed delay. Sending eof does not mean transcripts
stop immediately, since how much audio is still in flight varies. Settle on
whichever comes first, the eos frame or the socket closing.
The connection drops after 15 seconds of silence. The server resets a 15 second read deadline every time it receives a message from you. If your audio source can go quiet for longer than that, expect to reconnect rather than holding the socket open.
Sample rate must match reality. The config frame declares it and nothing verifies it. Declaring 16 kHz while sending 8 kHz audio produces a poor transcript rather than an error.
Measuring latency
Streaming latency is the gap between how much audio you have sent and how much of it has come back as text. Measuring it means tracking two positions.
audio_cursor is how much audio you have sent, in seconds. Advance it as
you send each chunk.
transcript_cursor is how far the transcript has got, taken from the last
frame you received:
current_offset = 0
timestamps = message["segment_meta"]["timestamps"]
if timestamps:
current_offset = timestamps[-1]
transcript_cursor = message["segment_meta"]["start_time"] + current_offsetThe difference between the two gives you three useful numbers: the gap after the latest segment is processed is your minimum latency, the gap before processing starts on the current segment is your maximum, and weighting each segment’s midpoint latency by its duration gives you a meaningful average.
Errors
Connection-time failures come back as a normal HTTP response on the upgrade, so you get a readable status rather than a silent close.
| Status | Cause |
|---|---|
400 |
Malformed config frame, invalid transaction_id, or unknown model. |
401 |
Missing or incorrect API key. |
402 |
Credit balance exhausted. |
403 |
Account inactive, or the key lacks the required scope. |
503 |
Service unavailable or temporarily overloaded. |