What you can build
Streaming synthesis separates how much text you have from when audio starts. You can begin playing the first sentence while the rest is still being written.
- Speak an LLM’s output as it generates. Buffer tokens into sentences, push each one, and play audio continuously. Time to first sound is one sentence rather than one full answer.
- Long-form narration without a long wait for the first byte.
- Live voice agents where a caller hears a response start immediately.
- Multi-turn sessions. One connection carries many text frames, so a whole conversation’s audio flows over a single socket.
How it works
One WebSocket to wss://tts.navana.ai/v1. You open with a hello frame that
locks in the session’s language, voice, and format. Then you send text frames
as text becomes available, and finish with end.
Send the key as an X-API-Key header on the upgrade, or as an auth_token
field in the hello frame, as the examples below do.
sequenceDiagram
participant C as Your backend
participant B as Bodhi
C->>B: WebSocket upgrade
C->>B: {"type":"hello", ..., "auth_token":"..."}
B-->>C: {"type":"ready","sample_rate":24000,...}
loop per text frame
C->>B: {"type":"text","seq":n,"target_text":"..."}
B-->>C: binary audio chunks
end
C->>B: {"type":"end"}
B-->>C: {"type":"done"}
B->>C: close
The protocol
1. Connect and send hello
The first frame authenticates the connection and fixes its settings. None of them can change afterwards.
import asyncio, json, os
from websockets.asyncio.client import connect
async def speak():
async with connect("wss://tts.navana.ai/v1") as socket:
await socket.send(json.dumps({
"type": "hello",
"lang": "hi",
"voice": "default_female",
"output_format": "24000:pcm16",
"auth_token": os.environ["BODHI_API_KEY"],
}))
...| Field | Type | Required | Description |
|---|---|---|---|
type |
string | Yes | "hello" |
auth_token |
string | Yes | Your API key. |
lang |
string | No | Language code for the whole session. Defaults to hi. |
voice |
string | No | Voice id for the whole session. Defaults to default_female. |
output_format |
string | No | "<sample_rate>:<encoding>", for example "24000:pcm16". Defaults to "24000:float32". |
num_step |
number | No | Flow-matching steps, 1 to 100. Omit to use the voice’s default. |
Only auth_token is genuinely required, but send the rest explicitly. The
default format is float32, and a field the server does not recognise closes the
connection rather than being ignored. The full field list is in the
reference.
The server replies once it is ready:
{
"type": "ready",
"session_id": "…",
"protocol_version": 2,
"sample_rate": 24000,
"encoding": "pcm16",
"bytes_per_sample": 2,
"max_text_len": 10000,
"idle_timeout_s": 180
}2. Send text frames
One per unit of text. You own seq, so increment it yourself. The server
splits your text into sentence chunks internally and streams back one audio
message per chunk, and seq tags which text frame the replies belong to.
{ "type": "text", "seq": 0, "target_text": "आपके खाते में पाँच हज़ार रुपये हैं।" }Audio comes back as binary WebSocket frames, raw samples in the negotiated format. Chunks concatenate with no seams, so append or play them directly.
3. Send end
{ "type": "end" }The server finishes the outstanding audio and sends {"type": "done"}, then
closes.
Frame reference
You send:
| Frame | Shape |
|---|---|
hello |
{"type":"hello","lang":…,"voice":…,"output_format":…,"auth_token":…} |
text |
{"type":"text","seq":n,"target_text":"…"} |
end |
{"type":"end"} |
You receive:
| Frame | Shape | Meaning |
|---|---|---|
ready |
{"type":"ready","sample_rate":…,"encoding":…,"max_text_len":…,"idle_timeout_s":…, …} |
Session open, start sending text. |
audio |
{"type":"audio","seq":…,"chunk_index":…,"chunk_total":…,"is_last_chunk":…,"bytes":…,"duration_ms":…,"inference_ms":…} |
Describes the binary frame that follows. |
| binary | Raw audio bytes | Generated audio in the negotiated format. |
done |
{"type":"done","session_id":…,"user_requests":…,"total_audio_bytes":…, …} |
All audio for the text you sent has been delivered, with session totals. |
error |
{"type":"error", …} |
The session failed. |
Each binary chunk is preceded by an audio frame describing it. If you only
want the bytes you can ignore that metadata, but do expect text frames
interleaved with the binary ones rather than treating every text frame as an
error.
A complete client
import asyncio, json, os
from websockets.asyncio.client import connect
URL = "wss://tts.navana.ai/v1"
async def speak(sentences, lang="hi", voice="default_female", rate=24000):
async with connect(URL) as socket:
await socket.send(json.dumps({
"type": "hello",
"lang": lang,
"voice": voice,
"output_format": f"{rate}:pcm16",
"auth_token": os.environ["BODHI_API_KEY"],
}))
ready = json.loads(await socket.recv())
sample_rate = ready["sample_rate"]
for seq, sentence in enumerate(sentences):
await socket.send(json.dumps({
"type": "text", "seq": seq, "target_text": sentence,
}))
await socket.send(json.dumps({"type": "end"}))
audio = bytearray()
async for message in socket:
if isinstance(message, (bytes, bytearray)):
audio += message
continue
frame = json.loads(message)
if frame["type"] == "error":
raise RuntimeError(frame)
if frame["type"] == "done":
break
return bytes(audio), sample_rateimport WebSocket from 'ws'
export function speak(
sentences: string[],
{ lang = 'hi', voice = 'default_female', rate = 24000 } = {},
): Promise<{ audio: Buffer; sampleRate: number }> {
const socket = new WebSocket('wss://tts.navana.ai/v1')
const chunks: Buffer[] = []
let sampleRate = rate
return new Promise((resolve, reject) => {
socket.on('error', reject)
socket.on('open', () => {
socket.send(JSON.stringify({
type: 'hello',
lang,
voice,
output_format: `${rate}:pcm16`,
auth_token: process.env.BODHI_API_KEY,
}))
})
socket.on('message', (data, isBinary) => {
if (isBinary) {
chunks.push(data as Buffer)
return
}
const frame = JSON.parse(String(data))
switch (frame.type) {
case 'ready':
sampleRate = frame.sample_rate
sentences.forEach((target_text, seq) =>
socket.send(JSON.stringify({ type: 'text', seq, target_text })))
socket.send(JSON.stringify({ type: 'end' }))
break
case 'error':
reject(new Error(JSON.stringify(frame)))
break
case 'done':
socket.close()
resolve({ audio: Buffer.concat(chunks), sampleRate })
break
}
})
})
}package main
// go get github.com/gorilla/websocket
import (
"encoding/json"
"fmt"
"net/http"
"os"
"strconv"
"github.com/gorilla/websocket"
)
func speak(sentences []string, lang, voice string, rate int) ([]byte, int, error) {
header := http.Header{} // no auth header: the key goes in the hello frame
conn, _, err := websocket.DefaultDialer.Dial("wss://tts.navana.ai/v1", header)
if err != nil {
return nil, 0, err
}
defer conn.Close()
if err := conn.WriteJSON(map[string]any{
"type": "hello",
"lang": lang,
"voice": voice,
"output_format": strconv.Itoa(rate) + ":pcm16",
"auth_token": os.Getenv("BODHI_API_KEY"),
}); err != nil {
return nil, 0, err
}
var ready struct {
Type string `json:"type"`
SampleRate int `json:"sample_rate"`
}
if err := conn.ReadJSON(&ready); err != nil {
return nil, 0, err
}
for seq, text := range sentences {
conn.WriteJSON(map[string]any{"type": "text", "seq": seq, "target_text": text})
}
conn.WriteJSON(map[string]string{"type": "end"})
var audio []byte
for {
msgType, data, err := conn.ReadMessage()
if err != nil {
return nil, 0, err
}
if msgType == websocket.BinaryMessage {
audio = append(audio, data...)
continue
}
var frame struct {
Type string `json:"type"`
}
json.Unmarshal(data, &frame)
if frame.Type == "error" {
return nil, 0, fmt.Errorf("server error: %s", data)
}
if frame.Type == "done" {
return audio, ready.SampleRate, nil
}
}
}
func main() {
audio, rate, err := speak(
[]string{"आपके खाते में पाँच हज़ार रुपये हैं।"}, "hi", "default_female", 24000)
if err != nil {
panic(err)
}
fmt.Println(len(audio), "bytes at", rate, "Hz")
}Important considerations
Billing is by characters sent, not audio produced. Closing the socket early does not refund text you already sent.
Buffer to sentence boundaries. Sending a text frame per token produces
choppy prosody, since the engine synthesizes what you give it. Accumulate until
a sentence terminator, then send.
One voice and language per connection. hello fixes both. To change
either, open a new connection.
Keep credit headroom. Usage is reported while the session runs, so a
balance that empties mid-call ends the session with an insufficient_funds
frame at the next frame boundary. Audio already in flight is delivered, but the
call is over. Budget headroom rather than running to the floor.
Using the Python SDK
If you are writing Python, the SDK speaks this protocol for you — the hello
handshake, the sequence numbers, and the end frame:
pip install bodhi-api-sdkimport asyncio, os
from bodhi import BodhiTTSClient
async def main():
client = BodhiTTSClient(api_key=os.environ["BODHI_API_KEY"])
async for chunk in client.stream(
"आपके खाते में पाँच हज़ार रुपये हैं।",
lang="hi",
voice="default_female",
sample_rate=8000,
encoding="pcm16",
):
print(len(chunk)) # feed this to your audio device
asyncio.run(main())Chunks arrive in order, so the first plays while the rest is still being synthesized. Pass a list of strings instead of one to send several utterances down the same connection — that is how you speak an LLM’s sentences as they arrive:
async for chunk in client.stream(["पहला वाक्य।", "दूसरा वाक्य।"], lang="hi"):
play(chunk)stream_to_file(text, "speech.wav", ...) collects the whole utterance into a
playable file instead. speed is deliberately absent: the hello frame rejects
it, exactly as described above.
Errors
There are no HTTP statuses here past the upgrade. Failures arrive as an error
frame, and every one of them is fatal: the server closes the connection
immediately after.
{ "type": "error", "code": "text_too_long", "message": "…", "fatal": true, "seq": 3 }The ones worth handling explicitly:
code |
What to do |
|---|---|
unauthorized |
Check auth_token is in the hello frame and is the whole key. A revoked key or an empty balance also lands here. |
insufficient_funds |
The balance ran out mid-session. Top up, then reconnect. |
protocol_violation |
You sent something the protocol does not allow, often an unknown hello field. Fix the client; retrying will not help. |
text_too_long |
Split the text. The ceiling is max_text_len from ready. |
idle_timeout |
The connection sat idle too long. Open a new one when you next have text. |
server_overloaded |
Back off and retry. |
The full list is in the reference.