wss://tts.navana.ai/v1A WebSocket that accepts text frames and returns generated audio chunks as they are produced.
Authentication
This endpoint takes no auth header. The key travels in the opening hello
frame instead, as auth_token. See Authentication.
Protocol
sequenceDiagram
participant C as Client
participant B as Bodhi
C->>B: upgrade
C->>B: {"type":"hello", ..., "auth_token": "..."}
B-->>C: {"type":"ready","sample_rate":24000}
loop per text frame
C->>B: {"type":"text","seq":n,"target_text":"..."}
B-->>C: binary audio chunks
end
C->>B: {"type":"end"}
B-->>C: {"type":"done"}
B->>C: close
Client to server
hello
The first frame. It authenticates the connection and locks in language, voice, and output format for the whole session. None of them can change afterwards.
{
"type": "hello",
"lang": "hi",
"voice": "default_female",
"output_format": "24000:pcm16",
"auth_token": "bd_xxxxxxxxxxxx.xxxxxxxx…"
}| Field | Type | Required | Description |
|---|---|---|---|
type |
string | Yes | "hello" |
auth_token |
string | Yes | Your API key. This is how the connection authenticates. |
lang |
string | No | Language code. Defaults to hi. See Voices and languages. |
voice |
string | No | Voice id for the whole session. Defaults to default_female for the language. |
output_format |
string | No | "<sample_rate>:<encoding>", for example "24000:pcm16". Defaults to "24000:float32". |
num_step |
number | No | Flow-matching steps the engine runs, from 1 to 100. Omit to use the voice’s own default. |
g2p_overrides_json |
string | No | Pronunciation overrides for the session, as a pre-serialized JSON string. The hello schema takes scalars only, which is why this differs from the non-streaming endpoint’s plain g2p_overrides object. |
client_session_id |
string | No | Your own id for this session. The server records it alongside its own session_id, so passing one makes a session traceable from your logs. |
text
One per unit of text. You own seq, so increment it yourself. The server
splits the text into sentence chunks internally and streams back one audio
message per chunk, and seq tags which text frame the replies belong to.
{ "type": "text", "seq": 0, "target_text": "आपके खाते में पाँच हज़ार रुपये हैं।" }end
{ "type": "end" }Server to client
| Frame | Shape | Meaning |
|---|---|---|
ready |
JSON object, see below | Session is open and the settings are locked in. |
audio |
JSON object, see below | Describes the binary frame that comes next. |
| binary | Raw audio bytes | Generated audio in the negotiated format. Chunks concatenate with no seams. |
done |
JSON object, see below | All audio for the text you sent has been delivered. |
error |
{"type":"error","code":…,"message":…,"fatal":true} |
The session failed. Carries seq when the failure belongs to a specific text frame. |
Audio arrives as a pair of frames. A JSON audio frame describing the
chunk, immediately followed by the binary chunk itself. A client that only
wants the bytes can ignore the metadata, but it has to expect a text frame
between binary ones rather than treating every text frame as an error.
ready
{
"type": "ready",
"session_id": "…",
"protocol_version": 2,
"sample_rate": 24000,
"encoding": "pcm16",
"bytes_per_sample": 2,
"max_text_len": 10000,
"idle_timeout_s": 180
}This frame is worth reading rather than skipping past. It carries four things you would otherwise have to guess:
| Field | Why it matters |
|---|---|
sample_rate, encoding |
What the engine settled on. Trust these over what you asked for when sizing buffers or writing a WAV header. |
bytes_per_sample |
Lets you convert a byte count to a duration without a table of your own. |
max_text_len |
The per-frame character ceiling. Exceed it and the frame is refused with text_too_long. |
idle_timeout_s |
How long the connection may sit with no frame before the server closes it. |
session_id |
The server’s id for this session. Log it. |
audio
{
"type": "audio",
"seq": 0,
"chunk_index": 1,
"chunk_total": 3,
"is_last_chunk": false,
"bytes": 48000,
"duration_ms": 1000,
"inference_ms": 180
}| Field | Description |
|---|---|
seq |
The text frame this chunk belongs to, echoing the seq you sent. |
chunk_index, chunk_total |
Position of this chunk within that text frame’s audio. The server splits long text into sentence chunks itself. |
is_last_chunk |
true on the final chunk for that seq. |
bytes, duration_ms |
Size and playing time of the binary frame that follows. |
inference_ms |
How long the engine took to generate this chunk, for latency tracking. |
done
{
"type": "done",
"session_id": "…",
"user_requests": 1,
"internal_chunks": 3,
"total_audio_bytes": 144000,
"total_inference_ms": 540,
"session_duration_ms": 1200
}user_requests counts the text frames you sent, internal_chunks the
chunks the server split them into. The rest are session totals, useful for
logging without instrumenting your own side.
Example
import asyncio, json, os
from websockets.asyncio.client import connect
async def speak(sentences, lang="hi", voice="default_female", rate=24000):
async with connect("wss://tts.navana.ai/v1") as socket:
await socket.send(json.dumps({
"type": "hello",
"lang": lang,
"voice": voice,
"output_format": f"{rate}:pcm16",
"auth_token": os.environ["BODHI_API_KEY"],
}))
ready = json.loads(await socket.recv())
sample_rate = ready["sample_rate"]
for seq, sentence in enumerate(sentences):
await socket.send(json.dumps({
"type": "text", "seq": seq, "target_text": sentence,
}))
await socket.send(json.dumps({"type": "end"}))
audio = bytearray()
async for message in socket:
if isinstance(message, (bytes, bytearray)):
audio += message
continue
frame = json.loads(message)
if frame["type"] == "error":
raise RuntimeError(frame)
if frame["type"] == "done":
break
return bytes(audio), sample_rateimport WebSocket from 'ws'
export function speak(
sentences: string[],
{ lang = 'hi', voice = 'default_female', rate = 24000 } = {},
): Promise<{ audio: Buffer; sampleRate: number }> {
const socket = new WebSocket('wss://tts.navana.ai/v1')
const chunks: Buffer[] = []
let sampleRate = rate
return new Promise((resolve, reject) => {
socket.on('error', reject)
socket.on('open', () => {
socket.send(JSON.stringify({
type: 'hello',
lang,
voice,
output_format: `${rate}:pcm16`,
auth_token: process.env.BODHI_API_KEY,
}))
})
socket.on('message', (data, isBinary) => {
if (isBinary) {
chunks.push(data as Buffer)
return
}
const frame = JSON.parse(String(data))
switch (frame.type) {
case 'ready':
sampleRate = frame.sample_rate
sentences.forEach((target_text, seq) =>
socket.send(JSON.stringify({ type: 'text', seq, target_text })))
socket.send(JSON.stringify({ type: 'end' }))
break
case 'error':
reject(new Error(JSON.stringify(frame)))
break
case 'done':
socket.close()
resolve({ audio: Buffer.concat(chunks), sampleRate })
break
}
})
})
}package main
// go get github.com/gorilla/websocket
import (
"encoding/json"
"fmt"
"net/http"
"os"
"strconv"
"github.com/gorilla/websocket"
)
func speak(sentences []string, lang, voice string, rate int) ([]byte, int, error) {
header := http.Header{} // no auth header: the key goes in the hello frame
conn, _, err := websocket.DefaultDialer.Dial("wss://tts.navana.ai/v1", header)
if err != nil {
return nil, 0, err
}
defer conn.Close()
if err := conn.WriteJSON(map[string]any{
"type": "hello",
"lang": lang,
"voice": voice,
"output_format": strconv.Itoa(rate) + ":pcm16",
"auth_token": os.Getenv("BODHI_API_KEY"),
}); err != nil {
return nil, 0, err
}
var ready struct {
Type string `json:"type"`
SampleRate int `json:"sample_rate"`
}
if err := conn.ReadJSON(&ready); err != nil {
return nil, 0, err
}
for seq, text := range sentences {
conn.WriteJSON(map[string]any{"type": "text", "seq": seq, "target_text": text})
}
conn.WriteJSON(map[string]string{"type": "end"})
var audio []byte
for {
msgType, data, err := conn.ReadMessage()
if err != nil {
return nil, 0, err
}
if msgType == websocket.BinaryMessage {
audio = append(audio, data...)
continue
}
var frame struct {
Type string `json:"type"`
}
json.Unmarshal(data, &frame)
if frame.Type == "error" {
return nil, 0, fmt.Errorf("server error: %s", data)
}
if frame.Type == "done" {
return audio, ready.SampleRate, nil
}
}
}
func main() {
audio, rate, err := speak(
[]string{"आपके खाते में पाँच हज़ार रुपये हैं।"}, "hi", "default_female", 24000)
if err != nil {
panic(err)
}
fmt.Println(len(audio), "bytes at", rate, "Hz")
}Errors
Failures arrive as an error frame, not an HTTP status, because by the time
anything can go wrong the connection is already upgraded. Every error frame is
fatal: the server closes the connection straight after sending it.
{ "type": "error", "code": "text_too_long", "message": "…", "fatal": true, "seq": 3 }code |
Cause |
|---|---|
unauthorized |
The auth_token in the hello frame was not accepted. Wrong, revoked, out of credit, or missing the tts:synth scope all land here. |
insufficient_funds |
The balance ran out while the session was open. |
protocol_violation |
A frame broke the protocol: wrong first frame type, an unknown hello field, an unsupported lang, a bad output_format. |
bad_json |
A text frame was not valid JSON. |
text_too_long |
A text frame exceeded the max_text_len from ready. |
unknown_voice |
The voice id does not exist for that language. |
bad_request |
The engine rejected the request for another client-side reason. |
idle_timeout |
No hello arrived in time, or the session sat with no frame for idle_timeout_s. |
server_overloaded |
The server is at its session cap. Back off and retry. |
triton_5xx |
An upstream failure on our side. |
The close code that follows tells you the category: 1008 for anything you
sent, 1011 for a failure on our side, 1013 for overload, and 1000 for a
clean close such as an idle timeout.
Notes
- Billed on the text you send, the same as non-streaming. Streaming costs no more for the same input, it just changes when you receive the audio. Usage is reported as the session runs, not only at the end.
- Buffer to sentence boundaries. Sending a
textframe per token produces choppy prosody, since the engine synthesizes what you give it. Accumulate until a sentence terminator, then send. - One voice and language per connection. Changing either means opening a new connection.
- Trust
ready.sample_rateover the rate you asked for when sizing playback buffers or writing a WAV header.