---
title: "Streaming synthesis"
description: "Push text into a WebSocket and receive audio chunks as they are generated, so playback starts before the text is finished."
---

> Documentation Index
> Fetch the complete documentation index at: https://docs.navana.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Streaming synthesis

## What you can build

Streaming synthesis separates how much text you have from when audio starts.
You can begin playing the first sentence while the rest is still being written.

- **Speak an LLM's output as it generates.** Buffer tokens into sentences, push
  each one, and play audio continuously. Time to first sound is one sentence
  rather than one full answer.
- **Long-form narration** without a long wait for the first byte.
- **Live voice agents** where a caller hears a response start immediately.
- **Multi-turn sessions.** One connection carries many text frames, so a whole
  conversation's audio flows over a single socket.

## How it works

One WebSocket to `wss://tts.navana.ai/v1`. You open with a `hello` frame that
locks in the session's language, voice, and format. Then you send `text` frames
as text becomes available, and finish with `end`.

Send the key as an `X-API-Key` header on the upgrade, or as an `auth_token`
field in the `hello` frame, as the examples below do.

<Diagram
  title="Hello, ready, then text frames interleaved with returned audio chunks, then end and done."
  code={`sequenceDiagram
  participant C as Your backend
  participant B as Bodhi
  C->>B: WebSocket upgrade
  C->>B: {"type":"hello", ..., "auth_token":"..."}
  B-->>C: {"type":"ready","sample_rate":24000,...}
  loop per text frame
    C->>B: {"type":"text","seq":n,"target_text":"..."}
    B-->>C: binary audio chunks
  end
  C->>B: {"type":"end"}
  B-->>C: {"type":"done"}
  B->>C: close`}
/>

## The protocol

### 1. Connect and send `hello`

The first frame authenticates the connection and fixes its settings. None of
them can change afterwards.

```python
import asyncio, json, os
from websockets.asyncio.client import connect

async def speak():
async with connect("wss://tts.navana.ai/v1") as socket:
    await socket.send(json.dumps({
        "type": "hello",
        "lang": "hi",
        "voice": "default_female",
        "output_format": "24000:pcm16",
        "auth_token": os.environ["BODHI_API_KEY"],
    }))
    ...
```

| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `type` | string | Yes | `"hello"` |
| `auth_token` | string | Yes | Your API key. |
| `lang` | string | No | Language code for the whole session. Defaults to `hi`. |
| `voice` | string | No | Voice id for the whole session. Defaults to `default_female`. |
| `output_format` | string | No | `"<sample_rate>:<encoding>"`, for example `"24000:pcm16"`. Defaults to `"24000:float32"`. |
| `num_step` | number | No | Flow-matching steps, `1` to `100`. Omit to use the voice's default. |

Only `auth_token` is genuinely required, but send the rest explicitly. The
default format is float32, and a field the server does not recognise closes the
connection rather than being ignored. The full field list is in the
[reference](/api-reference/synthesize-streaming).

The server replies once it is ready:

```json
{
  "type": "ready",
  "session_id": "…",
  "protocol_version": 2,
  "sample_rate": 24000,
  "encoding": "pcm16",
  "bytes_per_sample": 2,
  "max_text_len": 10000,
  "idle_timeout_s": 180
}
```

> **Read ready, do not skip it**
>
> `sample_rate` and `encoding` are what the engine settled on, so size your
> playback buffer and any WAV header from those rather than from what you sent.
> `max_text_len` is the per-frame character ceiling, and `idle_timeout_s` is how
> long the connection may sit with no frame before the server closes it.

### 2. Send `text` frames

One per unit of text. You own `seq`, so increment it yourself. The server
splits your text into sentence chunks internally and streams back one audio
message per chunk, and `seq` tags which text frame the replies belong to.

```json
{ "type": "text", "seq": 0, "target_text": "आपके खाते में पाँच हज़ार रुपये हैं।" }
```

Audio comes back as **binary** WebSocket frames, raw samples in the negotiated
format. Chunks concatenate with no seams, so append or play them directly.

### 3. Send `end`

```json
{ "type": "end" }
```

The server finishes the outstanding audio and sends `{"type": "done"}`, then
closes.

## Frame reference

**You send:**

| Frame | Shape |
| --- | --- |
| `hello` | `{"type":"hello","lang":…,"voice":…,"output_format":…,"auth_token":…}` |
| `text` | `{"type":"text","seq":n,"target_text":"…"}` |
| `end` | `{"type":"end"}` |

**You receive:**

| Frame | Shape | Meaning |
| --- | --- | --- |
| `ready` | `{"type":"ready","sample_rate":…,"encoding":…,"max_text_len":…,"idle_timeout_s":…, …}` | Session open, start sending text. |
| `audio` | `{"type":"audio","seq":…,"chunk_index":…,"chunk_total":…,"is_last_chunk":…,"bytes":…,"duration_ms":…,"inference_ms":…}` | Describes the binary frame that follows. |
| binary | Raw audio bytes | Generated audio in the negotiated format. |
| `done` | `{"type":"done","session_id":…,"user_requests":…,"total_audio_bytes":…, …}` | All audio for the text you sent has been delivered, with session totals. |
| `error` | `{"type":"error", …}` | The session failed. |

Each binary chunk is preceded by an `audio` frame describing it. If you only
want the bytes you can ignore that metadata, but do expect text frames
interleaved with the binary ones rather than treating every text frame as an
error.

## A complete client

```python title="synthesize_stream.py"
import asyncio, json, os
from websockets.asyncio.client import connect

URL = "wss://tts.navana.ai/v1"

async def speak(sentences, lang="hi", voice="default_female", rate=24000):
async with connect(URL) as socket:
    await socket.send(json.dumps({
        "type": "hello",
        "lang": lang,
        "voice": voice,
        "output_format": f"{rate}:pcm16",
        "auth_token": os.environ["BODHI_API_KEY"],
    }))

    ready = json.loads(await socket.recv())
    sample_rate = ready["sample_rate"]

    for seq, sentence in enumerate(sentences):
        await socket.send(json.dumps({
            "type": "text", "seq": seq, "target_text": sentence,
        }))
    await socket.send(json.dumps({"type": "end"}))

    audio = bytearray()
    async for message in socket:
        if isinstance(message, (bytes, bytearray)):
            audio += message
            continue
        frame = json.loads(message)
        if frame["type"] == "error":
            raise RuntimeError(frame)
        if frame["type"] == "done":
            break

    return bytes(audio), sample_rate
```

```ts
import WebSocket from 'ws'

export function speak(
  sentences: string[],
  { lang = 'hi', voice = 'default_female', rate = 24000 } = {},
): Promise<{ audio: Buffer; sampleRate: number }> {
  const socket = new WebSocket('wss://tts.navana.ai/v1')
  const chunks: Buffer[] = []
  let sampleRate = rate

  return new Promise((resolve, reject) => {
socket.on('error', reject)

socket.on('open', () => {
  socket.send(JSON.stringify({
    type: 'hello',
    lang,
    voice,
    output_format: `${rate}:pcm16`,
    auth_token: process.env.BODHI_API_KEY,
  }))
})

socket.on('message', (data, isBinary) => {
  if (isBinary) {
    chunks.push(data as Buffer)
    return
  }
  const frame = JSON.parse(String(data))
  switch (frame.type) {
    case 'ready':
      sampleRate = frame.sample_rate
      sentences.forEach((target_text, seq) =>
        socket.send(JSON.stringify({ type: 'text', seq, target_text })))
      socket.send(JSON.stringify({ type: 'end' }))
      break
    case 'error':
      reject(new Error(JSON.stringify(frame)))
      break
    case 'done':
      socket.close()
      resolve({ audio: Buffer.concat(chunks), sampleRate })
      break
  }
})
  })
}
```

```go
package main

// go get github.com/gorilla/websocket
import (
	"encoding/json"
	"fmt"
	"net/http"
	"os"
	"strconv"

	"github.com/gorilla/websocket"
)

func speak(sentences []string, lang, voice string, rate int) ([]byte, int, error) {
	header := http.Header{} // no auth header: the key goes in the hello frame
	conn, _, err := websocket.DefaultDialer.Dial("wss://tts.navana.ai/v1", header)
	if err != nil {
		return nil, 0, err
	}
	defer conn.Close()

	if err := conn.WriteJSON(map[string]any{
		"type":          "hello",
		"lang":          lang,
		"voice":         voice,
		"output_format": strconv.Itoa(rate) + ":pcm16",
		"auth_token":    os.Getenv("BODHI_API_KEY"),
	}); err != nil {
		return nil, 0, err
	}

	var ready struct {
		Type       string `json:"type"`
		SampleRate int    `json:"sample_rate"`
	}
	if err := conn.ReadJSON(&ready); err != nil {
		return nil, 0, err
	}

	for seq, text := range sentences {
		conn.WriteJSON(map[string]any{"type": "text", "seq": seq, "target_text": text})
	}
	conn.WriteJSON(map[string]string{"type": "end"})

	var audio []byte
	for {
		msgType, data, err := conn.ReadMessage()
		if err != nil {
			return nil, 0, err
		}
		if msgType == websocket.BinaryMessage {
			audio = append(audio, data...)
			continue
		}
		var frame struct {
			Type string `json:"type"`
		}
		json.Unmarshal(data, &frame)
		if frame.Type == "error" {
			return nil, 0, fmt.Errorf("server error: %s", data)
		}
		if frame.Type == "done" {
			return audio, ready.SampleRate, nil
		}
	}
}

func main() {
	audio, rate, err := speak(
		[]string{"आपके खाते में पाँच हज़ार रुपये हैं।"}, "hi", "default_female", 24000)
	if err != nil {
		panic(err)
	}
	fmt.Println(len(audio), "bytes at", rate, "Hz")
}
```

## Important considerations

**Billing is by characters sent, not audio produced.** Closing the socket early
does not refund text you already sent.

**Buffer to sentence boundaries.** Sending a `text` frame per token produces
choppy prosody, since the engine synthesizes what you give it. Accumulate until
a sentence terminator, then send.

**One voice and language per connection.** `hello` fixes both. To change
either, open a new connection.

**Keep credit headroom.** Usage is reported while the session runs, so a
balance that empties mid-call ends the session with an `insufficient_funds`
frame at the next frame boundary. Audio already in flight is delivered, but the
call is over. Budget headroom rather than running to the floor.

## Using the Python SDK

If you are writing Python, the SDK speaks this protocol for you — the `hello`
handshake, the sequence numbers, and the `end` frame:

```bash
pip install bodhi-api-sdk
```

```python
import asyncio, os
from bodhi import BodhiTTSClient

async def main():
client = BodhiTTSClient(api_key=os.environ["BODHI_API_KEY"])

async for chunk in client.stream(
    "आपके खाते में पाँच हज़ार रुपये हैं।",
    lang="hi",
    voice="default_female",
    sample_rate=8000,
    encoding="pcm16",
):
    print(len(chunk))      # feed this to your audio device

asyncio.run(main())
```

Chunks arrive in order, so the first plays while the rest is still being
synthesized. Pass a list of strings instead of one to send several utterances
down the same connection — that is how you speak an LLM's sentences as they
arrive:

```python
async for chunk in client.stream(["पहला वाक्य।", "दूसरा वाक्य।"], lang="hi"):
play(chunk)
```

`stream_to_file(text, "speech.wav", ...)` collects the whole utterance into a
playable file instead. `speed` is deliberately absent: the `hello` frame rejects
it, exactly as described above.

## Errors

There are no HTTP statuses here past the upgrade. Failures arrive as an `error`
frame, and every one of them is fatal: the server closes the connection
immediately after.

```json
{ "type": "error", "code": "text_too_long", "message": "…", "fatal": true, "seq": 3 }
```

The ones worth handling explicitly:

| `code` | What to do |
| --- | --- |
| `unauthorized` | Check `auth_token` is in the `hello` frame and is the whole key. A revoked key or an empty balance also lands here. |
| `insufficient_funds` | The balance ran out mid-session. Top up, then reconnect. |
| `protocol_violation` | You sent something the protocol does not allow, often an unknown `hello` field. Fix the client; retrying will not help. |
| `text_too_long` | Split the text. The ceiling is `max_text_len` from `ready`. |
| `idle_timeout` | The connection sat idle too long. Open a new one when you next have text. |
| `server_overloaded` | Back off and retry. |

The full list is in the
[reference](/api-reference/synthesize-streaming#errors).

## Related

- [API reference](/api-reference/synthesize-streaming) — Exact frame shapes and fields.
- [Non-streaming synthesis](/text-to-speech/non-streaming) — One call, complete audio.
- [Voices and languages](/text-to-speech/voices) — Languages and voice ids.

Source: https://docs.navana.ai/text-to-speech/streaming/index.mdx
