---
title: "Stream synthesis"
description: "WebSocket synthesis at wss://tts.navana.ai/v1, returning audio chunks as text arrives."
---

> Documentation Index
> Fetch the complete documentation index at: https://docs.navana.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Stream synthesis

A WebSocket that accepts text frames and returns generated audio chunks as they
are produced.

## Authentication

This endpoint takes **no auth header**. The key travels in the opening `hello`
frame instead, as `auth_token`. See [Authentication](/authentication).

## Protocol

<Diagram
  title="Hello, ready, then text frames interleaved with binary audio, then end and done."
  code={`sequenceDiagram
  participant C as Client
  participant B as Bodhi
  C->>B: upgrade
  C->>B: {"type":"hello", ..., "auth_token": "..."}
  B-->>C: {"type":"ready","sample_rate":24000}
  loop per text frame
    C->>B: {"type":"text","seq":n,"target_text":"..."}
    B-->>C: binary audio chunks
  end
  C->>B: {"type":"end"}
  B-->>C: {"type":"done"}
  B->>C: close`}
/>

### Client to server

#### `hello`

The first frame. It authenticates the connection and locks in language, voice,
and output format for the whole session. None of them can change afterwards.

```json
{
  "type": "hello",
  "lang": "hi",
  "voice": "default_female",
  "output_format": "24000:pcm16",
  "auth_token": "bd_xxxxxxxxxxxx.xxxxxxxx…"
}
```

| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `type` | string | Yes | `"hello"` |
| `auth_token` | string | Yes | Your API key. This is how the connection authenticates. |
| `lang` | string | No | Language code. Defaults to `hi`. See [Voices and languages](/text-to-speech/voices). |
| `voice` | string | No | Voice id for the whole session. Defaults to `default_female` for the language. |
| `output_format` | string | No | `"<sample_rate>:<encoding>"`, for example `"24000:pcm16"`. Defaults to `"24000:float32"`. |
| `num_step` | number | No | Flow-matching steps the engine runs, from `1` to `100`. Omit to use the voice's own default. |
| `g2p_overrides_json` | string | No | Pronunciation overrides for the session, as a **pre-serialized JSON string**. The hello schema takes scalars only, which is why this differs from the non-streaming endpoint's plain `g2p_overrides` object. |
| `client_session_id` | string | No | Your own id for this session. The server records it alongside its own `session_id`, so passing one makes a session traceable from your logs. |

> **Unknown fields are rejected**
>
> The `hello` frame is validated against exactly this set. A field the server
> does not recognise, including a typo like `use_fast_mode`, closes the
> connection with a `protocol_violation` rather than being ignored. That is
> deliberate: a silently dropped knob is harder to debug than a refused
> connection.
>
> Note also that `speed` is a non-streaming field only. It is not a `hello`
> knob, and sending it will be refused.

#### `text`

One per unit of text. You own `seq`, so increment it yourself. The server
splits the text into sentence chunks internally and streams back one audio
message per chunk, and `seq` tags which text frame the replies belong to.

```json
{ "type": "text", "seq": 0, "target_text": "आपके खाते में पाँच हज़ार रुपये हैं।" }
```

#### `end`

```json
{ "type": "end" }
```

### Server to client

| Frame | Shape | Meaning |
| --- | --- | --- |
| `ready` | JSON object, see below | Session is open and the settings are locked in. |
| `audio` | JSON object, see below | Describes the binary frame that comes next. |
| binary | Raw audio bytes | Generated audio in the negotiated format. Chunks concatenate with no seams. |
| `done` | JSON object, see below | All audio for the text you sent has been delivered. |
| `error` | `{"type":"error","code":…,"message":…,"fatal":true}` | The session failed. Carries `seq` when the failure belongs to a specific text frame. |

**Audio arrives as a pair of frames.** A JSON `audio` frame describing the
chunk, immediately followed by the binary chunk itself. A client that only
wants the bytes can ignore the metadata, but it has to expect a text frame
between binary ones rather than treating every text frame as an error.

#### `ready`

```json
{
  "type": "ready",
  "session_id": "…",
  "protocol_version": 2,
  "sample_rate": 24000,
  "encoding": "pcm16",
  "bytes_per_sample": 2,
  "max_text_len": 10000,
  "idle_timeout_s": 180
}
```

This frame is worth reading rather than skipping past. It carries four things
you would otherwise have to guess:

| Field | Why it matters |
| --- | --- |
| `sample_rate`, `encoding` | What the engine settled on. Trust these over what you asked for when sizing buffers or writing a WAV header. |
| `bytes_per_sample` | Lets you convert a byte count to a duration without a table of your own. |
| `max_text_len` | The per-frame character ceiling. Exceed it and the frame is refused with `text_too_long`. |
| `idle_timeout_s` | How long the connection may sit with no frame before the server closes it. |
| `session_id` | The server's id for this session. Log it. |

#### `audio`

```json
{
  "type": "audio",
  "seq": 0,
  "chunk_index": 1,
  "chunk_total": 3,
  "is_last_chunk": false,
  "bytes": 48000,
  "duration_ms": 1000,
  "inference_ms": 180
}
```

| Field | Description |
| --- | --- |
| `seq` | The `text` frame this chunk belongs to, echoing the `seq` you sent. |
| `chunk_index`, `chunk_total` | Position of this chunk within that text frame's audio. The server splits long text into sentence chunks itself. |
| `is_last_chunk` | `true` on the final chunk for that `seq`. |
| `bytes`, `duration_ms` | Size and playing time of the binary frame that follows. |
| `inference_ms` | How long the engine took to generate this chunk, for latency tracking. |

#### `done`

```json
{
  "type": "done",
  "session_id": "…",
  "user_requests": 1,
  "internal_chunks": 3,
  "total_audio_bytes": 144000,
  "total_inference_ms": 540,
  "session_duration_ms": 1200
}
```

`user_requests` counts the `text` frames you sent, `internal_chunks` the
chunks the server split them into. The rest are session totals, useful for
logging without instrumenting your own side.

## Example

```python
import asyncio, json, os
from websockets.asyncio.client import connect

async def speak(sentences, lang="hi", voice="default_female", rate=24000):
async with connect("wss://tts.navana.ai/v1") as socket:
    await socket.send(json.dumps({
        "type": "hello",
        "lang": lang,
        "voice": voice,
        "output_format": f"{rate}:pcm16",
        "auth_token": os.environ["BODHI_API_KEY"],
    }))

    ready = json.loads(await socket.recv())
    sample_rate = ready["sample_rate"]

    for seq, sentence in enumerate(sentences):
        await socket.send(json.dumps({
            "type": "text", "seq": seq, "target_text": sentence,
        }))
    await socket.send(json.dumps({"type": "end"}))

    audio = bytearray()
    async for message in socket:
        if isinstance(message, (bytes, bytearray)):
            audio += message
            continue
        frame = json.loads(message)
        if frame["type"] == "error":
            raise RuntimeError(frame)
        if frame["type"] == "done":
            break

    return bytes(audio), sample_rate
```

```ts
import WebSocket from 'ws'

export function speak(
  sentences: string[],
  { lang = 'hi', voice = 'default_female', rate = 24000 } = {},
): Promise<{ audio: Buffer; sampleRate: number }> {
  const socket = new WebSocket('wss://tts.navana.ai/v1')
  const chunks: Buffer[] = []
  let sampleRate = rate

  return new Promise((resolve, reject) => {
socket.on('error', reject)

socket.on('open', () => {
  socket.send(JSON.stringify({
    type: 'hello',
    lang,
    voice,
    output_format: `${rate}:pcm16`,
    auth_token: process.env.BODHI_API_KEY,
  }))
})

socket.on('message', (data, isBinary) => {
  if (isBinary) {
    chunks.push(data as Buffer)
    return
  }
  const frame = JSON.parse(String(data))
  switch (frame.type) {
    case 'ready':
      sampleRate = frame.sample_rate
      sentences.forEach((target_text, seq) =>
        socket.send(JSON.stringify({ type: 'text', seq, target_text })))
      socket.send(JSON.stringify({ type: 'end' }))
      break
    case 'error':
      reject(new Error(JSON.stringify(frame)))
      break
    case 'done':
      socket.close()
      resolve({ audio: Buffer.concat(chunks), sampleRate })
      break
  }
})
  })
}
```

```go
package main

// go get github.com/gorilla/websocket
import (
	"encoding/json"
	"fmt"
	"net/http"
	"os"
	"strconv"

	"github.com/gorilla/websocket"
)

func speak(sentences []string, lang, voice string, rate int) ([]byte, int, error) {
	header := http.Header{} // no auth header: the key goes in the hello frame
	conn, _, err := websocket.DefaultDialer.Dial("wss://tts.navana.ai/v1", header)
	if err != nil {
		return nil, 0, err
	}
	defer conn.Close()

	if err := conn.WriteJSON(map[string]any{
		"type":          "hello",
		"lang":          lang,
		"voice":         voice,
		"output_format": strconv.Itoa(rate) + ":pcm16",
		"auth_token":    os.Getenv("BODHI_API_KEY"),
	}); err != nil {
		return nil, 0, err
	}

	var ready struct {
		Type       string `json:"type"`
		SampleRate int    `json:"sample_rate"`
	}
	if err := conn.ReadJSON(&ready); err != nil {
		return nil, 0, err
	}

	for seq, text := range sentences {
		conn.WriteJSON(map[string]any{"type": "text", "seq": seq, "target_text": text})
	}
	conn.WriteJSON(map[string]string{"type": "end"})

	var audio []byte
	for {
		msgType, data, err := conn.ReadMessage()
		if err != nil {
			return nil, 0, err
		}
		if msgType == websocket.BinaryMessage {
			audio = append(audio, data...)
			continue
		}
		var frame struct {
			Type string `json:"type"`
		}
		json.Unmarshal(data, &frame)
		if frame.Type == "error" {
			return nil, 0, fmt.Errorf("server error: %s", data)
		}
		if frame.Type == "done" {
			return audio, ready.SampleRate, nil
		}
	}
}

func main() {
	audio, rate, err := speak(
		[]string{"आपके खाते में पाँच हज़ार रुपये हैं।"}, "hi", "default_female", 24000)
	if err != nil {
		panic(err)
	}
	fmt.Println(len(audio), "bytes at", rate, "Hz")
}
```

## Errors

Failures arrive as an `error` frame, not an HTTP status, because by the time
anything can go wrong the connection is already upgraded. Every error frame is
fatal: the server closes the connection straight after sending it.

```json
{ "type": "error", "code": "text_too_long", "message": "…", "fatal": true, "seq": 3 }
```

| `code` | Cause |
| --- | --- |
| `unauthorized` | The `auth_token` in the `hello` frame was not accepted. Wrong, revoked, out of credit, or missing the `tts:synth` scope all land here. |
| `insufficient_funds` | The balance ran out while the session was open. |
| `protocol_violation` | A frame broke the protocol: wrong first frame type, an unknown `hello` field, an unsupported `lang`, a bad `output_format`. |
| `bad_json` | A text frame was not valid JSON. |
| `text_too_long` | A `text` frame exceeded the `max_text_len` from `ready`. |
| `unknown_voice` | The `voice` id does not exist for that language. |
| `bad_request` | The engine rejected the request for another client-side reason. |
| `idle_timeout` | No `hello` arrived in time, or the session sat with no frame for `idle_timeout_s`. |
| `server_overloaded` | The server is at its session cap. Back off and retry. |
| `triton_5xx` | An upstream failure on our side. |

The close code that follows tells you the category: `1008` for anything you
sent, `1011` for a failure on our side, `1013` for overload, and `1000` for a
clean close such as an idle timeout.

> **Running out of credit mid-session**
>
> Usage is reported as the session runs, not only at the end. When a report
> comes back over budget, the server stops the session at the **next frame
> boundary**: audio already being synthesized finishes and is delivered, and
> the refusal arrives before the next `text` frame is served, as
> `insufficient_funds` followed by a `1008` close.
>
> So a stream never dies mid-word, but it can end a turn or so after the
> balance actually hits zero. Budget headroom rather than running to the floor.

## Notes

- **Billed on the text you send**, the same as non-streaming. Streaming costs
  no more for the same input, it just changes when you receive the audio.
  Usage is reported as the session runs, not only at the end.
- **Buffer to sentence boundaries.** Sending a `text` frame per token produces
  choppy prosody, since the engine synthesizes what you give it. Accumulate
  until a sentence terminator, then send.
- **One voice and language per connection.** Changing either means opening a
  new connection.
- **Trust `ready.sample_rate`** over the rate you asked for when sizing
  playback buffers or writing a WAV header.

## Related

- [Streaming synthesis guide](/text-to-speech/streaming) — Working clients and LLM integration.
- [Synthesize speech](/api-reference/synthesize) — One-shot synthesis instead.
- [Voices and languages](/text-to-speech/voices) — Languages and voice ids.

Source: https://docs.navana.ai/api-reference/synthesize-streaming/index.mdx
