Skip to content

Streaming synthesis

Push text into a WebSocket and receive audio chunks as they are generated, so playback starts before the text is finished.

What you can build

Streaming synthesis separates how much text you have from when audio starts. You can begin playing the first sentence while the rest is still being written.

  • Speak an LLM’s output as it generates. Buffer tokens into sentences, push each one, and play audio continuously. Time to first sound is one sentence rather than one full answer.
  • Long-form narration without a long wait for the first byte.
  • Live voice agents where a caller hears a response start immediately.
  • Multi-turn sessions. One connection carries many text frames, so a whole conversation’s audio flows over a single socket.

How it works

One WebSocket to wss://tts.navana.ai/v1. You open with a hello frame that locks in the session’s language, voice, and format. Then you send text frames as text becomes available, and finish with end.

Send the key as an X-API-Key header on the upgrade, or as an auth_token field in the hello frame, as the examples below do.

The protocol

1. Connect and send hello

The first frame authenticates the connection and fixes its settings. None of them can change afterwards.

import asyncio, json, os
from websockets.asyncio.client import connect

async def speak():
    async with connect("wss://tts.navana.ai/v1") as socket:
        await socket.send(json.dumps({
            "type": "hello",
            "lang": "hi",
            "voice": "default_female",
            "output_format": "24000:pcm16",
            "auth_token": os.environ["BODHI_API_KEY"],
        }))
        ...
Field Type Required Description
type string Yes "hello"
auth_token string Yes Your API key.
lang string No Language code for the whole session. Defaults to hi.
voice string No Voice id for the whole session. Defaults to default_female.
output_format string No "<sample_rate>:<encoding>", for example "24000:pcm16". Defaults to "24000:float32".
num_step number No Flow-matching steps, 1 to 100. Omit to use the voice’s default.

Only auth_token is genuinely required, but send the rest explicitly. The default format is float32, and a field the server does not recognise closes the connection rather than being ignored. The full field list is in the reference.

The server replies once it is ready:

{
  "type": "ready",
  "session_id": "…",
  "protocol_version": 2,
  "sample_rate": 24000,
  "encoding": "pcm16",
  "bytes_per_sample": 2,
  "max_text_len": 10000,
  "idle_timeout_s": 180
}

2. Send text frames

One per unit of text. You own seq, so increment it yourself. The server splits your text into sentence chunks internally and streams back one audio message per chunk, and seq tags which text frame the replies belong to.

{ "type": "text", "seq": 0, "target_text": "आपके खाते में पाँच हज़ार रुपये हैं।" }

Audio comes back as binary WebSocket frames, raw samples in the negotiated format. Chunks concatenate with no seams, so append or play them directly.

3. Send end

{ "type": "end" }

The server finishes the outstanding audio and sends {"type": "done"}, then closes.

Frame reference

You send:

Frame Shape
hello {"type":"hello","lang":…,"voice":…,"output_format":…,"auth_token":…}
text {"type":"text","seq":n,"target_text":"…"}
end {"type":"end"}

You receive:

Frame Shape Meaning
ready {"type":"ready","sample_rate":…,"encoding":…,"max_text_len":…,"idle_timeout_s":…, …} Session open, start sending text.
audio {"type":"audio","seq":…,"chunk_index":…,"chunk_total":…,"is_last_chunk":…,"bytes":…,"duration_ms":…,"inference_ms":…} Describes the binary frame that follows.
binary Raw audio bytes Generated audio in the negotiated format.
done {"type":"done","session_id":…,"user_requests":…,"total_audio_bytes":…, …} All audio for the text you sent has been delivered, with session totals.
error {"type":"error", …} The session failed.

Each binary chunk is preceded by an audio frame describing it. If you only want the bytes you can ignore that metadata, but do expect text frames interleaved with the binary ones rather than treating every text frame as an error.

A complete client

Important considerations

Billing is by characters sent, not audio produced. Closing the socket early does not refund text you already sent.

Buffer to sentence boundaries. Sending a text frame per token produces choppy prosody, since the engine synthesizes what you give it. Accumulate until a sentence terminator, then send.

One voice and language per connection. hello fixes both. To change either, open a new connection.

Keep credit headroom. Usage is reported while the session runs, so a balance that empties mid-call ends the session with an insufficient_funds frame at the next frame boundary. Audio already in flight is delivered, but the call is over. Budget headroom rather than running to the floor.

Using the Python SDK

If you are writing Python, the SDK speaks this protocol for you — the hello handshake, the sequence numbers, and the end frame:

pip install bodhi-api-sdk
import asyncio, os
from bodhi import BodhiTTSClient

async def main():
    client = BodhiTTSClient(api_key=os.environ["BODHI_API_KEY"])

    async for chunk in client.stream(
        "आपके खाते में पाँच हज़ार रुपये हैं।",
        lang="hi",
        voice="default_female",
        sample_rate=8000,
        encoding="pcm16",
    ):
        print(len(chunk))      # feed this to your audio device

asyncio.run(main())

Chunks arrive in order, so the first plays while the rest is still being synthesized. Pass a list of strings instead of one to send several utterances down the same connection — that is how you speak an LLM’s sentences as they arrive:

async for chunk in client.stream(["पहला वाक्य।", "दूसरा वाक्य।"], lang="hi"):
    play(chunk)

stream_to_file(text, "speech.wav", ...) collects the whole utterance into a playable file instead. speed is deliberately absent: the hello frame rejects it, exactly as described above.

Errors

There are no HTTP statuses here past the upgrade. Failures arrive as an error frame, and every one of them is fatal: the server closes the connection immediately after.

{ "type": "error", "code": "text_too_long", "message": "…", "fatal": true, "seq": 3 }

The ones worth handling explicitly:

code What to do
unauthorized Check auth_token is in the hello frame and is the whole key. A revoked key or an empty balance also lands here.
insufficient_funds The balance ran out mid-session. Top up, then reconnect.
protocol_violation You sent something the protocol does not allow, often an unknown hello field. Fix the client; retrying will not help.
text_too_long Split the text. The ceiling is max_text_len from ready.
idle_timeout The connection sat idle too long. Open a new one when you next have text.
server_overloaded Back off and retry.

The full list is in the reference.

Navigation

Type to search…

↑↓ navigate↵ selectEsc close