Skip to content

Stream synthesis

WebSocket synthesis at wss://tts.navana.ai/v1, returning audio chunks as text arrives.

WSSwss://tts.navana.ai/v1

A WebSocket that accepts text frames and returns generated audio chunks as they are produced.

Authentication

This endpoint takes no auth header. The key travels in the opening hello frame instead, as auth_token. See Authentication.

Protocol

Client to server

hello

The first frame. It authenticates the connection and locks in language, voice, and output format for the whole session. None of them can change afterwards.

{
  "type": "hello",
  "lang": "hi",
  "voice": "default_female",
  "output_format": "24000:pcm16",
  "auth_token": "bd_xxxxxxxxxxxx.xxxxxxxx…"
}
Field Type Required Description
type string Yes "hello"
auth_token string Yes Your API key. This is how the connection authenticates.
lang string No Language code. Defaults to hi. See Voices and languages.
voice string No Voice id for the whole session. Defaults to default_female for the language.
output_format string No "<sample_rate>:<encoding>", for example "24000:pcm16". Defaults to "24000:float32".
num_step number No Flow-matching steps the engine runs, from 1 to 100. Omit to use the voice’s own default.
g2p_overrides_json string No Pronunciation overrides for the session, as a pre-serialized JSON string. The hello schema takes scalars only, which is why this differs from the non-streaming endpoint’s plain g2p_overrides object.
client_session_id string No Your own id for this session. The server records it alongside its own session_id, so passing one makes a session traceable from your logs.

text

One per unit of text. You own seq, so increment it yourself. The server splits the text into sentence chunks internally and streams back one audio message per chunk, and seq tags which text frame the replies belong to.

{ "type": "text", "seq": 0, "target_text": "आपके खाते में पाँच हज़ार रुपये हैं।" }

end

{ "type": "end" }

Server to client

Frame Shape Meaning
ready JSON object, see below Session is open and the settings are locked in.
audio JSON object, see below Describes the binary frame that comes next.
binary Raw audio bytes Generated audio in the negotiated format. Chunks concatenate with no seams.
done JSON object, see below All audio for the text you sent has been delivered.
error {"type":"error","code":…,"message":…,"fatal":true} The session failed. Carries seq when the failure belongs to a specific text frame.

Audio arrives as a pair of frames. A JSON audio frame describing the chunk, immediately followed by the binary chunk itself. A client that only wants the bytes can ignore the metadata, but it has to expect a text frame between binary ones rather than treating every text frame as an error.

ready

{
  "type": "ready",
  "session_id": "…",
  "protocol_version": 2,
  "sample_rate": 24000,
  "encoding": "pcm16",
  "bytes_per_sample": 2,
  "max_text_len": 10000,
  "idle_timeout_s": 180
}

This frame is worth reading rather than skipping past. It carries four things you would otherwise have to guess:

Field Why it matters
sample_rate, encoding What the engine settled on. Trust these over what you asked for when sizing buffers or writing a WAV header.
bytes_per_sample Lets you convert a byte count to a duration without a table of your own.
max_text_len The per-frame character ceiling. Exceed it and the frame is refused with text_too_long.
idle_timeout_s How long the connection may sit with no frame before the server closes it.
session_id The server’s id for this session. Log it.

audio

{
  "type": "audio",
  "seq": 0,
  "chunk_index": 1,
  "chunk_total": 3,
  "is_last_chunk": false,
  "bytes": 48000,
  "duration_ms": 1000,
  "inference_ms": 180
}
Field Description
seq The text frame this chunk belongs to, echoing the seq you sent.
chunk_index, chunk_total Position of this chunk within that text frame’s audio. The server splits long text into sentence chunks itself.
is_last_chunk true on the final chunk for that seq.
bytes, duration_ms Size and playing time of the binary frame that follows.
inference_ms How long the engine took to generate this chunk, for latency tracking.

done

{
  "type": "done",
  "session_id": "…",
  "user_requests": 1,
  "internal_chunks": 3,
  "total_audio_bytes": 144000,
  "total_inference_ms": 540,
  "session_duration_ms": 1200
}

user_requests counts the text frames you sent, internal_chunks the chunks the server split them into. The rest are session totals, useful for logging without instrumenting your own side.

Example

Errors

Failures arrive as an error frame, not an HTTP status, because by the time anything can go wrong the connection is already upgraded. Every error frame is fatal: the server closes the connection straight after sending it.

{ "type": "error", "code": "text_too_long", "message": "…", "fatal": true, "seq": 3 }
code Cause
unauthorized The auth_token in the hello frame was not accepted. Wrong, revoked, out of credit, or missing the tts:synth scope all land here.
insufficient_funds The balance ran out while the session was open.
protocol_violation A frame broke the protocol: wrong first frame type, an unknown hello field, an unsupported lang, a bad output_format.
bad_json A text frame was not valid JSON.
text_too_long A text frame exceeded the max_text_len from ready.
unknown_voice The voice id does not exist for that language.
bad_request The engine rejected the request for another client-side reason.
idle_timeout No hello arrived in time, or the session sat with no frame for idle_timeout_s.
server_overloaded The server is at its session cap. Back off and retry.
triton_5xx An upstream failure on our side.

The close code that follows tells you the category: 1008 for anything you sent, 1011 for a failure on our side, 1013 for overload, and 1000 for a clean close such as an idle timeout.

Notes

  • Billed on the text you send, the same as non-streaming. Streaming costs no more for the same input, it just changes when you receive the audio. Usage is reported as the session runs, not only at the end.
  • Buffer to sentence boundaries. Sending a text frame per token produces choppy prosody, since the engine synthesizes what you give it. Accumulate until a sentence terminator, then send.
  • One voice and language per connection. Changing either means opening a new connection.
  • Trust ready.sample_rate over the rate you asked for when sizing playback buffers or writing a WAV header.
Navigation

Type to search…

↑↓ navigate↵ selectEsc close