Skip to content

Non-streaming synthesis

Send text and get the complete audio back in one call, with its sample rate and duration in the response headers.

What you can build

Non-streaming synthesis is the right tool whenever you already have the full text and want the full audio.

  • IVR and telephony prompts. Synthesize at 8 kHz directly so no resampling step sits between you and the call.
  • Notification audio. Render a message once, cache the bytes, replay them as often as you like.
  • Localised product voice. Same call, ten languages, one key.
  • Pre-rendered content. Generate audio ahead of time and serve it as a static file.

How it works

You post JSON to https://tts.navana.ai/tts/bytes with your API key in the X-API-Key header. The whole text is synthesized before anything returns, so the response arrives once, complete.

The response body is raw audio with no container, plus headers telling you what you got.

Quick example

Request body

Field Type Required Description
text string Yes The text to synthesize. Up to 10,000 characters, counted by Unicode code point.
lang string No Language code. Defaults to hi. See Voices and languages.
voice string No Voice id. Defaults to default_female. See Voices and languages.
output_format string No "<sample_rate>:<encoding>", for example "8000:pcm16". Defaults to "24000:float32".
speed number No Speaking rate, 0.25 to 4.0. Omit to use the voice’s own default.
num_step number No Flow-matching steps, 1 to 100. Omit to use the voice’s default.

Only text is required, but send output_format anyway. The default is 24000:float32, and float32 samples are twice the size of the pcm16 most playback paths expect.

There is no model parameter. Each language maps to a single backend.

Understanding the response

The body is raw audio bytes. Everything you need to interpret them is in the headers:

Header Example Meaning
X-Sample-Rate 8000 Sample rate of the returned audio.
X-Encoding pcm16 Encoding of the returned audio.
X-Audio-Duration-Ms 3480 Duration in milliseconds.

At 16-bit mono, one second is sample_rate × 2 bytes, so you can compute duration yourself if you need to:

duration_seconds = len(pcm) / (rate * 2)

Important considerations

The output has no container. It is raw samples, not a WAV or an MP3. Handing the bytes to an audio player will not work until you prepend a 44-byte WAV header, as the examples above do.

Characters are counted by code point, not byte. नमस्ते bills as six characters rather than the eighteen bytes of its UTF-8 encoding, so Indic scripts are not penalised for their encoding.

Three sample rates are accepted. 8000, 16000, and 24000. Anything else comes back as a 400 naming the unsupported rate.

There is no maximum text length, but nothing returns until all of it is synthesized. Past a few sentences, streaming gives you audio sooner for the same input.

Using the Python SDK

The SDK wraps this endpoint, and writes the WAV header the raw bytes lack:

pip install bodhi-api-sdk
import asyncio, os
from bodhi import BodhiTTSClient

async def main():
    client = BodhiTTSClient(api_key=os.environ["BODHI_API_KEY"])

    speech = await client.synthesize(
        "आपके खाते में पाँच हज़ार रुपये हैं।",
        lang="hi",
        voice="default_female",
        sample_rate=8000,
        encoding="pcm16",
    )
    print(speech.sample_rate, speech.duration_ms)
    speech.save("speech.wav")

asyncio.run(main())

synthesize returns a TTSAudio: .audio is the raw bytes, .sample_rate and .encoding are what the server actually used, and .save(path) writes a playable file. speed, num_step and guidance_scale are keyword arguments with the same meanings as above.

Errors

The body is {"error": "<message>"}.

Status Cause
400 Malformed request, a missing or over-long text, or an unsupported lang, voice, output_format, speed, or num_step. The message names the field.
401 The key was not accepted, for any reason.
500 Unexpected server error.
502 The synthesis backend was unreachable or returned a failure.

A refused key is always a 401 here, whether it is wrong, revoked, out of credit, or missing the tts:synth scope. Transcription splits those across 402 and 403, so do not share one error handler between the two services.

Complete list in Errors and debugging.

Navigation

Type to search…

↑↓ navigate↵ selectEsc close