Skip to content

Synthesize speech

POST https://tts.navana.ai/tts/bytes, to send text and receive complete audio in one request.

POSThttps://tts.navana.ai/tts/bytes

Synthesizes a block of text in one request and returns the complete audio.

Authentication

X-API-Key: <key>. Required. See Authentication.

Request

Content-Type: application/json

Field Type Required Description
text string Yes The text to synthesize. Up to 10,000 characters, counted by Unicode code point.
lang string No Language code, such as hi. Defaults to hi. See Voices and languages.
voice string No Voice id for this request. Defaults to default_female for the language.
output_format string No "<sample_rate>:<encoding>", for example "8000:pcm16". Sample rate is one of 8000, 16000, 24000; encoding is one of pcm16, float32. Defaults to "24000:float32", so name it explicitly if you want 16-bit PCM.
speed number No Speaking rate, from 0.25 to 4.0. Omit to use the voice’s own default.
num_step number No Flow-matching steps the engine runs, from 1 to 100. Omit to use the voice’s own default.
g2p_overrides object No Pronunciation overrides for this request, as {"word": "tokens"}.

Only text is required. Everything else has a default, and the default output format is 24000:float32 rather than PCM, which is the one to watch.

{
  "text": "आपके खाते में पाँच हज़ार रुपये हैं।",
  "lang": "hi",
  "output_format": "8000:pcm16"
}

Example request

Response

200 OK

The body is raw audio with no container, in whatever encoding you asked for via output_format. It is not a WAV and not an MP3, so handing the bytes straight to an audio player will not work. Prepend a 44-byte WAV header first, using the sample rate from X-Sample-Rate. See Non-streaming synthesis for a worked example.

Response headers

Header Example Description
X-Sample-Rate 8000 Sample rate of the returned audio.
X-Encoding pcm16 Encoding of the returned audio.
X-Audio-Duration-Ms 3480 Duration in milliseconds.
X-Request-Id r-4f2a91c8 Identifier for this request. Returned on every response, including errors.

Correlating with your own logs

Send your own X-Request-Id and it comes back unchanged, which ties an entry in your logs to the same request in ours. It has to be 1 to 64 characters of letters, digits, ., _, : or -; anything else is ignored and the server generates an id instead.

curl -X POST https://tts.navana.ai/tts/bytes \
  -H "X-API-Key: $BODHI_API_KEY" \
  -H "X-Request-Id: order-8842-retry-1" \
  -H "Content-Type: application/json" \
  -d '{"text":"नमस्ते","lang":"hi","output_format":"8000:pcm16"}' \
  --output speech.pcm --dump-header -

Errors

The body is {"error": "<message>"}.

Status Cause
400 Malformed request, a missing or over-long text, or an unsupported lang, voice, output_format, speed, or num_step. The message names the field.
401 The key was not accepted.
500 Unexpected server error.
502 The synthesis backend was unreachable or returned a failure.

Notes

  • Billed on the text you send, counted by Unicode code point. नमस्ते is six characters, not the eighteen bytes of its UTF-8 encoding, so Indic scripts are not penalised for their encoding.
  • A failed call is not billed. Usage is reported only after a 200 with audio in it.
  • Nothing returns until the whole text is synthesized. Past a few sentences, streaming gives you audio sooner for the same input.
  • Synthesize at the rate you need. Asking for 8000 directly avoids a resampling step in a telephony pipeline.
Navigation

Type to search…

↑↓ navigate↵ selectEsc close