What you can build
Non-streaming synthesis is the right tool whenever you already have the full text and want the full audio.
- IVR and telephony prompts. Synthesize at 8 kHz directly so no resampling step sits between you and the call.
- Notification audio. Render a message once, cache the bytes, replay them as often as you like.
- Localised product voice. Same call, ten languages, one key.
- Pre-rendered content. Generate audio ahead of time and serve it as a static file.
How it works
You post JSON to https://tts.navana.ai/tts/bytes with your API key in the
X-API-Key header. The whole text is synthesized before anything returns, so
the response arrives once, complete.
The response body is raw audio with no container, plus headers telling you what you got.
Quick example
curl -X POST https://tts.navana.ai/tts/bytes \
-H "X-API-Key: $BODHI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"text": "आपके खाते में पाँच हज़ार रुपये हैं।",
"lang": "hi",
"output_format": "8000:pcm16"
}' \
--output speech.pcm --dump-header -import os, wave, requests
response = requests.post(
"https://tts.navana.ai/tts/bytes",
headers={
"X-API-Key": os.environ["BODHI_API_KEY"],
"Content-Type": "application/json",
},
json={
"text": "आपके खाते में पाँच हज़ार रुपये हैं।",
"lang": "hi",
"output_format": "8000:pcm16",
},
timeout=120,
)
response.raise_for_status()
rate = int(response.headers["X-Sample-Rate"])
# Raw PCM is not playable on its own, so wrap it in a WAV container.
with wave.open("speech.wav", "wb") as f:
f.setnchannels(1)
f.setsampwidth(2)
f.setframerate(rate)
f.writeframes(response.content)const res = await fetch('https://tts.navana.ai/tts/bytes', {
method: 'POST',
headers: {
'X-API-Key': process.env.BODHI_API_KEY!,
'Content-Type': 'application/json',
},
body: JSON.stringify({
text: 'आपके खाते में पाँच हज़ार रुपये हैं।',
lang: 'hi',
output_format: '8000:pcm16',
}),
})
if (!res.ok) throw new Error(await res.text())
const pcm = new Uint8Array(await res.arrayBuffer())
const sampleRate = Number(res.headers.get('X-Sample-Rate'))package main
import (
"bytes"
"encoding/binary"
"encoding/json"
"io"
"net/http"
"os"
"strconv"
)
func main() {
body, _ := json.Marshal(map[string]any{
"text": "आपके खाते में पाँच हज़ार रुपये हैं।",
"lang": "hi",
"output_format": "8000:pcm16",
})
req, _ := http.NewRequest(http.MethodPost,
"https://tts.navana.ai/tts/bytes", bytes.NewReader(body))
req.Header.Set("X-API-Key", os.Getenv("BODHI_API_KEY"))
req.Header.Set("Content-Type", "application/json")
resp, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer resp.Body.Close()
pcm, _ := io.ReadAll(resp.Body)
rate, _ := strconv.Atoi(resp.Header.Get("X-Sample-Rate"))
// The body is raw PCM. Wrap it before anything can play it.
if err := writeWAV("speech.wav", pcm, rate); err != nil {
panic(err)
}
}
// writeWAV prepends the 44-byte canonical header for 16-bit mono PCM.
func writeWAV(path string, pcm []byte, sampleRate int) error {
f, err := os.Create(path)
if err != nil {
return err
}
defer f.Close()
var h bytes.Buffer
h.WriteString("RIFF")
binary.Write(&h, binary.LittleEndian, uint32(36+len(pcm)))
h.WriteString("WAVEfmt ")
binary.Write(&h, binary.LittleEndian, uint32(16)) // fmt chunk size
binary.Write(&h, binary.LittleEndian, uint16(1)) // PCM
binary.Write(&h, binary.LittleEndian, uint16(1)) // mono
binary.Write(&h, binary.LittleEndian, uint32(sampleRate)) // sample rate
binary.Write(&h, binary.LittleEndian, uint32(sampleRate*2)) // byte rate
binary.Write(&h, binary.LittleEndian, uint16(2)) // block align
binary.Write(&h, binary.LittleEndian, uint16(16)) // bits per sample
h.WriteString("data")
binary.Write(&h, binary.LittleEndian, uint32(len(pcm)))
if _, err := f.Write(h.Bytes()); err != nil {
return err
}
_, err = f.Write(pcm)
return err
}Request body
| Field | Type | Required | Description |
|---|---|---|---|
text |
string | Yes | The text to synthesize. Up to 10,000 characters, counted by Unicode code point. |
lang |
string | No | Language code. Defaults to hi. See Voices and languages. |
voice |
string | No | Voice id. Defaults to default_female. See Voices and languages. |
output_format |
string | No | "<sample_rate>:<encoding>", for example "8000:pcm16". Defaults to "24000:float32". |
speed |
number | No | Speaking rate, 0.25 to 4.0. Omit to use the voice’s own default. |
num_step |
number | No | Flow-matching steps, 1 to 100. Omit to use the voice’s default. |
Only text is required, but send output_format anyway. The default is
24000:float32, and float32 samples are twice the size of the pcm16 most
playback paths expect.
There is no model parameter. Each language maps to a single backend.
Understanding the response
The body is raw audio bytes. Everything you need to interpret them is in the headers:
| Header | Example | Meaning |
|---|---|---|
X-Sample-Rate |
8000 |
Sample rate of the returned audio. |
X-Encoding |
pcm16 |
Encoding of the returned audio. |
X-Audio-Duration-Ms |
3480 |
Duration in milliseconds. |
At 16-bit mono, one second is sample_rate × 2 bytes, so you can compute
duration yourself if you need to:
duration_seconds = len(pcm) / (rate * 2)Important considerations
The output has no container. It is raw samples, not a WAV or an MP3. Handing the bytes to an audio player will not work until you prepend a 44-byte WAV header, as the examples above do.
Characters are counted by code point, not byte. नमस्ते bills as six
characters rather than the eighteen bytes of its UTF-8 encoding, so Indic
scripts are not penalised for their encoding.
Three sample rates are accepted. 8000, 16000, and 24000. Anything
else comes back as a 400 naming the unsupported rate.
There is no maximum text length, but nothing returns until all of it is synthesized. Past a few sentences, streaming gives you audio sooner for the same input.
Using the Python SDK
The SDK wraps this endpoint, and writes the WAV header the raw bytes lack:
pip install bodhi-api-sdkimport asyncio, os
from bodhi import BodhiTTSClient
async def main():
client = BodhiTTSClient(api_key=os.environ["BODHI_API_KEY"])
speech = await client.synthesize(
"आपके खाते में पाँच हज़ार रुपये हैं।",
lang="hi",
voice="default_female",
sample_rate=8000,
encoding="pcm16",
)
print(speech.sample_rate, speech.duration_ms)
speech.save("speech.wav")
asyncio.run(main())synthesize returns a TTSAudio: .audio is the raw bytes, .sample_rate and
.encoding are what the server actually used, and .save(path) writes a
playable file. speed, num_step and guidance_scale are keyword arguments
with the same meanings as above.
Errors
The body is {"error": "<message>"}.
| Status | Cause |
|---|---|
400 |
Malformed request, a missing or over-long text, or an unsupported lang, voice, output_format, speed, or num_step. The message names the field. |
401 |
The key was not accepted, for any reason. |
500 |
Unexpected server error. |
502 |
The synthesis backend was unreachable or returned a failure. |
A refused key is always a 401 here, whether it is wrong, revoked, out of
credit, or missing the tts:synth scope. Transcription splits those across
402 and 403, so do not share one error handler between the two services.
Complete list in Errors and debugging.