https://tts.navana.ai/tts/bytesSynthesizes a block of text in one request and returns the complete audio.
Authentication
X-API-Key: <key>. Required. See Authentication.
Request
Content-Type: application/json
| Field | Type | Required | Description |
|---|---|---|---|
text |
string | Yes | The text to synthesize. Up to 10,000 characters, counted by Unicode code point. |
lang |
string | No | Language code, such as hi. Defaults to hi. See Voices and languages. |
voice |
string | No | Voice id for this request. Defaults to default_female for the language. |
output_format |
string | No | "<sample_rate>:<encoding>", for example "8000:pcm16". Sample rate is one of 8000, 16000, 24000; encoding is one of pcm16, float32. Defaults to "24000:float32", so name it explicitly if you want 16-bit PCM. |
speed |
number | No | Speaking rate, from 0.25 to 4.0. Omit to use the voice’s own default. |
num_step |
number | No | Flow-matching steps the engine runs, from 1 to 100. Omit to use the voice’s own default. |
g2p_overrides |
object | No | Pronunciation overrides for this request, as {"word": "tokens"}. |
Only text is required. Everything else has a default, and the default output
format is 24000:float32 rather than PCM, which is the one to watch.
{
"text": "आपके खाते में पाँच हज़ार रुपये हैं।",
"lang": "hi",
"output_format": "8000:pcm16"
}Example request
curl -X POST https://tts.navana.ai/tts/bytes \
-H "X-API-Key: $BODHI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"text":"आपके खाते में पाँच हज़ार रुपये हैं।","lang":"hi","output_format":"8000:pcm16"}' \
--output speech.pcm --dump-header -import os, wave, requests
response = requests.post(
"https://tts.navana.ai/tts/bytes",
headers={
"X-API-Key": os.environ["BODHI_API_KEY"],
"Content-Type": "application/json",
},
json={
"text": "आपके खाते में पाँच हज़ार रुपये हैं।",
"lang": "hi",
"output_format": "8000:pcm16",
},
timeout=120,
)
response.raise_for_status()
rate = int(response.headers["X-Sample-Rate"])
# The body is raw PCM. Wrap it before anything can play it.
with wave.open("speech.wav", "wb") as f:
f.setnchannels(1)
f.setsampwidth(2)
f.setframerate(rate)
f.writeframes(response.content)const res = await fetch('https://tts.navana.ai/tts/bytes', {
method: 'POST',
headers: {
'X-API-Key': process.env.BODHI_API_KEY!,
'Content-Type': 'application/json',
},
body: JSON.stringify({
text: 'आपके खाते में पाँच हज़ार रुपये हैं।',
lang: 'hi',
output_format: '8000:pcm16',
}),
})
if (!res.ok) throw new Error(await res.text())
const pcm = new Uint8Array(await res.arrayBuffer())
const sampleRate = Number(res.headers.get('X-Sample-Rate'))package main
import (
"bytes"
"encoding/binary"
"encoding/json"
"io"
"net/http"
"os"
"strconv"
)
func main() {
body, _ := json.Marshal(map[string]any{
"text": "आपके खाते में पाँच हज़ार रुपये हैं।",
"lang": "hi",
"output_format": "8000:pcm16",
})
req, _ := http.NewRequest(http.MethodPost,
"https://tts.navana.ai/tts/bytes", bytes.NewReader(body))
req.Header.Set("X-API-Key", os.Getenv("BODHI_API_KEY"))
req.Header.Set("Content-Type", "application/json")
resp, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer resp.Body.Close()
pcm, _ := io.ReadAll(resp.Body)
rate, _ := strconv.Atoi(resp.Header.Get("X-Sample-Rate"))
// The body is raw PCM. Wrap it before anything can play it.
if err := writeWAV("speech.wav", pcm, rate); err != nil {
panic(err)
}
}
// writeWAV prepends the 44-byte canonical header for 16-bit mono PCM.
func writeWAV(path string, pcm []byte, sampleRate int) error {
f, err := os.Create(path)
if err != nil {
return err
}
defer f.Close()
var h bytes.Buffer
h.WriteString("RIFF")
binary.Write(&h, binary.LittleEndian, uint32(36+len(pcm)))
h.WriteString("WAVEfmt ")
binary.Write(&h, binary.LittleEndian, uint32(16)) // fmt chunk size
binary.Write(&h, binary.LittleEndian, uint16(1)) // PCM
binary.Write(&h, binary.LittleEndian, uint16(1)) // mono
binary.Write(&h, binary.LittleEndian, uint32(sampleRate)) // sample rate
binary.Write(&h, binary.LittleEndian, uint32(sampleRate*2)) // byte rate
binary.Write(&h, binary.LittleEndian, uint16(2)) // block align
binary.Write(&h, binary.LittleEndian, uint16(16)) // bits per sample
h.WriteString("data")
binary.Write(&h, binary.LittleEndian, uint32(len(pcm)))
if _, err := f.Write(h.Bytes()); err != nil {
return err
}
_, err = f.Write(pcm)
return err
}Response
200 OK
The body is raw audio with no container, in whatever encoding you asked for
via output_format. It is not a WAV and not an MP3, so handing the bytes
straight to an audio player will not work. Prepend a 44-byte WAV header first,
using the sample rate from X-Sample-Rate. See
Non-streaming synthesis for a worked example.
Response headers
| Header | Example | Description |
|---|---|---|
X-Sample-Rate |
8000 |
Sample rate of the returned audio. |
X-Encoding |
pcm16 |
Encoding of the returned audio. |
X-Audio-Duration-Ms |
3480 |
Duration in milliseconds. |
X-Request-Id |
r-4f2a91c8 |
Identifier for this request. Returned on every response, including errors. |
Correlating with your own logs
Send your own X-Request-Id and it comes back unchanged, which ties an entry
in your logs to the same request in ours. It has to be 1 to 64 characters of
letters, digits, ., _, : or -; anything else is ignored and the server
generates an id instead.
curl -X POST https://tts.navana.ai/tts/bytes \
-H "X-API-Key: $BODHI_API_KEY" \
-H "X-Request-Id: order-8842-retry-1" \
-H "Content-Type: application/json" \
-d '{"text":"नमस्ते","lang":"hi","output_format":"8000:pcm16"}' \
--output speech.pcm --dump-header -Errors
The body is {"error": "<message>"}.
| Status | Cause |
|---|---|
400 |
Malformed request, a missing or over-long text, or an unsupported lang, voice, output_format, speed, or num_step. The message names the field. |
401 |
The key was not accepted. |
500 |
Unexpected server error. |
502 |
The synthesis backend was unreachable or returned a failure. |
Notes
- Billed on the text you send, counted by Unicode code point.
नमस्तेis six characters, not the eighteen bytes of its UTF-8 encoding, so Indic scripts are not penalised for their encoding. - A failed call is not billed. Usage is reported only after a
200with audio in it. - Nothing returns until the whole text is synthesized. Past a few sentences, streaming gives you audio sooner for the same input.
- Synthesize at the rate you need. Asking for
8000directly avoids a resampling step in a telephony pipeline.