Bodhi is a speech API platform built for Indian languages. It does two things, transcription and synthesis, and each is available non-streaming or streaming behind the same API key.
Speech to Text
Post a recording and get the whole transcript back, or stream audio over a WebSocket and read partial transcripts as the words are still being spoken.
Text to Speech
Send a block of text and get the complete audio back, or stream text in and get audio out sentence by sentence while the rest is still being written.
What you can build
Transcribe call recordings at scale. Post a file, get back the transcript plus per-segment timings and confidence. Route only the low-confidence segments to a human and accept the rest automatically.
Live-caption a call while it happens. Stream microphone or telephony audio
over a WebSocket and read partial transcripts as they refine into complete
ones. Because partials arrive before a sentence is finished, your own code can
act on what was said without waiting for the caller to stop talking.
Speak an assistant’s reply as it is written. Push an LLM’s output into the
synthesis socket sentence by sentence and play the first one while the model is
still producing the last. The same voice works across every language, so
supporting a new one means changing lang and nothing else.
What makes it different
Indian languages, telephony-first
Models are tuned per language for financial-services vocabulary, meaning account numbers, balances, and transaction terms. They are trained at 8 kHz, the sample rate phone calls arrive at.
Built for code-switching
Most models are bilingual with English, so code-switching mid-sentence is handled. Gujarati and Odia are monolingual, English has a dedicated English-only model, and there is a Hinglish model too.
One voice across ten languages
A synthesis voice is an identity rather than a recording, so the same voice id reads Hindi, Tamil, or Bengali. No per-language mapping table to keep.
One key, metered per use
One key from the platform authorizes every call, and every call draws down the same credit balance. No per-request fees and no separate invoice.
How it fits together
The platform app is where you create an API key and watch what you spend. Your integration does not call it. It calls the speech APIs directly, and each call is authorized and metered against your account as it runs.
flowchart LR P[Platform app] -->|issues API key| K[(Your API key)] K --> A[Your backend] A -->|API key| S[stt.navana.ai] A -->|API key| T[tts.navana.ai] S -->|usage| P T -->|usage| P
Start here
Quickstart
Key to first transcript in about five minutes.
Authentication
Where the key goes for each API.
Speech to Text
Streaming and non-streaming transcription, models and languages.
Text to Speech
Streaming and non-streaming synthesis, voices and languages.
API Reference
Every endpoint, parameter, response field, and error.