Skip to content

Introduction

Speech-to-text and text-to-speech for Indian languages, streaming or non-streaming, metered behind a single API key.

Bodhi is a speech API platform built for Indian languages. It does two things, transcription and synthesis, and each is available non-streaming or streaming behind the same API key.

Speech to Text

Post a recording and get the whole transcript back, or stream audio over a WebSocket and read partial transcripts as the words are still being spoken.

Text to Speech

Send a block of text and get the complete audio back, or stream text in and get audio out sentence by sentence while the rest is still being written.

What you can build

Transcribe call recordings at scale. Post a file, get back the transcript plus per-segment timings and confidence. Route only the low-confidence segments to a human and accept the rest automatically.

Live-caption a call while it happens. Stream microphone or telephony audio over a WebSocket and read partial transcripts as they refine into complete ones. Because partials arrive before a sentence is finished, your own code can act on what was said without waiting for the caller to stop talking.

Speak an assistant’s reply as it is written. Push an LLM’s output into the synthesis socket sentence by sentence and play the first one while the model is still producing the last. The same voice works across every language, so supporting a new one means changing lang and nothing else.

What makes it different

Indian languages, telephony-first

Models are tuned per language for financial-services vocabulary, meaning account numbers, balances, and transaction terms. They are trained at 8 kHz, the sample rate phone calls arrive at.

Built for code-switching

Most models are bilingual with English, so code-switching mid-sentence is handled. Gujarati and Odia are monolingual, English has a dedicated English-only model, and there is a Hinglish model too.

One voice across ten languages

A synthesis voice is an identity rather than a recording, so the same voice id reads Hindi, Tamil, or Bengali. No per-language mapping table to keep.

One key, metered per use

One key from the platform authorizes every call, and every call draws down the same credit balance. No per-request fees and no separate invoice.

How it fits together

The platform app is where you create an API key and watch what you spend. Your integration does not call it. It calls the speech APIs directly, and each call is authorized and metered against your account as it runs.

Start here

Navigation

Type to search…

↑↓ navigate↵ selectEsc close