---
title: "Introduction"
description: "Speech-to-text and text-to-speech for Indian languages, streaming or non-streaming, metered behind a single API key."
---

> Documentation Index
> Fetch the complete documentation index at: https://docs.navana.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Introduction

Bodhi is a speech API platform built for Indian languages. It does two things,
transcription and synthesis, and each is available non-streaming or streaming
behind the same API key.

- **Speech to Text** — Post a recording and get the whole transcript back, or stream audio over a
WebSocket and read partial transcripts as the words are still being spoken.
- **Text to Speech** — Send a block of text and get the complete audio back, or stream text in and
get audio out sentence by sentence while the rest is still being written.

## What you can build

**Transcribe call recordings at scale.** Post a file, get back the transcript
plus per-segment timings and confidence. Route only the low-confidence
segments to a human and accept the rest automatically.

**Live-caption a call while it happens.** Stream microphone or telephony audio
over a WebSocket and read `partial` transcripts as they refine into `complete`
ones. Because partials arrive before a sentence is finished, your own code can
act on what was said without waiting for the caller to stop talking.

**Speak an assistant's reply as it is written.** Push an LLM's output into the
synthesis socket sentence by sentence and play the first one while the model is
still producing the last. The same voice works across every language, so
supporting a new one means changing `lang` and nothing else.

## What makes it different

- **Indian languages, telephony-first** — Models are tuned per language for financial-services vocabulary, meaning
account numbers, balances, and transaction terms. They are trained at
8 kHz, the sample rate phone calls arrive at.
- **Built for code-switching** — Most models are bilingual with English, so code-switching mid-sentence is
handled. Gujarati and Odia are monolingual, English has a dedicated
English-only model, and there is a Hinglish model too.
- **One voice across ten languages** — A synthesis voice is an identity rather than a recording, so the same voice
id reads Hindi, Tamil, or Bengali. No per-language mapping table to keep.
- **One key, metered per use** — One key from the platform authorizes every call, and every call draws down
the same credit balance. No per-request fees and no separate invoice.

## How it fits together

The platform app is where you create an API key and watch what you spend. Your
integration does not call it. It calls the speech APIs directly, and each call
is authorized and metered against your account as it runs.

<Diagram
  title="The platform issues a key and tracks spend; your backend calls the speech API directly with that key."
  code={`flowchart LR
  P[Platform app] -->|issues API key| K[(Your API key)]
  K --> A[Your backend]
  A -->|API key| S[stt.navana.ai]
  A -->|API key| T[tts.navana.ai]
  S -->|usage| P
  T -->|usage| P`}
/>

## Start here

- [Quickstart](/quickstart) — Key to first transcript in about five minutes.
- [Authentication](/authentication) — Where the key goes for each API.
- [Speech to Text](/speech-to-text) — Streaming and non-streaming transcription, models and languages.
- [Text to Speech](/text-to-speech) — Streaming and non-streaming synthesis, voices and languages.
- [API Reference](/api-reference) — Every endpoint, parameter, response field, and error.

Source: https://docs.navana.ai/introduction/index.mdx
