What you can build
Non-streaming transcription turns a recording into a transcript in a single
request. With aux turned on you also get per-segment timings and confidence
scores, which supports more than search over text:
- Call analytics. Segment boundaries give you talk time, pauses, and turn structure without extra processing.
- Reviewable transcripts. Surface only the low-confidence segments to a human and accept the rest.
- Playback alignment. Map each segment to a timestamp so a reader can click text and hear the audio.
- Bulk backfill. Transcribe an archive of recordings one file at a time.
How it works
You post an audio file and a model name to https://stt.navana.ai/api/transcribe
with your API key in the X-Api-Key header. The whole file is transcribed
before anything returns, so the response arrives once, complete.
Every request needs a transaction_id that you generate. It has to be a valid
UUID, and it is how you tie a request to your own logs.
Quick example
curl -X POST https://stt.navana.ai/api/transcribe \
-H "X-Api-Key: $BODHI_API_KEY" \
-F "model=hi-banking-v2-8khz" \
-F "transaction_id=$(uuidgen)" \
-F "aux=true" \
-F "audio_file=@recording.wav"import os, uuid, requests
response = requests.post(
"https://stt.navana.ai/api/transcribe",
headers={"X-Api-Key": os.environ["BODHI_API_KEY"]},
data={
"model": "hi-banking-v2-8khz",
"transaction_id": str(uuid.uuid4()),
"aux": "true",
},
files={"audio_file": ("recording.wav", open("recording.wav", "rb"), "audio/wav")},
timeout=120,
)
response.raise_for_status()
result = response.json()
print(result["text"])import { openAsBlob } from 'node:fs'
const form = new FormData()
form.set('model', 'hi-banking-v2-8khz')
form.set('transaction_id', crypto.randomUUID())
form.set('aux', 'true')
form.set('audio_file', await openAsBlob('recording.wav'), 'recording.wav')
const res = await fetch('https://stt.navana.ai/api/transcribe', {
method: 'POST',
headers: { 'X-Api-Key': process.env.BODHI_API_KEY! },
body: form,
})
if (!res.ok) throw new Error(await res.text())
const { text } = await res.json()
console.log(text)package main
import (
"bytes"
"encoding/json"
"fmt"
"io"
"mime/multipart"
"net/http"
"os"
"github.com/google/uuid"
)
func main() {
file, err := os.Open("recording.wav")
if err != nil {
panic(err)
}
defer file.Close()
var body bytes.Buffer
form := multipart.NewWriter(&body)
form.WriteField("model", "hi-banking-v2-8khz")
form.WriteField("transaction_id", uuid.NewString())
form.WriteField("aux", "true")
part, err := form.CreateFormFile("audio_file", "recording.wav")
if err != nil {
panic(err)
}
if _, err := io.Copy(part, file); err != nil {
panic(err)
}
form.Close()
req, _ := http.NewRequest(http.MethodPost,
"https://stt.navana.ai/api/transcribe", &body)
req.Header.Set("X-Api-Key", os.Getenv("BODHI_API_KEY"))
req.Header.Set("Content-Type", form.FormDataContentType())
resp, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer resp.Body.Close()
var out struct {
Text string `json:"text"`
CallID string `json:"call_id"`
Status string `json:"status"`
}
json.NewDecoder(resp.Body).Decode(&out)
fmt.Println(out.Text)
}Understanding the response
{
"text": "बिल्कुल आपकी पूरी सहायता की जाएगी",
"call_id": "f99e79ca-6a2b-42f1-8d6b-fd0855e791de",
"status": "success"
}text is the full transcript and status is success or error. When a
request fails, the body carries error and message instead.
Setting aux=true adds an aux_info object with the segment breakdown,
overall confidence, and server processing time. See the
API reference for every field.
Using confidence
Confidence comes back both overall and per segment. A practical pattern is to accept high-confidence output automatically and queue the rest:
REVIEW_THRESHOLD = 0.75
aux = result.get("aux_info") or {}
segments = aux.get("segments_meta") or []
needs_review = [s for s in segments if s.get("confidence", 1) < REVIEW_THRESHOLD]
if needs_review:
queue_for_human_review(result["call_id"], needs_review)Read aux_info defensively. It is only present when you ask for it, and a
model may not return every field inside it.
Using the Python SDK
The SDK transcribes a file in one call:
pip install bodhi-api-sdkimport asyncio, os
from bodhi import BodhiClient, TranscriptionConfig, LiveTranscriptionEvents
async def main():
client = BodhiClient(api_key=os.environ["BODHI_API_KEY"])
sentences = []
async def on_transcript(response):
if response.type == "complete":
sentences.append(response.text)
client.on(LiveTranscriptionEvents.Transcript, on_transcript)
await client.transcribe_local_file(
"recording.wav",
config=TranscriptionConfig(model="hi-banking-v2-8khz"),
)
print(" ".join(sentences))
asyncio.run(main())There is also transcribe_remote_url for audio the server can fetch itself.
Important considerations
WAV is the safest format. It is read directly. MP3, M4A, FLAC, OGG, and WebM are decoded server-side, so they work too. Models are trained at 8 kHz, which is what telephony delivers, so mono at that rate is what to aim for:
ffmpeg -i input.mp3 -ac 1 -ar 8000 -c:a pcm_s16le output.wavAt 16-bit mono, one second is sample_rate × 2 bytes, so 8 kHz audio runs
about 16 KB per second or 1 MB per minute. Useful for sizing uploads before
you send them.
Match the model to the language. There is no auto-detection. A Hindi model on Tamil audio returns a poor transcript rather than an error. See Models and languages.
transaction_id must be a valid UUID. Anything else is rejected as a 400.
Inverse-text-normalize numbers with parse_number. It converts spoken
numbers into digits, so “पच्चीस लाख” becomes 2500000. It only affects Hindi,
Malayalam, Kannada, Gujarati, and Marathi models, and is ignored silently on
the rest.
Bias recognition toward specific phrases with hotwords. Product names and
domain jargon are the usual case. See
Advanced features.
Long files take time. Nothing returns until the whole file is transcribed, so set a generous client timeout.
Errors
| Status | Cause |
|---|---|
400 |
Malformed request, invalid transaction_id, or unknown model. |
401 |
Missing or incorrect API key. |
402 |
Credit balance exhausted. Top up before retrying. |
403 |
Account inactive, or the key lacks the required scope. |
500 |
Unexpected server error. |
503 |
Service unavailable or temporarily overloaded. |
Complete list in Errors and debugging.