POST
https://stt.navana.ai/api/transcribeTranscribes an audio file in a single request. Best when you already have the whole recording, such as a stored call or a voicemail.
Authentication
X-Api-Key: <key>. Required. See Authentication.
Request
Content-Type: multipart/form-data
| Field | Type | Required | Description |
|---|---|---|---|
model |
string | Yes | Transcription model, which determines the language. There is no auto-detection. See Models and languages. |
transaction_id |
string | Yes | An id you generate, to correlate this request with your own logs. Must be a valid UUID. |
audio_file |
file | Yes | The audio to transcribe, mono. WAV is read directly; MP3, M4A, FLAC, OGG, and WebM are decoded server-side. |
aux |
boolean | No | true to receive aux_info, which adds per-segment timings, tokens, and confidence scores. |
parse_number |
boolean | No | Turn on inverse text normalisation, converting spoken form to written form, so “पच्चीस लाख” becomes 2500000. Beta, and only affects Hindi, Malayalam, Kannada, Gujarati, and Marathi models, silently ignored on the rest. See Inverse text normalisation. |
hotwords |
array | No | Context biasing, as [{"phrase": "बोधी", "score": 2.5}]. Boosts recognition of domain-specific or uncommon phrases. score is optional and defaults to 1.5; around 2.5 is recommended for longer phrases. |
Example request
curl -X POST https://stt.navana.ai/api/transcribe \
-H "X-Api-Key: $BODHI_API_KEY" \
-F "model=hi-banking-v2-8khz" \
-F "transaction_id=$(uuidgen)" \
-F "aux=true" \
-F "audio_file=@recording.wav"import os, uuid, requests
response = requests.post(
"https://stt.navana.ai/api/transcribe",
headers={"X-Api-Key": os.environ["BODHI_API_KEY"]},
data={
"model": "hi-banking-v2-8khz",
"transaction_id": str(uuid.uuid4()),
"aux": "true",
},
files={"audio_file": ("recording.wav", open("recording.wav", "rb"), "audio/wav")},
timeout=120,
)
response.raise_for_status()
print(response.json()["text"])import { openAsBlob } from 'node:fs'
const form = new FormData()
form.set('model', 'hi-banking-v2-8khz')
form.set('transaction_id', crypto.randomUUID())
form.set('aux', 'true')
form.set('audio_file', await openAsBlob('recording.wav'), 'recording.wav')
const res = await fetch('https://stt.navana.ai/api/transcribe', {
method: 'POST',
headers: { 'X-Api-Key': process.env.BODHI_API_KEY! },
body: form,
})
if (!res.ok) throw new Error(await res.text())
const { text } = await res.json()
console.log(text)package main
import (
"bytes"
"encoding/json"
"fmt"
"io"
"mime/multipart"
"net/http"
"os"
"github.com/google/uuid"
)
func main() {
file, err := os.Open("recording.wav")
if err != nil {
panic(err)
}
defer file.Close()
var body bytes.Buffer
form := multipart.NewWriter(&body)
form.WriteField("model", "hi-banking-v2-8khz")
form.WriteField("transaction_id", uuid.NewString())
form.WriteField("aux", "true")
part, err := form.CreateFormFile("audio_file", "recording.wav")
if err != nil {
panic(err)
}
if _, err := io.Copy(part, file); err != nil {
panic(err)
}
form.Close()
req, _ := http.NewRequest(http.MethodPost,
"https://stt.navana.ai/api/transcribe", &body)
req.Header.Set("X-Api-Key", os.Getenv("BODHI_API_KEY"))
req.Header.Set("Content-Type", form.FormDataContentType())
resp, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer resp.Body.Close()
var out struct {
Text string `json:"text"`
CallID string `json:"call_id"`
Status string `json:"status"`
}
json.NewDecoder(resp.Body).Decode(&out)
fmt.Println(out.Text)
}Response
200 OK · application/json
{
"text": "बिल्कुल आपकी पूरी सहायता की जाएगी",
"call_id": "f99e79ca-6a2b-42f1-8d6b-fd0855e791de",
"status": "success"
}| Field | Type | Description |
|---|---|---|
text |
string | The transcript of the full audio. |
call_id |
string | Identifier for this transcription. |
status |
string | success or error. |
error |
string | Summary of the error. Present on failures only. |
message |
string | Detailed error message. Present on failures only. |
With aux: true
Setting aux adds an aux_info object carrying diagnostics and the segment
breakdown:
| Field | Type | Description |
|---|---|---|
aux_info.request_time |
number | Server processing time, in seconds. |
aux_info.received_request_time |
string | UTC timestamp of when the server received the request. |
aux_info.confidence |
number | Confidence for the whole audio, 0 to 1. |
aux_info.segments_meta |
array | Per-segment breakdown, with tokens, timestamps, and confidence scores. |
Confidence is the practical reason to turn aux on. A common pattern is to
accept high-confidence output automatically and route the rest to a human:
REVIEW_THRESHOLD = 0.75
aux = result.get("aux_info") or {}
if (aux.get("confidence") or 1) < REVIEW_THRESHOLD:
queue_for_human_review(result["call_id"], result["text"])Errors
| Status | Cause |
|---|---|
400 |
A malformed request, an invalid transaction_id, or a model that does not exist. |
401 |
Missing or incorrect API key. |
402 |
The account’s credit balance is exhausted. |
403 |
The account is inactive, or the key lacks the required scope. |
500 |
Unexpected server error. |
503 |
The service is unavailable or temporarily overloaded. |
Notes
- Match the model to the language. There is no auto-detection, so a Hindi model on Tamil audio returns a poor transcript rather than an error.
transaction_idmust parse as a UUID. A non-UUID value is rejected as a400.- Long files take time. Nothing returns until the whole file is transcribed, so set a generous client timeout.