Skip to content

Transcribe a file

POST https://stt.navana.ai/api/transcribe, to upload an audio file and receive a transcript in one request.

POSThttps://stt.navana.ai/api/transcribe

Transcribes an audio file in a single request. Best when you already have the whole recording, such as a stored call or a voicemail.

Authentication

X-Api-Key: <key>. Required. See Authentication.

Request

Content-Type: multipart/form-data

Field Type Required Description
model string Yes Transcription model, which determines the language. There is no auto-detection. See Models and languages.
transaction_id string Yes An id you generate, to correlate this request with your own logs. Must be a valid UUID.
audio_file file Yes The audio to transcribe, mono. WAV is read directly; MP3, M4A, FLAC, OGG, and WebM are decoded server-side.
aux boolean No true to receive aux_info, which adds per-segment timings, tokens, and confidence scores.
parse_number boolean No Turn on inverse text normalisation, converting spoken form to written form, so “पच्चीस लाख” becomes 2500000. Beta, and only affects Hindi, Malayalam, Kannada, Gujarati, and Marathi models, silently ignored on the rest. See Inverse text normalisation.
hotwords array No Context biasing, as [{"phrase": "बोधी", "score": 2.5}]. Boosts recognition of domain-specific or uncommon phrases. score is optional and defaults to 1.5; around 2.5 is recommended for longer phrases.

Example request

Response

200 OK · application/json

{
  "text": "बिल्कुल आपकी पूरी सहायता की जाएगी",
  "call_id": "f99e79ca-6a2b-42f1-8d6b-fd0855e791de",
  "status": "success"
}
Field Type Description
text string The transcript of the full audio.
call_id string Identifier for this transcription.
status string success or error.
error string Summary of the error. Present on failures only.
message string Detailed error message. Present on failures only.

With aux: true

Setting aux adds an aux_info object carrying diagnostics and the segment breakdown:

Field Type Description
aux_info.request_time number Server processing time, in seconds.
aux_info.received_request_time string UTC timestamp of when the server received the request.
aux_info.confidence number Confidence for the whole audio, 0 to 1.
aux_info.segments_meta array Per-segment breakdown, with tokens, timestamps, and confidence scores.

Confidence is the practical reason to turn aux on. A common pattern is to accept high-confidence output automatically and route the rest to a human:

REVIEW_THRESHOLD = 0.75

aux = result.get("aux_info") or {}
if (aux.get("confidence") or 1) < REVIEW_THRESHOLD:
    queue_for_human_review(result["call_id"], result["text"])

Errors

Status Cause
400 A malformed request, an invalid transaction_id, or a model that does not exist.
401 Missing or incorrect API key.
402 The account’s credit balance is exhausted.
403 The account is inactive, or the key lacks the required scope.
500 Unexpected server error.
503 The service is unavailable or temporarily overloaded.

Notes

  • Match the model to the language. There is no auto-detection, so a Hindi model on Tamil audio returns a poor transcript rather than an error.
  • transaction_id must parse as a UUID. A non-UUID value is rejected as a 400.
  • Long files take time. Nothing returns until the whole file is transcribed, so set a generous client timeout.
Navigation

Type to search…

↑↓ navigate↵ selectEsc close