---
title: "Transcribe a file"
description: "POST https://stt.navana.ai/api/transcribe, to upload an audio file and receive a transcript in one request."
---

> Documentation Index
> Fetch the complete documentation index at: https://docs.navana.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Transcribe a file

Transcribes an audio file in a single request. Best when you already have the
whole recording, such as a stored call or a voicemail.

## Authentication

`X-Api-Key: <key>`. Required. See [Authentication](/authentication).

## Request

`Content-Type: multipart/form-data`

| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `model` | string | Yes | Transcription model, which determines the language. There is no auto-detection. See [Models and languages](/speech-to-text/models). |
| `transaction_id` | string | Yes | An id you generate, to correlate this request with your own logs. **Must be a valid UUID.** |
| `audio_file` | file | Yes | The audio to transcribe, mono. WAV is read directly; MP3, M4A, FLAC, OGG, and WebM are decoded server-side. |
| `aux` | boolean | No | `true` to receive `aux_info`, which adds per-segment timings, tokens, and confidence scores. |
| `parse_number` | boolean | No | Turn on inverse text normalisation, converting spoken form to written form, so "पच्चीस लाख" becomes `2500000`. **Beta**, and **only affects Hindi, Malayalam, Kannada, Gujarati, and Marathi models**, silently ignored on the rest. See [Inverse text normalisation](/speech-to-text/advanced-features#inverse-text-normalisation). |
| `hotwords` | array | No | Context biasing, as `[{"phrase": "बोधी", "score": 2.5}]`. Boosts recognition of domain-specific or uncommon phrases. `score` is optional and defaults to `1.5`; around `2.5` is recommended for longer phrases. |

### Example request

```bash
curl -X POST https://stt.navana.ai/api/transcribe \
  -H "X-Api-Key: $BODHI_API_KEY" \
  -F "model=hi-banking-v2-8khz" \
  -F "transaction_id=$(uuidgen)" \
  -F "aux=true" \
  -F "audio_file=@recording.wav"
```

```python
import os, uuid, requests

response = requests.post(
"https://stt.navana.ai/api/transcribe",
headers={"X-Api-Key": os.environ["BODHI_API_KEY"]},
data={
    "model": "hi-banking-v2-8khz",
    "transaction_id": str(uuid.uuid4()),
    "aux": "true",
},
files={"audio_file": ("recording.wav", open("recording.wav", "rb"), "audio/wav")},
timeout=120,
)
response.raise_for_status()
print(response.json()["text"])
```

```ts
import { openAsBlob } from 'node:fs'

const form = new FormData()
form.set('model', 'hi-banking-v2-8khz')
form.set('transaction_id', crypto.randomUUID())
form.set('aux', 'true')
form.set('audio_file', await openAsBlob('recording.wav'), 'recording.wav')

const res = await fetch('https://stt.navana.ai/api/transcribe', {
  method: 'POST',
  headers: { 'X-Api-Key': process.env.BODHI_API_KEY! },
  body: form,
})
if (!res.ok) throw new Error(await res.text())

const { text } = await res.json()
console.log(text)
```
```go
package main

import (
	"bytes"
	"encoding/json"
	"fmt"
	"io"
	"mime/multipart"
	"net/http"
	"os"

	"github.com/google/uuid"
)

func main() {
	file, err := os.Open("recording.wav")
	if err != nil {
		panic(err)
	}
	defer file.Close()

	var body bytes.Buffer
	form := multipart.NewWriter(&body)
	form.WriteField("model", "hi-banking-v2-8khz")
	form.WriteField("transaction_id", uuid.NewString())
	form.WriteField("aux", "true")
	part, err := form.CreateFormFile("audio_file", "recording.wav")
	if err != nil {
		panic(err)
	}
	if _, err := io.Copy(part, file); err != nil {
		panic(err)
	}
	form.Close()

	req, _ := http.NewRequest(http.MethodPost,
		"https://stt.navana.ai/api/transcribe", &body)
	req.Header.Set("X-Api-Key", os.Getenv("BODHI_API_KEY"))
	req.Header.Set("Content-Type", form.FormDataContentType())

	resp, err := http.DefaultClient.Do(req)
	if err != nil {
		panic(err)
	}
	defer resp.Body.Close()

	var out struct {
		Text   string `json:"text"`
		CallID string `json:"call_id"`
		Status string `json:"status"`
	}
	json.NewDecoder(resp.Body).Decode(&out)
	fmt.Println(out.Text)
}
```

## Response

`200 OK` · `application/json`

```json
{
  "text": "बिल्कुल आपकी पूरी सहायता की जाएगी",
  "call_id": "f99e79ca-6a2b-42f1-8d6b-fd0855e791de",
  "status": "success"
}
```

| Field | Type | Description |
| --- | --- | --- |
| `text` | string | The transcript of the full audio. |
| `call_id` | string | Identifier for this transcription. |
| `status` | string | `success` or `error`. |
| `error` | string | Summary of the error. Present on failures only. |
| `message` | string | Detailed error message. Present on failures only. |

### With `aux: true`

Setting `aux` adds an `aux_info` object carrying diagnostics and the segment
breakdown:

| Field | Type | Description |
| --- | --- | --- |
| `aux_info.request_time` | number | Server processing time, in seconds. |
| `aux_info.received_request_time` | string | UTC timestamp of when the server received the request. |
| `aux_info.confidence` | number | Confidence for the whole audio, `0` to `1`. |
| `aux_info.segments_meta` | array | Per-segment breakdown, with tokens, timestamps, and confidence scores. |

Confidence is the practical reason to turn `aux` on. A common pattern is to
accept high-confidence output automatically and route the rest to a human:

```python
REVIEW_THRESHOLD = 0.75

aux = result.get("aux_info") or {}
if (aux.get("confidence") or 1) < REVIEW_THRESHOLD:
queue_for_human_review(result["call_id"], result["text"])
```

## Errors

| Status | Cause |
| --- | --- |
| `400` | A malformed request, an invalid `transaction_id`, or a model that does not exist. |
| `401` | Missing or incorrect API key. |
| `402` | The account's credit balance is exhausted. |
| `403` | The account is inactive, or the key lacks the required scope. |
| `500` | Unexpected server error. |
| `503` | The service is unavailable or temporarily overloaded. |

## Notes

- **Match the model to the language.** There is no auto-detection, so a Hindi
  model on Tamil audio returns a poor transcript rather than an error.
- **`transaction_id` must parse as a UUID.** A non-UUID value is rejected as a
  `400`.
- **Long files take time.** Nothing returns until the whole file is
  transcribed, so set a generous client timeout.

## Related

- [Non-streaming transcription](/speech-to-text/non-streaming) — Patterns, confidence handling, conversion.
- [Stream transcription](/api-reference/transcribe-streaming) — Live transcription instead.
- [Models and languages](/speech-to-text/models) — Pick the right model.

Source: https://docs.navana.ai/api-reference/transcribe/index.mdx
