Skip to content

Models and languages

The recommended transcription model for each of 10 Indian languages, plus a Hinglish compatibility code.

Pass a model name to model in the form field for non-streaming, in the config frame for streaming. There is no auto-detection: the model you name determines the language.

Available models

Language Code Model Notes
Bengali bn bn-banking-v2-8khz Bilingual with English (code-switching)
English en en-banking-v2-8khz English-only (dedicated model)
Gujarati gu gu-banking-v2-8khz Monolingual
Hindi hi hi-banking-v2-8khz Bilingual with English (code-switching)
Hinglish hi-en hi-en-banking-v2-8khz Bilingual with English, retained for compatibility
Kannada kn kn-banking-v2-8khz Bilingual with English (code-switching)
Malayalam ml ml-banking-v2-8khz Bilingual with English (code-switching)
Marathi mr mr-banking-v2-8khz Bilingual with English (code-switching)
Odia or or-general-v3-8khz Monolingual
Tamil ta ta-banking-v2-8khz Bilingual with English (code-switching)
Telugu te te-banking-v2-8khz Bilingual with English (code-switching)

Odia has no banking model yet, so or-general-v3-8khz is its only option today.

Reading a model name

hi-banking-v2-8khz
│   │       │  └── training sample rate
│   │       └───── model version
│   └───────────── domain
└───────────────── language

Odia is the one exception to the version number: its model is or-general-v3-8khz, not v2.

Domain. These models are tuned for financial vocabulary: account numbers, balances, transaction terms, the code-switched English that shows up in Indian banking calls.

Sample rate. Every model is trained at 8 kHz, the rate telephony actually delivers. Feeding higher-rate audio doesn’t improve accuracy on a model trained at 8 kHz, and downsampling a 16 kHz recording to 8 kHz before transcribing usually matches the training distribution better.

Choosing a model

Start from the language

There’s no detection and no fallback. A Hindi model on Tamil audio returns a poor transcript rather than an error, so pick the model from your own metadata about the call.

Validate on your own audio

Run a sample set through the model and check per-segment confidence from aux_info. That’s a cheap, objective way to spot-check without labelled ground truth.

Navigation

Type to search…

↑↓ navigate↵ selectEsc close