Pass a model name to model in the form field for non-streaming, in the config
frame for streaming. There is no auto-detection: the model you name determines
the language.
Available models
| Language | Code | Model | Notes |
|---|---|---|---|
| Bengali | bn |
bn-banking-v2-8khz |
Bilingual with English (code-switching) |
| English | en |
en-banking-v2-8khz |
English-only (dedicated model) |
| Gujarati | gu |
gu-banking-v2-8khz |
Monolingual |
| Hindi | hi |
hi-banking-v2-8khz |
Bilingual with English (code-switching) |
| Hinglish | hi-en |
hi-en-banking-v2-8khz |
Bilingual with English, retained for compatibility |
| Kannada | kn |
kn-banking-v2-8khz |
Bilingual with English (code-switching) |
| Malayalam | ml |
ml-banking-v2-8khz |
Bilingual with English (code-switching) |
| Marathi | mr |
mr-banking-v2-8khz |
Bilingual with English (code-switching) |
| Odia | or |
or-general-v3-8khz |
Monolingual |
| Tamil | ta |
ta-banking-v2-8khz |
Bilingual with English (code-switching) |
| Telugu | te |
te-banking-v2-8khz |
Bilingual with English (code-switching) |
Odia has no banking model yet, so or-general-v3-8khz is its only option
today.
Reading a model name
hi-banking-v2-8khz
│ │ │ └── training sample rate
│ │ └───── model version
│ └───────────── domain
└───────────────── languageOdia is the one exception to the version number: its model is
or-general-v3-8khz, not v2.
Domain. These models are tuned for financial vocabulary: account numbers, balances, transaction terms, the code-switched English that shows up in Indian banking calls.
Sample rate. Every model is trained at 8 kHz, the rate telephony actually delivers. Feeding higher-rate audio doesn’t improve accuracy on a model trained at 8 kHz, and downsampling a 16 kHz recording to 8 kHz before transcribing usually matches the training distribution better.
Choosing a model
Start from the language
There’s no detection and no fallback. A Hindi model on Tamil audio returns a poor transcript rather than an error, so pick the model from your own metadata about the call.
Validate on your own audio
Run a sample set through the model and check per-segment confidence from
aux_info. That’s a cheap, objective way to spot-check without labelled
ground truth.