Beyond model and the audio itself, transcription takes a handful of optional
controls. Where you put them depends on the mode: as form fields for
non-streaming, or inside the config object
for streaming.
| Feature | Field | Non-streaming | Streaming |
|---|---|---|---|
| Context biasing | hotwords |
Yes | Yes |
| Inverse text normalisation | parse_number |
Yes | Yes |
| Auxiliary metadata | aux |
Yes | Yes |
| Exclude partial results | exclude_partial |
No | Yes |
| Endpoint silence threshold | endpoint_silence_duration |
No | Yes |
Context biasing with hotwords
hotwords boosts recognition of important or uncommon phrases: product names,
domain jargon, or anything the model is unlikely to have seen often. Pass an
array of objects, each with a phrase and an optional score.
hindi_hotwords = [
{"phrase": "बोधी"},
{"phrase": "स्पीच रिकग्निशन", "score": 2.5},
]
await socket.send(json.dumps({
"config": {
"sample_rate": 8000,
"transaction_id": str(uuid.uuid4()),
"model": "hi-banking-v2-8khz",
"hotwords": hindi_hotwords,
}
}))hindi_hotwords = [
{"phrase": "बोधी"},
{"phrase": "स्पीच रिकग्निशन", "score": 2.5},
]
requests.post(
"https://stt.navana.ai/api/transcribe",
headers={"X-Api-Key": os.environ["BODHI_API_KEY"]},
data={
"model": "hi-banking-v2-8khz",
"transaction_id": str(uuid.uuid4()),
"hotwords": json.dumps(hindi_hotwords),
},
files={"audio_file": open("recording.wav", "rb")},
)score defaults to 1.5. Longer phrases generally need a higher value, and
around 2.5 is a reasonable starting point for those.
Inverse text normalisation
parse_number turns on inverse text normalisation (ITN), which converts
spoken-form output into written form. “दो हज़ार पाँच सौ रुपये” comes back as
₹2,500 rather than as words, and the same applies to phone numbers,
addresses, dates, and quantities.
{
"config": {
"sample_rate": 8000,
"transaction_id": "a-uuid-you-generate",
"model": "hi-banking-v2-8khz",
"parse_number": true
}
}When to use it
Turn it on for human-facing output, such as call logs, QA review, and exported reports, where written form is what a reader expects.
Leave it off when the transcript feeds a downstream LLM. The model handles spoken-form numbers perfectly well, and the extra conversion is one more thing that can be wrong.
The API leaves ITN off unless you ask for it. The playground turns it on by default for the languages that support it.
Language support
| Language | Code | ITN |
|---|---|---|
| Gujarati | gu |
Yes |
| Hindi | hi |
Yes |
| Kannada | kn |
Yes |
| Malayalam | ml |
Yes |
| Marathi | mr |
Yes |
| Bengali | bn |
No |
| English | en |
No |
| Odia | or |
No |
| Tamil | ta |
No |
| Telugu | te |
No |
On a model whose language is not supported, parse_number is ignored and
numbers come back as words.
Limitations
Coverage is partial. Five of the ten languages today, per the table above.
Accuracy is highest on monolingual input. Code-mixed speech, meaning a native language interleaved with English, degrades the output. That includes the Hinglish model, which runs the Hindi normaliser.
Auxiliary metadata
aux adds diagnostic and per-segment detail to the response. On the
non-streaming endpoint it produces an aux_info object carrying
request_time, received_request_time, an overall confidence, and a
segments_meta breakdown. On streaming it populates segment_meta on each
frame, with tokens, timestamps, confidence, and per-word confidence on final
frames.
Confidence is the practical reason to turn it on, since it lets you accept high-confidence output automatically and route the rest to review. The exact fields are on the non-streaming and streaming reference pages.
Exclude partial results
Streaming sends partial frames as a segment refines, then a complete frame
once it settles. If you only act on final text, set exclude_partial to true
and the partials are suppressed entirely.
{
"config": {
"sample_rate": 8000,
"transaction_id": "a-uuid-you-generate",
"model": "hi-banking-v2-8khz",
"exclude_partial": true
}
}Defaults to false. Leave it off if you are showing live captions, since
partials are what make captions feel immediate.
Endpoint silence threshold
endpoint_silence_duration controls how long a pause has to last before the
recognizer decides a segment has ended and finalizes it.
{
"config": {
"sample_rate": 8000,
"transaction_id": "a-uuid-you-generate",
"model": "hi-banking-v2-8khz",
"endpoint_silence_duration": 0.5
}
}Accepts 0.44 to 1.2 seconds and defaults to 0.44. A shorter threshold
finalizes sooner, which suits a voice agent that needs to respond quickly. A
longer one avoids splitting a speaker who pauses mid-thought.