Skip to content

Advanced features

Context biasing, number parsing, auxiliary metadata, and the controls that shape when a segment ends.

Beyond model and the audio itself, transcription takes a handful of optional controls. Where you put them depends on the mode: as form fields for non-streaming, or inside the config object for streaming.

Feature Field Non-streaming Streaming
Context biasing hotwords Yes Yes
Inverse text normalisation parse_number Yes Yes
Auxiliary metadata aux Yes Yes
Exclude partial results exclude_partial No Yes
Endpoint silence threshold endpoint_silence_duration No Yes

Context biasing with hotwords

hotwords boosts recognition of important or uncommon phrases: product names, domain jargon, or anything the model is unlikely to have seen often. Pass an array of objects, each with a phrase and an optional score.

score defaults to 1.5. Longer phrases generally need a higher value, and around 2.5 is a reasonable starting point for those.

Inverse text normalisation

parse_number turns on inverse text normalisation (ITN), which converts spoken-form output into written form. “दो हज़ार पाँच सौ रुपये” comes back as ₹2,500 rather than as words, and the same applies to phone numbers, addresses, dates, and quantities.

{
  "config": {
    "sample_rate": 8000,
    "transaction_id": "a-uuid-you-generate",
    "model": "hi-banking-v2-8khz",
    "parse_number": true
  }
}

When to use it

Turn it on for human-facing output, such as call logs, QA review, and exported reports, where written form is what a reader expects.

Leave it off when the transcript feeds a downstream LLM. The model handles spoken-form numbers perfectly well, and the extra conversion is one more thing that can be wrong.

The API leaves ITN off unless you ask for it. The playground turns it on by default for the languages that support it.

Language support

Language Code ITN
Gujarati gu Yes
Hindi hi Yes
Kannada kn Yes
Malayalam ml Yes
Marathi mr Yes
Bengali bn No
English en No
Odia or No
Tamil ta No
Telugu te No

On a model whose language is not supported, parse_number is ignored and numbers come back as words.

Limitations

Coverage is partial. Five of the ten languages today, per the table above.

Accuracy is highest on monolingual input. Code-mixed speech, meaning a native language interleaved with English, degrades the output. That includes the Hinglish model, which runs the Hindi normaliser.

Auxiliary metadata

aux adds diagnostic and per-segment detail to the response. On the non-streaming endpoint it produces an aux_info object carrying request_time, received_request_time, an overall confidence, and a segments_meta breakdown. On streaming it populates segment_meta on each frame, with tokens, timestamps, confidence, and per-word confidence on final frames.

Confidence is the practical reason to turn it on, since it lets you accept high-confidence output automatically and route the rest to review. The exact fields are on the non-streaming and streaming reference pages.

Exclude partial results

Streaming sends partial frames as a segment refines, then a complete frame once it settles. If you only act on final text, set exclude_partial to true and the partials are suppressed entirely.

{
  "config": {
    "sample_rate": 8000,
    "transaction_id": "a-uuid-you-generate",
    "model": "hi-banking-v2-8khz",
    "exclude_partial": true
  }
}

Defaults to false. Leave it off if you are showing live captions, since partials are what make captions feel immediate.

Endpoint silence threshold

endpoint_silence_duration controls how long a pause has to last before the recognizer decides a segment has ended and finalizes it.

{
  "config": {
    "sample_rate": 8000,
    "transaction_id": "a-uuid-you-generate",
    "model": "hi-banking-v2-8khz",
    "endpoint_silence_duration": 0.5
  }
}

Accepts 0.44 to 1.2 seconds and defaults to 0.44. A shorter threshold finalizes sooner, which suits a voice agent that needs to respond quickly. A longer one avoids splitting a speaker who pauses mid-thought.

Navigation

Type to search…

↑↓ navigate↵ selectEsc close