---
title: "Advanced features"
description: "Context biasing, number parsing, auxiliary metadata, and the controls that shape when a segment ends."
---

> Documentation Index
> Fetch the complete documentation index at: https://docs.navana.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Advanced features

Beyond `model` and the audio itself, transcription takes a handful of optional
controls. Where you put them depends on the mode: as form fields for
[non-streaming](/speech-to-text/non-streaming), or inside the `config` object
for [streaming](/speech-to-text/streaming).

| Feature | Field | Non-streaming | Streaming |
| --- | --- | --- | --- |
| Context biasing | `hotwords` | Yes | Yes |
| Inverse text normalisation | `parse_number` | Yes | Yes |
| Auxiliary metadata | `aux` | Yes | Yes |
| Exclude partial results | `exclude_partial` | No | Yes |
| Endpoint silence threshold | `endpoint_silence_duration` | No | Yes |

## Context biasing with hotwords

`hotwords` boosts recognition of important or uncommon phrases: product names,
domain jargon, or anything the model is unlikely to have seen often. Pass an
array of objects, each with a `phrase` and an optional `score`.

```python
hindi_hotwords = [
{"phrase": "बोधी"},
{"phrase": "स्पीच रिकग्निशन", "score": 2.5},
]

await socket.send(json.dumps({
"config": {
    "sample_rate": 8000,
    "transaction_id": str(uuid.uuid4()),
    "model": "hi-banking-v2-8khz",
    "hotwords": hindi_hotwords,
}
}))
```

```python
hindi_hotwords = [
{"phrase": "बोधी"},
{"phrase": "स्पीच रिकग्निशन", "score": 2.5},
]

requests.post(
"https://stt.navana.ai/api/transcribe",
headers={"X-Api-Key": os.environ["BODHI_API_KEY"]},
data={
    "model": "hi-banking-v2-8khz",
    "transaction_id": str(uuid.uuid4()),
    "hotwords": json.dumps(hindi_hotwords),
},
files={"audio_file": open("recording.wav", "rb")},
)
```

`score` defaults to `1.5`. Longer phrases generally need a higher value, and
around `2.5` is a reasonable starting point for those.

## Inverse text normalisation

`parse_number` turns on **inverse text normalisation (ITN)**, which converts
spoken-form output into written form. "दो हज़ार पाँच सौ रुपये" comes back as
`₹2,500` rather than as words, and the same applies to phone numbers,
addresses, dates, and quantities.

```json
{
  "config": {
"sample_rate": 8000,
"transaction_id": "a-uuid-you-generate",
"model": "hi-banking-v2-8khz",
"parse_number": true
  }
}
```

> **Beta**
>
> ITN is in beta. Its behaviour and accuracy may change between releases.

### When to use it

**Turn it on for human-facing output**, such as call logs, QA review, and
exported reports, where written form is what a reader expects.

**Leave it off when the transcript feeds a downstream LLM.** The model handles
spoken-form numbers perfectly well, and the extra conversion is one more thing
that can be wrong.

The API leaves ITN off unless you ask for it. The playground turns it on by
default for the languages that support it.

### Language support

| Language | Code | ITN |
| --- | --- | --- |
| Gujarati | `gu` | Yes |
| Hindi | `hi` | Yes |
| Kannada | `kn` | Yes |
| Malayalam | `ml` | Yes |
| Marathi | `mr` | Yes |
| Bengali | `bn` | No |
| English | `en` | No |
| Odia | `or` | No |
| Tamil | `ta` | No |
| Telugu | `te` | No |

On a model whose language is not supported, `parse_number` is ignored and
numbers come back as words.

### Limitations

**Coverage is partial.** Five of the ten languages today, per the table above.

**Accuracy is highest on monolingual input.** Code-mixed speech, meaning a
native language interleaved with English, degrades the output. That includes
the Hinglish model, which runs the Hindi normaliser.

## Auxiliary metadata

`aux` adds diagnostic and per-segment detail to the response. On the
non-streaming endpoint it produces an `aux_info` object carrying
`request_time`, `received_request_time`, an overall `confidence`, and a
`segments_meta` breakdown. On streaming it populates `segment_meta` on each
frame, with tokens, timestamps, confidence, and per-word confidence on final
frames.

Confidence is the practical reason to turn it on, since it lets you accept
high-confidence output automatically and route the rest to review. The exact
fields are on the
[non-streaming](/api-reference/transcribe) and
[streaming](/api-reference/transcribe-streaming) reference pages.

## Exclude partial results

Streaming sends `partial` frames as a segment refines, then a `complete` frame
once it settles. If you only act on final text, set `exclude_partial` to `true`
and the partials are suppressed entirely.

```json
{
  "config": {
"sample_rate": 8000,
"transaction_id": "a-uuid-you-generate",
"model": "hi-banking-v2-8khz",
"exclude_partial": true
  }
}
```

Defaults to `false`. Leave it off if you are showing live captions, since
partials are what make captions feel immediate.

## Endpoint silence threshold

`endpoint_silence_duration` controls how long a pause has to last before the
recognizer decides a segment has ended and finalizes it.

```json
{
  "config": {
"sample_rate": 8000,
"transaction_id": "a-uuid-you-generate",
"model": "hi-banking-v2-8khz",
"endpoint_silence_duration": 0.5
  }
}
```

Accepts `0.44` to `1.2` seconds and defaults to `0.44`. A shorter threshold
finalizes sooner, which suits a voice agent that needs to respond quickly. A
longer one avoids splitting a speaker who pauses mid-thought.

## Related

- [Non-streaming transcription](/speech-to-text/non-streaming) — Where these go as form fields.
- [Streaming transcription](/speech-to-text/streaming) — Where these go in the config frame.
- [Models and languages](/speech-to-text/models) — Which model to pair them with.

Source: https://docs.navana.ai/speech-to-text/advanced-features/index.mdx
