Skip to content

Stream transcription

WebSocket transcription at wss://stt.navana.ai, with partial and final results.

WSSwss://stt.navana.ai

A WebSocket that accepts PCM audio frames and returns transcript frames in real time.

Authentication

X-Api-Key: <key> on the upgrade request. See Authentication.

Protocol

1. Config frame, client to server

The first message, as JSON text. Nothing is transcribed until it arrives.

{
  "config": {
    "sample_rate": 8000,
    "transaction_id": "a-uuid-you-generate",
    "model": "hi-banking-v2-8khz"
  }
}
Field Type Required Description
sample_rate number Yes Sample rate of the PCM frames you will send. Must match your audio, since nothing verifies it.
transaction_id string Yes An id you generate, to correlate this session with your own logs. Must be a valid UUID.
model string Yes Transcription model. See Models and languages.
aux boolean No true to receive segment_meta, which carries timings, confidence, tokens, and per-word confidence on final frames.
parse_number boolean No Turn on inverse text normalisation, converting spoken form to written form, so “पच्चीस लाख” becomes 2500000. Beta, and only affects Hindi, Malayalam, Kannada, Gujarati, and Marathi models, silently ignored on the rest. See Inverse text normalisation.
hotwords array No Context biasing, as [{"phrase": "बोधी", "score": 2.5}]. Boosts recognition of domain-specific or uncommon phrases. score is optional and defaults to 1.5; around 2.5 is recommended for longer phrases.
exclude_partial boolean No true to receive only "complete" frames, suppressing partials entirely. Defaults to false.
endpoint_silence_duration number No Seconds of silence before an in-progress segment is finalized. Accepts 0.44 to 1.2, defaults to 0.44.

2. Audio frames, client to server

Binary WebSocket frames carrying 16-bit signed PCM, mono, at the sample rate you declared in the config frame.

There is no required frame size. Smaller frames lower latency, larger ones reduce overhead, and around 100 ms is a reasonable default. At 8 kHz mono that is 1600 bytes per frame, since one second of audio is sample_rate × 2 bytes.

3. Transcript frames, server to client

{
  "call_id": "0f8c1b2e-4a55-4c7e-9d31-1b6a0e2f7c44",
  "segment_id": 0,
  "eos": false,
  "type": "partial",
  "text": "the words recognised so far",
  "segment_meta": {
    "start_time": 0,
    "confidence": 0.91,
    "tokens": ["the", "words", "recognised", "so", "far"],
    "timestamps": [0.2, 0.5, 0.8, 1.1, 1.3],
    "words": [
      { "word": "the", "confidence": 0.94 },
      { "word": "words", "confidence": 0.91 }
    ]
  }
}
Field Type Description
call_id string Identifier for this streaming connection.
segment_id number Which speech segment this frame is about. Successive frames with the same id replace each other, so do not append.
eos boolean true on the frame that ends the stream. This is how you know nothing more is coming.
type string "partial" while the segment is still refining, "complete" once it is final.
text string The transcript processed so far for this segment.
segment_meta object Present when you set aux: true.
segment_meta.start_time number Where this segment starts, in seconds from the start of the audio.
segment_meta.confidence number Segment-level confidence, 0 to 1.
segment_meta.tokens string[] The individual tokens recognized from the audio, in order.
segment_meta.timestamps number[] When each token was detected, in seconds. Pairs positionally with tokens.
segment_meta.words array Only populated when type is "complete". Per-word breakdown as [{"word": "…", "confidence": 0.91}].

4. EOF frame, client to server

{ "eof": 1 }

Signals that you have finished sending audio. Remaining transcripts flush, and the final frame carries eos: true.

Example

Errors

The upgrade is rejected with a normal HTTP response, so you get a readable status rather than a silent close.

Status Cause
400 Malformed config frame, an invalid transaction_id, or a model that does not exist.
401 Missing or incorrect API key.
402 The account’s credit balance is exhausted.
403 The account is inactive, or the key lacks the required scope.
500 Unexpected server error.
503 The service is unavailable or temporarily overloaded.

Notes

  • eos is the end-of-stream signal. Wait for it rather than a fixed delay after your eof frame, since how much audio is still being transcribed when you stop sending varies. Settle your client on whichever comes first, the eos frame or the socket closing, so a session can never leave it waiting.
  • Sample rate is not verified. Declaring one rate and sending another produces a poor transcript rather than an error.
  • parse_number only affects five languages, namely Hindi, Malayalam, Kannada, Gujarati, and Marathi.
Navigation

Type to search…

↑↓ navigate↵ selectEsc close