---
title: "Audio formats"
description: "What the APIs accept, what they return, and how to convert between them."
---

> Documentation Index
> Fetch the complete documentation index at: https://docs.navana.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Audio formats

Audio on this platform is **raw 16-bit signed PCM, mono** in most directions.
The exception is uploading a file for transcription, where several compressed
formats are accepted and decoded for you. Everything else, meaning every
synthesis response and every streaming frame in both directions, is headerless
PCM samples.

| Direction | Format |
| --- | --- |
| Transcription, file upload | WAV, MP3, FLAC, OGG, or WebM, mono |
| Transcription, streaming send | Raw PCM16, mono, as binary frames |
| Synthesis, one-shot response | Raw PCM, no container |
| Synthesis, streaming response | Raw PCM, as binary frames |

WAV is read directly. The other formats are decoded server-side before
transcription, so they cost a little more time up front. AAC audio — `.m4a` or
`.mp4` — is not accepted; convert it first. Converting to 8 kHz
mono WAV yourself avoids that and matches what the models are trained on:

```bash title="Terminal"
ffmpeg -i input.mp3 -ac 1 -ar 8000 -c:a pcm_s16le output.wav
```

## Sample rates

**Transcription** models are trained at 8 kHz, the rate telephony actually
delivers. You declare your sample rate in the stream's config frame and send
audio at that rate.

**Synthesis** takes the rate as part of `output_format`, written as
`"<sample_rate>:<encoding>"`, for example `"8000:pcm16"`.

| Rate | Typical use |
| --- | --- |
| `8000` | Telephony, matching what a phone call carries |
| `16000` | Voice applications and speech pipelines |
| `24000` | Highest quality this engine produces |

## Encodings

The second half of `output_format` is the encoding. It decides how many bytes
each sample takes, which is what the playback code below has to match.

| Encoding | Bytes per sample | Use it when |
| --- | --- | --- |
| `pcm16` | 2 | Almost always — 16-bit PCM is what players, telephony stacks and the helpers below expect |
| `float32` | 4 | You are feeding a float pipeline that would otherwise convert back |

> **The default is not pcm16**
>
> Omitting `output_format` gives you `"24000:float32"`, not 16-bit PCM. Four-byte
> float samples read as 16-bit produce noise, so a request that leaves the field
> out and then uses the helpers below will not play back correctly. Name the
> format explicitly — `"8000:pcm16"` for telephony, `"24000:pcm16"` otherwise.

## Playing back what synthesis returns

Synthesis returns **raw PCM with no header**. Handing those bytes straight to
an audio player or an `<audio>` element will not work, because nothing in them
says what sample rate or bit depth they are.

One-shot synthesis reports the rate in response headers:

| Header | Meaning |
| --- | --- |
| `X-Sample-Rate` | Sample rate of the returned audio |
| `X-Encoding` | Encoding of the returned audio |
| `X-Audio-Duration-Ms` | Duration in milliseconds |

Streaming synthesis reports it in the `ready` frame, as `sample_rate`.

To get a playable file, prepend a 44-byte RIFF/WAVE header:

```python
import wave

def pcm_to_wav(pcm: bytes, path: str, sample_rate: int) -> None:
with wave.open(path, "wb") as f:
    f.setnchannels(1)
    f.setsampwidth(2)   # 16-bit
    f.setframerate(sample_rate)
    f.writeframes(pcm)
```

```ts
export function pcmToWav(pcm: Uint8Array, sampleRate: number): Blob {
  const blockAlign = 2 // mono, 16-bit
  const header = new DataView(new ArrayBuffer(44))
  const write = (offset: number, s: string) => {
for (let i = 0; i < s.length; i++) header.setUint8(offset + i, s.charCodeAt(i))
  }

  write(0, 'RIFF')
  header.setUint32(4, 36 + pcm.byteLength, true)
  write(8, 'WAVE')
  write(12, 'fmt ')
  header.setUint32(16, 16, true)          // fmt chunk size
  header.setUint16(20, 1, true)           // PCM
  header.setUint16(22, 1, true)           // channels
  header.setUint32(24, sampleRate, true)
  header.setUint32(28, sampleRate * blockAlign, true)
  header.setUint16(32, blockAlign, true)
  header.setUint16(34, 16, true)          // bits per sample
  write(36, 'data')
  header.setUint32(40, pcm.byteLength, true)

  return new Blob([header.buffer, pcm], { type: 'audio/wav' })
}
```
```go
package audio

import (
	"bytes"
	"encoding/binary"
)

// pcmToWAV prepends the 44-byte mono 16-bit RIFF/WAVE header that makes raw
// PCM playable. The samples themselves are copied through untouched.
func pcmToWAV(pcm []byte, sampleRate int) []byte {
	const blockAlign = 2 // mono, 16-bit

	var buf bytes.Buffer
	write := func(v any) { binary.Write(&buf, binary.LittleEndian, v) }

	buf.WriteString("RIFF")
	write(uint32(36 + len(pcm)))
	buf.WriteString("WAVE")
	buf.WriteString("fmt ")
	write(uint32(16))                      // fmt chunk size
	write(uint16(1))                       // PCM
	write(uint16(1))                       // channels
	write(uint32(sampleRate))              // sample rate
	write(uint32(sampleRate * blockAlign)) // byte rate
	write(uint16(blockAlign))              // block align
	write(uint16(16))                      // bits per sample
	buf.WriteString("data")
	write(uint32(len(pcm)))
	buf.Write(pcm)

	return buf.Bytes()
}
```

Or with `ffmpeg`, if you have written the raw bytes to disk:

```bash title="Terminal"
ffmpeg -f s16le -ar 24000 -ac 1 -i output.pcm output.wav
```

> **Concatenating streamed chunks is safe**
>
> Because streamed audio is headerless PCM, chunks concatenate byte for byte
> with no seams. Append them as they arrive and wrap the whole buffer once at
> the end, or feed them straight into a live audio sink and skip the header
> entirely.

## Sizing audio

At 16-bit mono, one second of audio is `sample_rate × 2` bytes.

| Rate | Per second | Per minute |
| --- | --- | --- |
| 8 kHz | 16 KB | about 1 MB |
| 16 kHz | 32 KB | about 2 MB |
| 24 kHz | 48 KB | about 3 MB |

Source: https://docs.navana.ai/concepts/audio-formats/index.mdx
