Skip to content

Audio formats

What the APIs accept, what they return, and how to convert between them.

Audio on this platform is raw 16-bit signed PCM, mono in most directions. The exception is uploading a file for transcription, where several compressed formats are accepted and decoded for you. Everything else, meaning every synthesis response and every streaming frame in both directions, is headerless PCM samples.

Direction Format
Transcription, file upload WAV, MP3, FLAC, OGG, or WebM, mono
Transcription, streaming send Raw PCM16, mono, as binary frames
Synthesis, one-shot response Raw PCM, no container
Synthesis, streaming response Raw PCM, as binary frames

WAV is read directly. The other formats are decoded server-side before transcription, so they cost a little more time up front. AAC audio — .m4a or .mp4 — is not accepted; convert it first. Converting to 8 kHz mono WAV yourself avoids that and matches what the models are trained on:

Terminalbash
ffmpeg -i input.mp3 -ac 1 -ar 8000 -c:a pcm_s16le output.wav

Sample rates

Transcription models are trained at 8 kHz, the rate telephony actually delivers. You declare your sample rate in the stream’s config frame and send audio at that rate.

Synthesis takes the rate as part of output_format, written as "<sample_rate>:<encoding>", for example "8000:pcm16".

Rate Typical use
8000 Telephony, matching what a phone call carries
16000 Voice applications and speech pipelines
24000 Highest quality this engine produces

Encodings

The second half of output_format is the encoding. It decides how many bytes each sample takes, which is what the playback code below has to match.

Encoding Bytes per sample Use it when
pcm16 2 Almost always — 16-bit PCM is what players, telephony stacks and the helpers below expect
float32 4 You are feeding a float pipeline that would otherwise convert back

Playing back what synthesis returns

Synthesis returns raw PCM with no header. Handing those bytes straight to an audio player or an <audio> element will not work, because nothing in them says what sample rate or bit depth they are.

One-shot synthesis reports the rate in response headers:

Header Meaning
X-Sample-Rate Sample rate of the returned audio
X-Encoding Encoding of the returned audio
X-Audio-Duration-Ms Duration in milliseconds

Streaming synthesis reports it in the ready frame, as sample_rate.

To get a playable file, prepend a 44-byte RIFF/WAVE header:

Or with ffmpeg, if you have written the raw bytes to disk:

Terminalbash
ffmpeg -f s16le -ar 24000 -ac 1 -i output.pcm output.wav

Sizing audio

At 16-bit mono, one second of audio is sample_rate × 2 bytes.

Rate Per second Per minute
8 kHz 16 KB about 1 MB
16 kHz 32 KB about 2 MB
24 kHz 48 KB about 3 MB
Navigation

Type to search…

↑↓ navigate↵ selectEsc close