Audio on this platform is raw 16-bit signed PCM, mono in most directions. The exception is uploading a file for transcription, where several compressed formats are accepted and decoded for you. Everything else, meaning every synthesis response and every streaming frame in both directions, is headerless PCM samples.
| Direction | Format |
|---|---|
| Transcription, file upload | WAV, MP3, FLAC, OGG, or WebM, mono |
| Transcription, streaming send | Raw PCM16, mono, as binary frames |
| Synthesis, one-shot response | Raw PCM, no container |
| Synthesis, streaming response | Raw PCM, as binary frames |
WAV is read directly. The other formats are decoded server-side before
transcription, so they cost a little more time up front. AAC audio — .m4a or
.mp4 — is not accepted; convert it first. Converting to 8 kHz
mono WAV yourself avoids that and matches what the models are trained on:
ffmpeg -i input.mp3 -ac 1 -ar 8000 -c:a pcm_s16le output.wavSample rates
Transcription models are trained at 8 kHz, the rate telephony actually delivers. You declare your sample rate in the stream’s config frame and send audio at that rate.
Synthesis takes the rate as part of output_format, written as
"<sample_rate>:<encoding>", for example "8000:pcm16".
| Rate | Typical use |
|---|---|
8000 |
Telephony, matching what a phone call carries |
16000 |
Voice applications and speech pipelines |
24000 |
Highest quality this engine produces |
Encodings
The second half of output_format is the encoding. It decides how many bytes
each sample takes, which is what the playback code below has to match.
| Encoding | Bytes per sample | Use it when |
|---|---|---|
pcm16 |
2 | Almost always — 16-bit PCM is what players, telephony stacks and the helpers below expect |
float32 |
4 | You are feeding a float pipeline that would otherwise convert back |
Playing back what synthesis returns
Synthesis returns raw PCM with no header. Handing those bytes straight to
an audio player or an <audio> element will not work, because nothing in them
says what sample rate or bit depth they are.
One-shot synthesis reports the rate in response headers:
| Header | Meaning |
|---|---|
X-Sample-Rate |
Sample rate of the returned audio |
X-Encoding |
Encoding of the returned audio |
X-Audio-Duration-Ms |
Duration in milliseconds |
Streaming synthesis reports it in the ready frame, as sample_rate.
To get a playable file, prepend a 44-byte RIFF/WAVE header:
import wave
def pcm_to_wav(pcm: bytes, path: str, sample_rate: int) -> None:
with wave.open(path, "wb") as f:
f.setnchannels(1)
f.setsampwidth(2) # 16-bit
f.setframerate(sample_rate)
f.writeframes(pcm)export function pcmToWav(pcm: Uint8Array, sampleRate: number): Blob {
const blockAlign = 2 // mono, 16-bit
const header = new DataView(new ArrayBuffer(44))
const write = (offset: number, s: string) => {
for (let i = 0; i < s.length; i++) header.setUint8(offset + i, s.charCodeAt(i))
}
write(0, 'RIFF')
header.setUint32(4, 36 + pcm.byteLength, true)
write(8, 'WAVE')
write(12, 'fmt ')
header.setUint32(16, 16, true) // fmt chunk size
header.setUint16(20, 1, true) // PCM
header.setUint16(22, 1, true) // channels
header.setUint32(24, sampleRate, true)
header.setUint32(28, sampleRate * blockAlign, true)
header.setUint16(32, blockAlign, true)
header.setUint16(34, 16, true) // bits per sample
write(36, 'data')
header.setUint32(40, pcm.byteLength, true)
return new Blob([header.buffer, pcm], { type: 'audio/wav' })
}package audio
import (
"bytes"
"encoding/binary"
)
// pcmToWAV prepends the 44-byte mono 16-bit RIFF/WAVE header that makes raw
// PCM playable. The samples themselves are copied through untouched.
func pcmToWAV(pcm []byte, sampleRate int) []byte {
const blockAlign = 2 // mono, 16-bit
var buf bytes.Buffer
write := func(v any) { binary.Write(&buf, binary.LittleEndian, v) }
buf.WriteString("RIFF")
write(uint32(36 + len(pcm)))
buf.WriteString("WAVE")
buf.WriteString("fmt ")
write(uint32(16)) // fmt chunk size
write(uint16(1)) // PCM
write(uint16(1)) // channels
write(uint32(sampleRate)) // sample rate
write(uint32(sampleRate * blockAlign)) // byte rate
write(uint16(blockAlign)) // block align
write(uint16(16)) // bits per sample
buf.WriteString("data")
write(uint32(len(pcm)))
buf.Write(pcm)
return buf.Bytes()
}Or with ffmpeg, if you have written the raw bytes to disk:
ffmpeg -f s16le -ar 24000 -ac 1 -i output.pcm output.wavSizing audio
At 16-bit mono, one second of audio is sample_rate × 2 bytes.
| Rate | Per second | Per minute |
|---|---|---|
| 8 kHz | 16 KB | about 1 MB |
| 16 kHz | 32 KB | about 2 MB |
| 24 kHz | 48 KB | about 3 MB |