Skip to content

Pipecat

Drop Bodhi into the TTS slot of a Pipecat voice agent pipeline.

Pipecat builds voice agents as a pipeline of services. Bodhi provides the text-to-speech stage through the Python SDK, so synthesis behaves like any other Pipecat TTS service — audio streams out as it is generated, and an interruption cancels it mid-utterance.

Install

pip install "bodhi-api-sdk[tts]"

Needs pipecat-ai 1.4 or newer. Verified against 1.4.0, 1.8.1 and 1.11.0.

Add it to a pipeline

import os

from pipecat.pipeline.pipeline import Pipeline
from pipecat.transcriptions.language import Language

from bodhi.integrations.pipecat_tts import BodhiTTSService

tts = BodhiTTSService(
    api_key=os.environ["BODHI_API_KEY"],
    voice="bhavana",
    language=Language.HI,
)

pipeline = Pipeline([
    transport.input(),
    stt,
    context_aggregator.user(),
    llm,
    tts,
    transport.output(),
])

api_key is the only required argument. Everything else has a default, including url, which points at wss://tts.navana.ai/v1.

Arguments

Argument Required Description
api_key Yes Your platform API key, sent as the X-API-Key header.
voice No Voice id. Defaults to default_female. See Voices and languages.
language No Language to synthesize in, as a Pipecat Language. Defaults to Hindi.
sample_rate No Output rate in Hz: 8000, 16000 or 24000. Defaults to the pipeline’s own, falling back to 24000 when the pipeline asks for a rate Bodhi does not serve.
url No Override the endpoint, usually to point at a different deployment.
settings No Runtime-updatable synthesis settings, below.

Settings

Pass these through BodhiTTSService.Settings. They can be updated while the agent is running, and a change reconnects so the new hello takes effect.

Setting Description
voice Voice id. A cloned cv_ voice only works for the language it was cloned under.
language Language to synthesize in.
use_fast Use the distilled model where one is available. Faster, slightly lower quality.
num_step Flow-matching steps, 1 to 100. Higher is slower and cleaner. Omit for the voice’s own default.
guidance_scale Classifier-free guidance. Omit for the voice’s own default.

speed is deliberately absent: the streaming handshake rejects it. Use non-streaming synthesis if you need it.

How it maps onto Pipecat

Audio arrives as TTSAudioRawFrames between a TTSStartedFrame and a TTSStoppedFrame, which is the contract every Pipecat TTS service follows. Because the frames flow through Pipecat’s audio context, an interruption cancels the utterance in progress rather than letting it play to the end, so barge-in works the same as with any other provider.

One connection carries the whole session. Sentences are sent as they arrive from the LLM, and the first audio comes back while the rest is still being synthesized.

Navigation

Type to search…

↑↓ navigate↵ selectEsc close