Pipecat builds voice agents as a pipeline of services. Bodhi provides the text-to-speech stage through the Python SDK, so synthesis behaves like any other Pipecat TTS service — audio streams out as it is generated, and an interruption cancels it mid-utterance.
Install
pip install "bodhi-api-sdk[tts]"Needs pipecat-ai 1.4 or newer. Verified against 1.4.0, 1.8.1 and 1.11.0.
Add it to a pipeline
import os
from pipecat.pipeline.pipeline import Pipeline
from pipecat.transcriptions.language import Language
from bodhi.integrations.pipecat_tts import BodhiTTSService
tts = BodhiTTSService(
api_key=os.environ["BODHI_API_KEY"],
voice="bhavana",
language=Language.HI,
)
pipeline = Pipeline([
transport.input(),
stt,
context_aggregator.user(),
llm,
tts,
transport.output(),
])api_key is the only required argument. Everything else has a default,
including url, which points at wss://tts.navana.ai/v1.
Arguments
| Argument | Required | Description |
|---|---|---|
api_key |
Yes | Your platform API key, sent as the X-API-Key header. |
voice |
No | Voice id. Defaults to default_female. See Voices and languages. |
language |
No | Language to synthesize in, as a Pipecat Language. Defaults to Hindi. |
sample_rate |
No | Output rate in Hz: 8000, 16000 or 24000. Defaults to the pipeline’s own, falling back to 24000 when the pipeline asks for a rate Bodhi does not serve. |
url |
No | Override the endpoint, usually to point at a different deployment. |
settings |
No | Runtime-updatable synthesis settings, below. |
Settings
Pass these through BodhiTTSService.Settings. They can be updated while the
agent is running, and a change reconnects so the new hello takes effect.
| Setting | Description |
|---|---|
voice |
Voice id. A cloned cv_ voice only works for the language it was cloned under. |
language |
Language to synthesize in. |
use_fast |
Use the distilled model where one is available. Faster, slightly lower quality. |
num_step |
Flow-matching steps, 1 to 100. Higher is slower and cleaner. Omit for the voice’s own default. |
guidance_scale |
Classifier-free guidance. Omit for the voice’s own default. |
speed is deliberately absent: the streaming handshake rejects it. Use
non-streaming synthesis if you need it.
How it maps onto Pipecat
Audio arrives as TTSAudioRawFrames between a TTSStartedFrame and a
TTSStoppedFrame, which is the contract every Pipecat TTS service follows.
Because the frames flow through Pipecat’s audio context, an interruption
cancels the utterance in progress rather than letting it play to the end, so
barge-in works the same as with any other provider.
One connection carries the whole session. Sentences are sent as they arrive from the LLM, and the first audio comes back while the rest is still being synthesized.