Skip to content

Pipecat

Drop Bodhi into the STT slot of a Pipecat voice agent pipeline.

Pipecat builds voice agents as a pipeline of services. Bodhi provides the speech-to-text stage through the Python SDK, so transcription behaves like any other Pipecat STT service.

Install

pip install "bodhi-api-sdk[stt]"

Needs pipecat-ai 1.4 or newer. Verified against 1.4.0, 1.8.1 and 1.11.0.

Add it to a pipeline

import os

from pipecat.pipeline.pipeline import Pipeline

from bodhi.integrations.pipecat_stt import BodhiHotword, BodhiSTTService

stt = BodhiSTTService(
    api_key=os.environ["BODHI_API_KEY"],
    model="hi-banking-v2-8khz",
    settings=BodhiSTTService.Settings(
        parse_number=True,
        endpoint_silence_duration=0.6,
        hotwords=[BodhiHotword("बजाज फिनसर्व", 2.0)],
    ),
)

pipeline = Pipeline([
    transport.input(),
    stt,
    context_aggregator.user(),
    llm,
    tts,
    transport.output(),
])

api_key and model are the only required arguments. Everything else has a default, including url, which points at wss://stt.navana.ai.

Arguments

Argument Required Description
api_key Yes Your platform API key, sent as the x-api-key header.
model Yes Any transcription model. See Models and languages.
url No Override the endpoint, usually to point at a different deployment.
sample_rate No Input rate in Hz. Defaults to the pipeline’s own.
language No Language tag for emitted frames. Defaults to the model’s prefix.
interim_results No Emit InterimTranscriptionFrames from partials. Defaults to on.
aux No Attach timing and confidence metadata to each frame.
min_confidence No Drop finals below this utterance confidence. Defaults to 0.5; pass 0 to keep everything.
transaction_id No Correlation id for the session. A fresh UUID per connection when omitted.
settings No Runtime-updatable recognition settings, below.

Settings

Pass these through BodhiSTTService.Settings. They can be updated while the agent is running.

Setting Description
model The transcription model. Changing it reconnects the session.
language Language used to tag emitted frames. Presentational only, since Bodhi infers the language from the model.
hotwords Phrase boosting, as BodhiHotword(phrase, score). score is optional, and roughly 1.0 to 3.0 is useful. See Advanced features.
parse_number Normalise spoken numbers, dates, and currency in the returned text.
endpoint_silence_duration Trailing silence before an utterance is finalised, in seconds. Clamped to 0.44 to 1.2, and left to the server’s own 0.44 when unset.

How it maps onto Pipecat

Partial transcripts arrive as InterimTranscriptionFrames and endpointed finals as TranscriptionFrames, which is the contract every Pipecat STT service follows. Interruption handling and turn taking work unchanged, so swapping another provider for Bodhi does not touch the rest of the pipeline.

Navigation

Type to search…

↑↓ navigate↵ selectEsc close