---
title: "Pipecat"
description: "Drop Bodhi into the STT slot of a Pipecat voice agent pipeline."
---

> Documentation Index
> Fetch the complete documentation index at: https://docs.navana.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Pipecat

[Pipecat](https://github.com/pipecat-ai/pipecat) builds voice agents as a
pipeline of services. Bodhi provides the speech-to-text stage through the
Python SDK, so transcription behaves like any other Pipecat STT service.

## Install

```bash
pip install "bodhi-api-sdk[stt]"
```

Needs pipecat-ai 1.4 or newer. Verified against 1.4.0, 1.8.1 and 1.11.0.

> **Uninstall bodhi-sdk first**
>
> `bodhi-api-sdk` and the older `bodhi-sdk` both install a module named
> `bodhi`, so they cannot be installed side by side. Run
> `pip uninstall bodhi-sdk` before installing this one, or you end up with a
> mix of the two.

## Add it to a pipeline

```python
import os

from pipecat.pipeline.pipeline import Pipeline

from bodhi.integrations.pipecat_stt import BodhiHotword, BodhiSTTService

stt = BodhiSTTService(
api_key=os.environ["BODHI_API_KEY"],
model="hi-banking-v2-8khz",
settings=BodhiSTTService.Settings(
    parse_number=True,
    endpoint_silence_duration=0.6,
    hotwords=[BodhiHotword("बजाज फिनसर्व", 2.0)],
),
)

pipeline = Pipeline([
transport.input(),
stt,
context_aggregator.user(),
llm,
tts,
transport.output(),
])
```

`api_key` and `model` are the only required arguments. Everything else has a
default, including `url`, which points at `wss://stt.navana.ai`.

## Arguments

| Argument | Required | Description |
| --- | --- | --- |
| `api_key` | Yes | Your platform API key, sent as the `x-api-key` header. |
| `model` | Yes | Any transcription model. See [Models and languages](/speech-to-text/models). |
| `url` | No | Override the endpoint, usually to point at a different deployment. |
| `sample_rate` | No | Input rate in Hz. Defaults to the pipeline's own. |
| `language` | No | Language tag for emitted frames. Defaults to the model's prefix. |
| `interim_results` | No | Emit `InterimTranscriptionFrame`s from partials. Defaults to on. |
| `aux` | No | Attach timing and confidence metadata to each frame. |
| `min_confidence` | No | Drop finals below this utterance confidence. Defaults to `0.5`; pass `0` to keep everything. |
| `transaction_id` | No | Correlation id for the session. A fresh UUID per connection when omitted. |
| `settings` | No | Runtime-updatable recognition settings, below. |

### Settings

Pass these through `BodhiSTTService.Settings`. They can be updated while the
agent is running.

| Setting | Description |
| --- | --- |
| `model` | The transcription model. Changing it reconnects the session. |
| `language` | Language used to tag emitted frames. Presentational only, since Bodhi infers the language from the model. |
| `hotwords` | Phrase boosting, as `BodhiHotword(phrase, score)`. `score` is optional, and roughly `1.0` to `3.0` is useful. See [Advanced features](/speech-to-text/advanced-features). |
| `parse_number` | Normalise spoken numbers, dates, and currency in the returned text. |
| `endpoint_silence_duration` | Trailing silence before an utterance is finalised, in seconds. Clamped to `0.44` to `1.2`, and left to the server's own `0.44` when unset. |

## How it maps onto Pipecat

Partial transcripts arrive as `InterimTranscriptionFrame`s and endpointed
finals as `TranscriptionFrame`s, which is the contract every Pipecat STT
service follows. Interruption handling and turn taking work unchanged, so
swapping another provider for Bodhi does not touch the rest of the pipeline.

## Related

- [Python SDK](/speech-to-text/streaming) — Using the SDK directly, without Pipecat.
- [Models and languages](/speech-to-text/models) — Every model, and how to pick one.
- [Advanced features](/speech-to-text/advanced-features) — Hotwords, number parsing, endpointing.

Source: https://docs.navana.ai/speech-to-text/pipecat/index.mdx
