Framework
pipecat-ai/pipecat avatar
pipecat-ai/pipecat

Pipecat: an open framework for voice agents and realtime multimodal AI

Open Source framework for voice agents, multimodal apps, and realtime AI. Maintained by Daily and the community.

15,980 stars2,781 forksPythonBSD-2-Clause

At a glance

What is it?
Pipecat is a BSD-2-Clause Python framework for building voice agents, multimodal apps and realtime AI, chaining speech-to-text, an LLM and text-to-speech into a low-latency pipeline you can point at many vendors. It is maintained by Daily and the community.
Who is it for?
Adopt Pipecat if you are building realtime voice or multimodal AI agents and want an open, vendor-neutral pipeline that chains speech-to-text, an LLM and text-to-speech with the freedom to swap each. Do not expect it to remove latency tuning or the cost of the providers it orchestrates, which are your responsibility.
Can I use it commercially?
Yes. BSD-2-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Pipecat is for

Building a voice agent that listens, thinks and speaks in real time means wiring together speech-to-text, a language model and text-to-speech, handling audio streaming, interruptions and latency along the way. Pipecat is an open framework that provides that plumbing. It is a Python framework for voice agents, multimodal apps and realtime AI, maintained by Daily and the community, and it lets you compose a pipeline of processing stages that move audio and text between services. The audience is developers building conversational voice applications, phone agents, in-app assistants, live multimodal experiences, who want an open, vendor-flexible framework rather than a single closed product. Its scope is the realtime pipeline: capturing audio, transcribing it, running it through an LLM, synthesizing speech, and managing the turn-taking and streaming that make a spoken conversation feel responsive rather than laggy.

A pipeline of pluggable services

The mechanism is a pipeline of composable processors through which frames of audio and text flow. You assemble stages, a transport for the audio connection, a speech-to-text service, an LLM service, a text-to-speech service, and Pipecat streams data between them while handling the realtime concerns: partial transcripts, letting a user interrupt the bot, and joining fragments so speech sounds continuous. Its defining trait is breadth of integrations: it supports many vendors for each stage, so you can mix, say, one provider's transcription with another's model and a third's voice, and swap any of them without rewriting the app. The changelog shows this depth, for example a text-to-speech service using a vendor's continuation API so fragments within one turn share context and prosody does not reset. That vendor-neutral, swappable-stage design is what separates a framework from a single-vendor SDK.

Installing Pipecat

Pipecat is a Python package on PyPI, installed with pip, typically with extras selecting the service integrations you need:

bash
pip install pipecat-ai

Because it integrates many vendors, you install the extras for the specific speech-to-text, LLM and text-to-speech providers you plan to use, and provide their API keys, and an env.example in the repository shows the configuration. There is also a Pipecat CLI that ships with the package. The examples directory is the fastest path to a working agent: it contains runnable pipelines you adapt, and the first real use is running one of those examples end to end, speaking to it and hearing it respond, which confirms your transport, chosen services and keys are wired correctly before you build your own pipeline on top.

Where a realtime framework demands work

The limitations are those of realtime, multi-vendor software. Latency is the constant concern: a natural conversation needs the whole loop, capture to transcription to model to speech, to complete fast, and each vendor you choose adds its own latency and failure modes, so tuning and provider selection matter and are your responsibility. It orchestrates external services rather than running models itself, so you bring and pay for the speech-to-text, LLM and text-to-speech providers, and the cost and reliability of a deployed agent track those choices. Building a production voice agent, handling interruptions, errors, telephony and scale, is real engineering that the framework supports but does not eliminate. And the breadth of integrations means learning which combination fits your latency and quality needs. These are inherent to the problem, not defects of Pipecat, which exists precisely to manage this complexity.

Pipecat versus LiveKit Agents

A frequent comparison is LiveKit Agents, another framework for realtime AI agents. Both let you build voice agents from pluggable speech, model and voice components, so the difference is largely in transport and ecosystem. LiveKit Agents is tightly integrated with LiveKit's own realtime transport and platform, which is a strength if you are on LiveKit. Pipecat is transport-flexible and vendor-neutral by design, maintained by Daily but not locked to one provider, with a wide catalog of integrations across the pipeline, which suits teams that want to choose or change each component and transport. The choice depends on your stack. Pick LiveKit Agents if you are building on LiveKit and want that tight integration; pick Pipecat when you want an open framework that lets you assemble and swap speech, model, voice and transport pieces freely, which is the flexibility its design centers on.

BSD-2-Clause and active development

Pipecat is BSD-2-Clause licensed, a permissive license, so it is freely usable including commercially. The last push was on 2026-09-15, with a v1.10.0 in mid-September 2026, and the changelog shows steady work across its many service integrations, which matters because the value of a vendor-neutral framework is keeping those integrations current as providers evolve. It is documented at pipecat.ai with an examples directory and a CLI. Adopt it when you are building realtime voice or multimodal agents and want an open, swappable-vendor pipeline rather than a closed product, install it with the extras for your chosen providers, run an example end to end to validate latency and wiring, and treat provider selection as a first-class decision since it drives both cost and the responsiveness of your agent.

Editorial conclusion

Adopt Pipecat if you are building realtime voice or multimodal AI agents and want an open, vendor-neutral pipeline that chains speech-to-text, an LLM and text-to-speech with the freedom to swap each. Do not expect it to remove latency tuning or the cost of the providers it orchestrates, which are your responsibility. Install it with pip install pipecat-ai plus the extras for your chosen services, run an example end to end to validate wiring and latency, and treat provider choice as central since it drives responsiveness and cost.

Frequently asked questions

What is Pipecat?

Pipecat is a BSD-2-Clause Python framework, maintained by Daily and the community, for building voice agents, multimodal apps and realtime AI by chaining speech-to-text, an LLM and text-to-speech into a low-latency, vendor-flexible pipeline.

How do I install it?

Install the PyPI package with pip install pipecat-ai, adding the extras for the specific speech-to-text, LLM and text-to-speech providers you use, and supply their API keys. The repository's examples are the fastest way to a working agent.

How does it compare to LiveKit Agents?

Both build voice agents from pluggable components. LiveKit Agents integrates tightly with LiveKit's transport and platform; Pipecat is transport-flexible and vendor-neutral, with a wide catalog of integrations you can swap freely.

Official sources

  1. License: BSD-2-Clause
  2. pipecat-ai/pipecat on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/pipecat-ai-pipecat.svg)](https://hysenlabs.com/projects/pipecat-ai-pipecat)