Open-source project
rapidaai/voice-ai avatar
rapidaai/voice-ai

Rapida voice-ai: a self-hosted Go orchestration layer for real-time voice agents

Rapida is an open-source, end-to-end voice AI orchestration platform for building real-time conversational voice agents with audio streaming, STT, TTS, VAD, multi-channel integration, agent state management, and observability.

738 stars121 forksGoNOASSERTION

At a glance

What is it?
Rapida is an open-source Go platform that wires audio streaming, STT, TTS, VAD, telephony and agent state into one orchestration service you can run with Docker Compose. The appeal is ownership of the audio path; the cost is a 16GB stack and a README that stops short of production operations.
Who is it for?
Adopt Rapida if you are building white-label or private voice agents and need the audio path, credentials and call logs inside your own infrastructure; the Go services, gRPC protos and per-service YAML configs give you that control. Do not adopt it if you want a hosted agent builder with a managed SLA, or if your machine cannot spare 16GB of RAM, since the README lists that as a prerequisite for all services.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Rapida voice-ai solves, and for whom

Most voice agent stacks force a choice: a hosted agent builder that owns your prompts, keys and call recordings, or a pile of SDKs you glue together yourself. Rapida positions itself against both. The README frames the project around three principles, ownership, control and scale, and names two audiences: agencies that need ownership over white-label client deployments, and enterprises that need scale, control and deploy-anywhere flexibility. The repository backs that framing with structure rather than marketing alone. There is a platform (the docker-compose stack with UI, gateway and four APIs) and a framework (the Go packages under pkg/, the protobuf definitions under protos/, and the SDK examples in examples/golang, examples/nodejs, examples/python, examples/react and examples/widget).

That split matters when you evaluate it. If you want a drop-in voice bot for a personal project, the surface area here is far larger than you need. If you are an agency shipping the same agent to five clients under five brands, the per-service YAML configuration and the self-hosted deployment boundary are the actual product. The repository also carries SECURITY.md, CONTRIBUTING.md and DEVELOPMENT_PROCESS.md at the top level, which suggests the maintainers expect outside contributors and security reports rather than treating the code as a demo.

How the orchestration actually fits together

The architecture is a set of Go services behind an nginx gateway, communicating over gRPC. The README states the project is written in Go and uses gRPC for bidirectional communication, and go.mod confirms the transport and media libraries: grpc-go, grpc-web for browser clients, pion/webrtc, pion/rtp, pion/rtcp and pion/dtls for real-time media, sipgo for SIP telephony, and gorilla/websocket alongside them. Speech providers are compiled in rather than proxied: the module requires deepgram-go-sdk, openai-go, anthropic-sdk-go, cohere-go, the Google speech and texttospeech clients, and the Microsoft cognitive-services-speech-sdk-go. That is a deliberate design choice. Provider selection happens in your YAML config, not in a separate adapter service you maintain.

State and storage are equally explicit in docker-compose.yml. PostgreSQL 15 and Redis 7 run as containers on an api-network bridge, with healthchecks that require four databases to answer pg_isready: web_db, assistant_db, integration_db and endpoint_db. That tells you the service boundaries are also data boundaries. The assistant API, endpoint API, integration API and web API each get their own database rather than sharing one schema. OpenSearch appears only in the knowledge variant of the stack, which is a sensible separation: retrieval is opt-in, not a cost every deployment pays. The tooling layer is visible in go.mod too, with mark3labs/mcp-go present, so MCP-style tool invocation is part of the dependency graph rather than an afterthought.

Installing Rapida with Docker Compose and just

The README gives a four-command quick start and states the prerequisites plainly: Docker and Docker Compose, the Just command runner, and 16GB or more of RAM for all services. Go 1.25.13, Node.js 22 and Yarn 1.22.22 are only needed if you intend to run `just ci`, so you can skip a local toolchain for a first look.

Clone the repository and run the setup and build recipe. The README shows these as separate steps in one block:

bash
git clone https://github.com/rapidaai/voice-ai.git && cd voice-ai
just setup-local build-all
just up-all

`just setup-local build-all` prepares the environment and builds images; `just up-all` starts the stack. After that, `docker compose ps` should list the containers, and the README names the endpoints you can reach: the UI on http://localhost:3000, the nginx API gateway on http://localhost:8080, the Assistant API on http://localhost:9007, the Endpoint API on http://localhost:9005 and the Integration API on http://localhost:9004. The Web API is internal-only by default and reachable only on the container network.

Before the stack is useful you need provider credentials. The README directs you to edit YAML files before starting, and lists them with their ports:

yaml
# docker/assistant-api/assistant.yml   (port 9007)
# docker/endpoint-api/endpoint.yml     (port 9005)
# docker/integration-api/integration.yml (port 9004)
# docker/web-api/web.yml               (port 9001)
# docker/document-api/config.yaml      (port 9010)

The README says to add API keys for OpenAI, Anthropic, Deepgram, Twilio and similar providers in these files. It does not enumerate the individual keys, so read each file rather than guessing names. To tear the stack down, `just down-all` stops the services, and `just logs-all` or a service-specific recipe such as `just logs-assistant` shows output when something fails to come up.

Where the quick start runs out of road

The README's troubleshooting section covers two cases: a port already in use, and services that will not start. That is a short list for a stack with a gateway, four APIs, a UI, PostgreSQL, Redis and an optional OpenSearch cluster. There is no documented rollback path, no stated upgrade procedure between the build tags, and no guidance on what happens to the PostgreSQL volume under ${HOME}/rapida-data/assets/db when you pull a newer build. The compose file pins images by digest, which is good for reproducibility, but it also means you are the one deciding when to move those digests.

The 16GB RAM figure is the other hard edge. The README states it as a prerequisite for all services, and it is not the kind of number you can negotiate with by closing browser tabs: you are running a substantial set of containers, and OpenSearch pushes it higher. A laptop with 8GB will not run this comfortably, and that is a legitimate reason to look elsewhere for evaluation work.

Finally, the project is candid about one piece of its own history: the document-api is described as deprecated and intentionally excluded from CI and release packaging. If you were planning to build retrieval on that service, the release notes are telling you not to. The knowledge path now runs through `just up-all-with-knowledge`, which brings up OpenSearch and a Document API on http://localhost:9010, so the naming in the repository is not perfectly consistent and you should confirm which component you are actually deploying.

Development workflow and the CI contract

Rapida's contribution path is more defined than most projects at this stage. `just ci` runs the same checks used by pull requests, and the README notes that it bootstraps pinned Python, commitlint and shellcheck tooling into local cache directories, with Docker still required for image, integration and smoke-test stages. The justfile imports separate recipe files (just/docker.just, just/images.just, just/run.just, just/development.just, just/ci.just), so the surface is organised rather than one monolithic script.

The release mechanism is worth understanding before you pin a version. According to the README, after Continuous Integration succeeds on a push to main, GitHub creates an immutable tag named build-YYYYMMDD-<short-merge-commit-sha>, and the matching prerelease carries packages for web-api, integration-api, endpoint-api, assistant-api and ui plus a SHA256SUMS file. The recent release list matches that pattern: build-20260909-ca64535, build-20260908-f8ae575 and build-20260908-9d6c116. Three builds in two days is a fast cadence, and it means the tags are merge artifacts rather than curated versions. If you need something stable to pin, you are choosing a commit and living with it.

For narrower work, the README documents starting single components: `just up-db`, `just up-ui`, `just up-assistant`, and `just rebuild-assistant` or `just rebuild-all` after code changes. Running without Docker is possible but the README is explicit that PostgreSQL, Redis and OpenSearch must be provided separately, and that nginx configuration should be copied from ./nginx/nginx.conf. It also requires a writable storage path at /opt/rapida-data/assets/workflow.

Rapida against LiveKit Agents and Pipecat

The closest comparisons in this space are LiveKit Agents and Pipecat. Both are open source and both target real-time voice, so the difference is in where the orchestration lives. LiveKit Agents is built on top of the LiveKit WebRTC infrastructure: you get a media server as the centre of gravity, and agents join rooms as participants. Pipecat is a Python framework that composes pipelines of STT, LLM and TTS processors in-process, which makes it easy to reason about and easy to extend with Python libraries.

Rapida takes a third position. The orchestration is a Go service mesh behind nginx, with gRPC between components and per-service PostgreSQL databases. The README's framing is deployment-oriented rather than library-oriented: white-label client deployments, private infrastructure, credentials you keep. That is a real difference in kind. With Pipecat you import a framework into your application; with Rapida you run a platform and point your channels at it. The trade-off is that you inherit more operational surface. A Pipecat pipeline failing is a Python stack trace in your process. A Rapida assistant-api failing is a container that needs its YAML checked and its database reachable.

If your team already lives in Python and wants to iterate on prompt and pipeline logic quickly, Pipecat is the more direct tool. If you need SIP telephony and multi-tenant deployment boundaries in one system, Rapida's inclusion of sipgo and per-service config is closer to the shape of that problem. Neither answer is universal, and the README does not publish latency comparisons, so do not pick on throughput claims you cannot verify.

Licence, maintenance and what a build tag costs you

The repository's licence is reported as NOASSERTION, which means the automated classifier could not map LICENSE.md to a recognised identifier. That is not a statement about what the licence says, only about what tooling could determine. Read LICENSE.md directly before you build a commercial product on it, and if the terms are unclear, that is a question for your own legal counsel rather than a blog post. The practical risk is that a permissive-looking project turns out to carry conditions you did not expect.

On maintenance, the last push was on 2026-09-14, and the most recent release tag is build-20260909-ca64535. The repository is not archived. That combination describes a project under current development, with the caveat that its releases are automated build tags rather than versioned milestones. Upgrading means choosing a tag, reading the diff, and accepting that the compose file pins image digests you may need to move yourself. There is no documented migration process for the PostgreSQL volumes, so a database schema change between builds is something you would discover rather than plan for.

The upgrade cost is therefore mostly yours: pin a build tag, snapshot ${HOME}/rapida-data/assets/db before you move, and check the release packages for the service you run. The project gives you SHA256SUMS for the artifacts, which is a genuine convenience, and it gives you nothing for the data layer.

Editorial conclusion

Adopt Rapida if you are building white-label or private voice agents and need the audio path, credentials and call logs inside your own infrastructure; the Go services, gRPC protos and per-service YAML configs give you that control. Do not adopt it if you want a hosted agent builder with a managed SLA, or if your machine cannot spare 16GB of RAM, since the README lists that as a prerequisite for all services. Before committing, verify three things yourself: that your model and telephony credentials work in docker/assistant-api/assistant.yml and the other YAML files, that the NOASSERTION licence in LICENSE.md grants the rights your deployment needs, and that the knowledge services you depend on are not the deprecated document-api that CI and release packaging exclude.

Frequently asked questions

How do I install Rapida voice-ai?

The README gives a four-command quick start: clone the repository, run `just setup-local build-all`, then `just up-all`. Docker and Docker Compose plus the Just command runner are prerequisites, and the README states 16GB or more of RAM is needed for all services.

How do I set up Rapida voice-ai for a first run?

After `just up-all`, the UI is on http://localhost:3000 and the nginx gateway on http://localhost:8080. Before the stack is useful you edit the YAML files under docker/, such as docker/assistant-api/assistant.yml on port 9007, and add API keys for providers like OpenAI, Anthropic, Deepgram and Twilio.

Is Rapida voice-ai free or paid?

The repository is open source and self-hostable, so the software itself carries no listed fee, but the README does not describe pricing for the managed option. Your real cost is infrastructure plus whatever you pay the STT, TTS and LLM providers whose keys you put in the config files.

How do I use Rapida voice-ai?

You run the platform stack and point your channels at it. The README documents the UI on port 3000 and the API gateway on port 8080, with per-service YAML files under docker/ holding provider credentials.

Is Rapida voice-ai safe to use?

The repository ships a SECURITY.md file and a CodeQL scanning workflow, which the README references. The README does not make broader security claims, so treat deployment hardening and credential handling as your responsibility.

Official sources

  1. Issues
  2. Project website
  3. rapidaai/voice-ai on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/rapidaai-voice-ai.svg)](https://hysenlabs.com/projects/rapidaai-voice-ai)