huggingface/speech-to-speech: a self-hosted voice agent pipeline behind an OpenAI Realtime WebSocket
GitHub describes it as Build local voice agents with open-source models. The repository metadata lists Python as its primary language. The metadata lists the Apache-2.0 license. This article stays within the project description and details documented in the GitHub repository README.
At a glance
- What is it?
- Hugging Face's speech-to-speech package wires Silero VAD, an interchangeable STT model, an OpenAI-compatible LLM and a local TTS into one server that speaks the OpenAI Realtime event set. It is a good fit when you want the whole loop on your own hardware, and a poor one when you want a hosted drop-in.
- Who is it for?
- Adopt it if you need a voice loop that stays on your own hardware and you are willing to pick STT and TTS backends per platform. Skip it if you want a managed endpoint, a voice changer, or a translation product: the README describes a conversation backend, not those.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What huggingface/speech-to-speech actually solves
The project packages the four stages of a spoken conversation into one process: voice activity detection, speech recognition, a language model call, and speech synthesis. Each stage runs in its own thread and the stages are joined by queues, so the output of one feeds the input of the next without a central scheduler. The README calls it a low-latency, fully modular voice-agent pipeline and writes the order as VAD -> STT -> LLM -> TTS.
The audience is narrow and fairly technical. You are expected to run a Python service, choose model backends, and either supply an OpenAI API key or stand up your own inference server. The README states the pipeline runs in production as the conversation backend for Reachy Mini robots, which tells you the intended shape of a deployment: a long-lived server that many clients connect to, not a one-shot command. If you want to click a button and hear a voice, this is not that project.
The four-thread cascade and the Realtime event surface
Silero VAD v5 decides where the user's turn begins and ends. The transcript then goes to whichever STT backend you selected, with optional live partial transcripts so the client can show words before the turn closes. The LLM slot speaks OpenAI-compatible protocols, which is the design decision that matters most: the same server can point at a hosted provider, at Hugging Face Inference Providers, or at a vLLM or llama.cpp server you run yourself. Text and tool calls stream out, and the TTS stage synthesizes audio back to the client.
What clients see is the core OpenAI Realtime GA event set, exposed over WebSocket and WebRTC. That is the integration contract. The README states the official OpenAI Agents SDK is tested over both stock transports, and points to its Realtime API section for the tested surface. Read that qualifier carefully: the core event set is implemented, not the entire specification. If your client depends on an event outside that set, the README does not promise it.
Running each stage in its own thread is what makes the swap story work, and it is also the source of the project's complexity. Timing, barge-in and turn boundaries are spread across threads rather than expressed in one place.
Installing speech-to-speech and holding a first conversation
The package requires Python 3.10 or newer. The default install covers the standard realtime path: Parakeet TDT for STT, an OpenAI-compatible API for the language model, and Qwen3-TTS for output, using the GGML backend by default on non-macOS platforms and mlx-audio on Apple Silicon. macOS and non-macOS dependencies are resolved through platform markers in pyproject.toml, so the same command works on both.
pip install speech-to-speech
export OPENAI_API_KEY=...
speech-to-speech serveThe server starts an OpenAI Realtime-compatible endpoint at ws://localhost:8765/v1/realtime. From a second terminal you can talk to it with the packaged client, which is the fastest way to confirm that VAD, STT, LLM and TTS are all responding.
speech-to-speech talk --url ws://127.0.0.1:8765/v1/realtimeIf you would rather start the server and the packaged microphone and speaker client together, one command does it.
speech-to-speech localKeeping the LLM on your own machine changes only the backend flags. The README gives this llama.cpp example for serving Gemma 4, followed by the matching server invocation.
llama-server -hf ggml-org/gemma-4-E4B-it-GGUF -np 2 -c 65536 -fa on --swa-fullspeech-to-speech serve \
--model_name "ggml-org/gemma-4-E4B-it-GGUF" \
--responses_api_base_url "http://127.0.0.1:8080/v1" \
--responses_api_api_key ""For a containerised setup, docker-compose.yml defines two services. The llama service runs ghcr.io/ggml-org/llama.cpp:server-cuda on port 8080 with the Gemma 4 GGUF weights, and the pipeline service runs speech-to-speech serve on port 8765 with --llm_backend responses-api pointing at http://llama:8080/v1. Both reserve an NVIDIA device. The Dockerfile builds from nvidia/cuda:12.8.1-cudnn-runtime-ubuntu24.04 and downloads the NLTK punkt_tab and averaged_perceptron_tagger_eng data during the build.
The CUDA and glibc trap in the Qwen3-TTS default
The default TTS path is where an install most often fails. On Linux the Qwen3-TTS GGML backend comes from faster-qwen3-tts[ggml], and the README states its default qwentts-cpp-python wheel on PyPI targets CUDA 12.8 and manylinux_2_39, which it illustrates with Ubuntu 24.04. If your CUDA runtime or glibc is older, the plain pip install can resolve to a wheel your machine cannot load.
The documented workaround is to install a matching wheel from the Hugging Face wheelhouse before installing speech-to-speech itself. The README lists three variants: 0.3.1+cu130, 0.3.1+cu124, and a CPU-only 0.3.1+cpu fallback, each with its own -f index URL. Note the ordering: the wheel goes in first, then the package. If you skip the first step, pip has already chosen a wheel by the time you notice.
There is a second escape hatch. Passing --qwen3_tts_backend torch switches to the previous CUDA-graphs implementation instead of GGML. The README does not say what performance difference to expect between the two, so treat the flag as a compatibility lever rather than a tuning knob.
Where speech-to-speech is the wrong tool
The README does not document rollback, and it does not describe a fallback path when the LLM backend is unreachable. In a cascade, a stalled stage is a stalled conversation, and the only recovery mechanism the documentation gives you is restarting the process. If you need a service-level guarantee around that, you are building it yourself.
Dependency conflicts are real and documented. DeepFilterNet, used for optional audio enhancement in VAD, requires numpy<2 and conflicts with Pocket TTS, which requires numpy>=2. The README's instruction is to install DeepFilterNet manually only in environments where Pocket TTS is not used. That is a genuine either-or, not a footnote.
Platform coverage is uneven by design. Parakeet TDT runs on CUDA and CPU through nano-parakeet and on Apple Silicon through MLX, while Lightning Whisper MLX is Apple Silicon only. A team standardising on one backend across Linux servers and Mac laptops may find the backend list splits along platform lines.
Finally, the name invites the wrong crowd. This is not a voice changer and not a translation product. It is a conversation backend that listens, thinks and answers. Searches for a free online speech-to-speech tool will land here and find a pip package and a server command.
Compared with hosted realtime APIs
The obvious alternative is a hosted realtime voice API, where you send audio and receive audio and never manage a model. The difference is not quality, it is where the loop runs and what you can change. A hosted API fixes the STT, LLM and TTS models behind one endpoint. Here every stage has multiple interchangeable backends selected by CLI flags, and the README states the code is designed for easy modification, with a focus on models available through Transformers and the Hugging Face Hub.
That flexibility has a cost the hosted option does not charge: you own the GPU, the wheel compatibility, the model downloads and the thread timing. The repository also ships an archive/ directory holding deprecated implementations such as MeloTTS that are no longer wired into the CLI, which is a reminder that backend churn is part of the project's normal life.
A second alternative is assembling the same four stages yourself from Silero VAD, a Whisper model and a TTS library. That gives you total control and no upgrade path. The value here is the Realtime event surface, the queue-based threading model, and the tested transport compatibility with the OpenAI Agents SDK, none of which you get for free by gluing libraries together.
Licence, maintenance and the cost of upgrading
The package is Apache-2.0, declared in pyproject.toml with license-files = ["LICENSE"]. That covers the pipeline code. It does not cover the model weights you pull at runtime: Parakeet TDT, Qwen3-TTS, Gemma 4 and the Whisper variants each carry their own terms, and the README links out to them rather than restating them. Check the model cards before shipping, and note that this is a description of what the repository declares, not legal advice.
The repository is not archived, and the last push was on 2026-08-05, which is recent. Recent releases run v0.2.10 on 2026-06-11, v0.2.11 on 2026-08-03 and v0.2.12 on 2026-08-05. The release cadence is uneven: two releases two days apart, then a two-month gap. pyproject.toml declares version 1.0.0 while the release tags sit at v0.2.12, so do not read the package version as the release number.
Upgrade cost concentrates in the optional extras and the wheel matrix. The pip extras are named individually (kokoro, pocket, chattts, omnivoice, faster-whisper, whisper-mlx, paraformer, mlx-lm), and each pulls its own dependency tree. The numpy split between DeepFilterNet and Pocket TTS means an upgrade that moves numpy can break one of the two. Pinning your extras and your qwentts-cpp-python wheel is the practical defence.
Editorial conclusion
Adopt it if you need a voice loop that stays on your own hardware and you are willing to pick STT and TTS backends per platform. Skip it if you want a managed endpoint, a voice changer, or a translation product: the README describes a conversation backend, not those. Before committing, install the package on your target CUDA or glibc combination, run speech-to-speech local, and confirm that the Qwen3-TTS wheel you install matches your runtime.
Frequently asked questions
Is there a free AI speech to speech tool available?
Yes. speech-to-speech is Apache-2.0 and installs with pip install speech-to-speech. It can run fully locally if you point the LLM slot at a vLLM or llama.cpp server instead of a hosted provider.
What is speech to speech?
In this project it means a pipeline that takes spoken audio in and produces spoken audio out, running through VAD, STT, an LLM and TTS in sequence. The README writes the order as VAD -> STT -> LLM -> TTS.
Which STT model is best for speech-to-speech?
The README does not rank the STT backends. It lists Parakeet TDT as the default, with Whisper through Transformers, Faster Whisper, Lightning Whisper MLX and MLX Audio Whisper available as alternatives, and notes that Lightning Whisper MLX is Apple Silicon only.
How do I do speech to speech with huggingface/speech-to-speech?
Install the package, set OPENAI_API_KEY, and run speech-to-speech serve, which starts an OpenAI Realtime-compatible endpoint at ws://localhost:8765/v1/realtime. A second command, speech-to-speech talk --url ws://127.0.0.1:8765/v1/realtime, connects the packaged client to it.
Is voice cloning illegal?
The README does not address voice cloning or its legality. It documents the pipeline components, the backends and the Realtime event surface, and leaves model licensing to the individual model cards it links to.
What is speech to speech AI?
Here it is a cascade of four models, Silero VAD v5, an STT backend, an OpenAI-compatible LLM and a TTS backend, each running in its own thread. The result is exposed to clients as the core OpenAI Realtime GA event set over WebSocket and WebRTC.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/huggingface-speech-to-speech)
Community notes