Model or dataset
collabora/WhisperLive avatar
collabora/WhisperLive

WhisperLive: a nearly-live Whisper server you run yourself

A nearly-live implementation of OpenAI's Whisper.

4,304 stars601 forksPythonMIT

At a glance

What is it?
WhisperLive wraps OpenAI's Whisper in a WebSocket server with three interchangeable inference backends. It is a good fit when you want streaming transcription on your own hardware and are willing to manage a Python server and a client.
Who is it for?
Adopt WhisperLive if you need streaming transcription on hardware you control and can run a Python server plus a client; the three backends let you move from CPU to NVIDIA or Intel acceleration without changing the client protocol. Do not adopt it if you want a hosted API, a finished desktop app, or zero setup, because the README's path is a venv, a pip install, and a run_server.py invocation.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 21 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What WhisperLive adds to Whisper

Whisper itself is a batch model. You hand it a file, it returns a transcript. WhisperLive's stated purpose is to turn that into a nearly-live transcription application: the README describes it as a real-time transcription application that converts speech input into text output, usable for both live microphone input and pre-recorded audio files. The gap it fills is the server and the streaming loop around the model, not the model. It ships a server (run_server.py), a client (run_client.py), an OpenAI-compatible REST path (client_openai.py), and browser extensions for Chrome and Firefox in the repository tree. The audience is developers and researchers who want live transcription on their own hardware and are comfortable running a Python service. It is not a consumer app, and the packaging metadata classifies it as Development Status 4 - Beta.

Backends and the server-side data flow

The server supports three backends: faster_whisper, tensorrt and openvino. That choice is the main architectural decision. faster_whisper is the default path and the one whose model loading behaviour the README explains in most detail. By default, when the server runs without a specified model, it instantiates a new Whisper model for every client connection. The README is explicit about the trade-off: this lets the server pick a model size based on what the client requests, but it means waiting for the model to load on connect and higher (V)RAM usage. If you serve a custom model with -fw or -trt, the server instantiates that model once and reuses it across connections. For TensorRT, the README recommends the Docker setup and says engines must be built before starting the server. The TensorRT backend uses a C++ session by default; the README notes that if you see repeated CrossAttentionMask warnings or crashes, the --trt_py_session flag switches to the Python session. OpenVINO targets Intel CPUs, iGPUs and dGPUs, and the README says the tested models are the ones OpenVINO uploads to Hugging Face.

Installing WhisperLive and running a first server

The README lists PortAudio as a required system dependency for microphone input via PyAudio, installed by a setup script that picks the right package per platform: portaudio19-dev on Debian/Ubuntu, portaudio-devel on Fedora, and Homebrew's portaudio on macOS. After that, the flow is a 3.12 virtual environment and a pip install.

bash
bash scripts/setup.sh
python3.12 -m venv whisper_env
source whisper_env/bin/activate
pip install whisper-live

With the package installed, the default server is a single command. The example below is the README's faster_whisper invocation, with a port, a client cap, and a connection time limit in seconds.

bash
python3 run_server.py --port 9090 \
                      --backend faster_whisper \
                      --max_clients 4 \
                      --max_connection_time 600

You should see the server start and listen on port 9090. To pin a model instead of letting each client request one, the README shows -fw with a path to a custom faster-whisper model and -c for a cache directory where auto-converted CTranslate2 models are saved.

bash
python3 run_server.py --port 9090 \
                      --backend faster_whisper \
                      -fw "/path/to/custom/faster/whisper/model" \
                      -c ~/.cache/whisper-live/

If you prefer an HTTP-style interface, the README documents an OpenAI REST mode, enabled with --enable_rest and restricted with --cors-origins. The server example uses port 9090 and the client is then invoked as `python3 client_openai.py $AUDIO_FILE`.

Client-side chunking and the streaming example

WhisperLive is not only a server. The repository includes run_client.py, client_openai.py, client_oldapi.py, and examples/manual_audio_chunking.py, which the README's table of contents calls the streaming client for manual audio chunking. That file is the place to look if you want to control how audio is cut before it is sent, rather than relying on the built-in client loop. The README also documents advanced features that live on the server side: word-level timestamps, custom vocabulary or hotwords, speaker diarization, batch inference, and raw PCM input. Taken together, the shape is a WebSocket protocol between a thin client and a model-holding server, with the client free to be a Python script, a browser extension, or your own code. The README does not document the wire protocol in the excerpt available, so treat the bundled clients as the reference implementation if you plan to write your own.

Where WhisperLive is the wrong choice

The default per-connection model loading is the sharpest limitation. If you run the server without -fw or -trt, every new client triggers a model instantiation, so the first seconds of a session are spent loading weights rather than transcribing, and memory grows with concurrency. The README presents this as a deliberate trade-off, not a bug, but it rules out the naive deployment for many short-lived connections. The default OMP_NUM_THREADS is 1, which the README says can be raised with --omp_num_threads; leaving it at the default on a multi-core CPU box underuses the machine. TensorRT is not a drop-in: the README says to build engines first and recommends Docker, and the C++ session has a documented failure mode that requires --trt_py_session. OpenVINO outside Docker requires Intel drivers and the OpenVINO runtime to be installed and configured, with the README pointing elsewhere for that. And the project is a Beta-classified library: if you need a supported SLA, a managed endpoint, or a desktop application with an installer, this is not that.

How it compares to faster-whisper alone

The closest alternative is using faster-whisper directly. faster-whisper is the inference library WhisperLive pins (faster-whisper==1.2.0 in setup.py) and the default backend. Using it alone gives you a Python API and no server: you write your own audio capture, your own chunking, your own concurrency handling, and your own transport. WhisperLive's contribution is exactly that layer, plus the ability to swap in TensorRT or OpenVINO behind the same client-facing interface. The difference matters when you have more than one consumer of transcripts, or when you want to move from CPU to an NVIDIA or Intel accelerator without rewriting the client. If you have a single script that transcribes one stream, the server adds a process and a protocol you do not need. If you have several clients, or you want to change accelerators later, the extra layer is the point.

Maintenance, licence, and upgrade cost

The repository is not archived, and the last push was on 2026-09-10. Releases are roughly quarterly: v0.8.0 on 2026-03-17, v0.9.0 on 2026-06-02, v0.10.0 on 2026-09-07. The licence is MIT, stated in both the repository and setup.py. MIT is permissive, so embedding the server in a commercial product is not the obstacle; the practical cost is the dependency chain. setup.py pins faster-whisper to 1.2.0 and carries version-conditional onnxruntime requirements, and the TensorRT path depends on NVIDIA's TensorRT-LLM repository plus separately built engines. Upgrading WhisperLive therefore means re-checking that pin, the backend wheels, and any prebuilt TensorRT engines, which do not carry across versions automatically. None of this is legal advice; if you redistribute, read the MIT text and the licences of the backends you enable.

Editorial conclusion

Adopt WhisperLive if you need streaming transcription on hardware you control and can run a Python server plus a client; the three backends let you move from CPU to NVIDIA or Intel acceleration without changing the client protocol. Do not adopt it if you want a hosted API, a finished desktop app, or zero setup, because the README's path is a venv, a pip install, and a run_server.py invocation. Before committing, verify that the backend you intend to use actually works on your machine: TensorRT requires building engines first, OpenVINO outside Docker requires Intel drivers and the OpenVINO runtime, and the README points to separate documents for both.

Frequently asked questions

What is WhisperLive?

It is a Python real-time transcription application built on OpenAI's Whisper, described in the README as a nearly-live implementation. It runs a server that holds the model and clients that send audio, and it can transcribe live microphone input or pre-recorded files.

How do I install WhisperLive?

The README's steps are to run scripts/setup.sh for PortAudio, create a Python 3.12 virtual environment, activate it, and run pip install whisper-live. On Debian/Ubuntu the script installs portaudio19-dev, on Fedora portaudio-devel, and on macOS it uses Homebrew.

How do I use WhisperLive?

Start the server with run_server.py, choosing a backend such as faster_whisper, then connect a client. The README shows a server on port 9090 with --max_clients 4 and --max_connection_time 600, and clients including run_client.py, client_openai.py for the REST interface, and the browser extensions.

Is Whisper AI safe to use?

The repository does not discuss safety or data handling. What it does document is that WhisperLive runs as a server you host yourself, with options such as --max_clients and --max_connection_time, so audio handling is determined by your own deployment rather than by a third-party service.

Is there a free app that can transcribe audio live?

WhisperLive is open source under the MIT licence and installs from pip as whisper-live, so there is no licence fee for the software itself. It is a server and client toolkit rather than a packaged app, so you supply the machine and the deployment.

Official sources

  1. collabora/WhisperLive on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/collabora-whisperlive.svg)](https://hysenlabs.com/projects/collabora-whisperlive)