Library / SDK
KoljaB/RealtimeSTT avatar
KoljaB/RealtimeSTT

RealtimeSTT: a Python speech-to-text library built around voice activity detection

A robust, efficient, low-latency speech-to-text library with advanced voice activity detection, wake word activation and instant transcription.

10,158 stars861 forksPythonMIT

At a glance

What is it?
RealtimeSTT packages VAD, wake words and streaming transcription behind one AudioToTextRecorder class. The install is small, but the engine choice is where the real work sits.
Who is it for?
RealtimeSTT fits Python projects that need a recorder loop with VAD, wake words and optional realtime text without assembling those parts yourself. It is the wrong choice if you only need one offline transcription of a finished file, since faster-whisper alone does that with fewer moving parts, and it is also wrong if you need Python 3.13, which the repository states is not a release target.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 12 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem RealtimeSTT solves for Python voice applications

Turning a microphone into text is not one problem, it is four: deciding when someone is speaking, deciding when they stopped, transcribing the segment, and doing something useful while the segment is still open. RealtimeSTT's README lists voice activity detection with WebRTC VAD and Silero VAD, final and realtime transcription with selectable engines, optional wake word activation through Porcupine or OpenWakeWord, and event callbacks for recording, VAD, realtime text, transcription and wake word state. Those are the four problems, wrapped in one class.

The audience is narrow but real. The README names assistants, dictation tools, browser streaming servers and prototypes. The common thread is a program that is already running and needs to react to speech as it arrives, rather than a batch job that receives a finished audio file. A dictation tool wants partial text on screen while the speaker is still talking. An assistant wants a wake word before it starts listening seriously. A browser streaming server wants a session per connected client and a bounded amount of inference behind it.

What RealtimeSTT does not try to be is a model. The default general-purpose path uses faster_whisper, and the README points at other engines through install extras when their optional dependencies and models are present. The library is the plumbing around the model: audio capture, segmentation, a worker process, callbacks, and a server if you want one.

How AudioToTextRecorder moves audio to text

The central object is AudioToTextRecorder. Constructing it starts the machinery; calling text() blocks until an utterance has been detected, recorded and transcribed, then returns the string. The README's microphone example is exactly that shape, with a context manager around it.

Underneath, the pieces are visible in the repository layout and the documented features. Audio arrives either from a microphone or from your own chunks when use_microphone=False. A VAD layer watches the stream and marks speech boundaries. Realtime transcription produces interim text while the turn is open, and final transcription produces the authoritative transcript at the end. Model work happens in a separate process, which is why the README insists on the if __name__ == "__main__": guard, especially on Windows.

The engine profiles in the README are the most interesting design decision. On CUDA the guidance is to keep the established faster_whisper setup. On CPU for production streaming on Linux x86-64, the recommended pairing is sherpa-onnx-nemotron-3.5-asr-streaming-0.6b-560ms-int8 for replaceable realtime text plus sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8 for the single authoritative final transcript. The README explains the reasoning directly: Nemotron processes only new audio frames during the turn, and Parakeet refines the complete turn once at finalization. That is a two-model split between latency and quality, and it avoids the pattern of repeatedly retranscribing a growing buffer.

For applications that are not local, the repository also ships a FastAPI server with versioned HTTP and WebSocket contracts, session isolation, bounded shared inference resources, authentication, and readiness and capabilities endpoints. That is a different deployment shape from the library, and it lives in RealtimeSTT_server/ with RealtimeSTT_server/PRODUCTION_SERVER.md as the guide.

Installing RealtimeSTT and running a first microphone transcription

The README's default install asks for the faster-whisper extra. The setup.py install guide also lists a recommended extra and a base install with no transcription engine at all, so pick deliberately: the base package includes microphone support, WebRTC VAD, recorder VAD logic and websocket dependencies, but not faster-whisper, Porcupine, OpenWakeWord or another optional ASR backend.

bash
pip install "RealtimeSTT[faster-whisper]"

On Linux, PortAudio headers come first. The README gives these two commands.

bash
sudo apt-get update
sudo apt-get install python3-dev portaudio19-dev

On macOS the README uses Homebrew instead.

bash
brew install portaudio

With the package in place, the shortest real use is the microphone example. It waits for speech, stops after the detected utterance, and prints the final transcript.

python
from RealtimeSTT import AudioToTextRecorder

if __name__ == "__main__":
    with AudioToTextRecorder() as recorder:
        print("Speak now")
        print(recorder.text())

Expect the first run to spend time on model download and initialization before "Speak now" means anything. For continuous dictation, the README passes a callback to text() so transcription completes asynchronously while the loop keeps listening.

python
from RealtimeSTT import AudioToTextRecorder


def process_text(text):
    print(text)


if __name__ == "__main__":
    recorder = AudioToTextRecorder()

    while True:
        recorder.text(process_text)

If your audio never touches a microphone, set use_microphone=False and feed 16-bit mono PCM chunks at 16 kHz, or pass the original sample rate so RealtimeSTT can resample. The README's external-audio example opens a file, calls feed_audio with original_sample_rate=16000, prints the transcript and calls recorder.shutdown().

Where RealtimeSTT gets in the way

The Python version gate is explicit and worth reading before anything else. The README states that the current CI matrix covers Python 3.11 and 3.12, and that Python 3.13 and newer are not release targets until dependency and CI gates are available. If your environment is already on 3.13, this is not a library you can adopt today without changing the interpreter.

The engine story is more fragmented than the quick-start suggests. The default path is faster_whisper, but the CPU production recommendation is a sherpa-onnx pairing with two pinned model bundles, installed separately through stt-install-sherpa-models with a --root directory. That means the CPU path is not one pip command; it is an extra plus a model download step, and the README defers the exact pinned model directories to RealtimeSTT_server/PRODUCTION_SERVER.md. Budget time for that if CPU streaming is your target.

The Kroko integration has a similar shape and an additional caveat. Installing it needs the kroko-builder and silero-onnx-cpu extras, then stt-install-kroko --build. The README points at public Community models for local testing and notes commercial model options for production licensing and higher-end models. If you need a licence for the model itself, that is a separate conversation from the MIT licence on the library.

Finally, the multiprocessing design leaks into your code. The if __name__ == "__main__": guard is not stylistic advice; without it, scripts can misbehave, particularly on Windows. Applications that embed RealtimeSTT inside an existing process model, a web framework worker, or a frozen executable should test that interaction early rather than at the end.

RealtimeSTT compared with calling faster-whisper directly

The honest alternative is not another library, it is the model backend on its own. faster-whisper is already the default engine here, and the requirements file pins faster-whisper==1.2.1. If your input is a finished audio file and your output is one transcript, calling faster-whisper directly gives you the same transcription with none of the recorder, VAD, callback and process machinery. You would write the segmenting yourself, but you would also own the failure modes.

The difference shows up the moment you need streaming. faster-whisper transcribes a buffer you hand it; it does not decide when a speaker started or stopped, does not emit interim text while a turn is open, and has no wake word layer. RealtimeSTT exists precisely to supply those four things, and the README's CPU profile is the clearest example of the trade-off it makes: rather than retranscribing a growing buffer with one model, it runs a small streaming model over new frames and a larger model once at finalization. That is a design you would have to invent yourself on top of raw faster-whisper.

The other comparison worth making is between RealtimeSTT as a library and RealtimeSTT as a server. The repository ships both. If your consumer is a browser or another machine, the packaged FastAPI server with its versioned HTTP and WebSocket contracts, session isolation and authentication endpoints may fit better than embedding the recorder in your own process. If your consumer is the same Python process that captures the audio, the library is the simpler surface.

Licence, maintenance and upgrade cost

The library is MIT licensed, which is permissive and places few obligations on how you redistribute it. That licence covers the RealtimeSTT code. It does not automatically cover every model you point it at, and the README makes the distinction visible in the Kroko section, where Community models are described for local testing and commercial model options for production licensing. Model licences are a separate review from the package licence, and nothing here should be read as legal advice.

On maintenance, the last push to the repository was on 2026-09-17, and the most recent release listed is v1.1.2 from 2026-08-30, following v1.1.1 and v1.1.0 within the same month. Three releases in four days suggests an active stretch rather than a settled API, so pinning a version is reasonable if you are building on top of the constructor parameters documented in docs/configuration.md.

The upgrade cost is concentrated in two places. First, the engine extras: the setup.py install guide lists many backends, including faster-whisper, whisper-cpp, transcribe-cpp, openai-whisper, sherpa-onnx, parakeet and omnilingual, and switching between them changes your dependency tree, your model downloads and sometimes your platform. Second, the pinned model bundles used by the CPU streaming profile. Those are installed outside pip through stt-install-sherpa-models, so they will not move when you bump the package version. Treat the package version and the model directory as two separate things to track.

Editorial conclusion

RealtimeSTT fits Python projects that need a recorder loop with VAD, wake words and optional realtime text without assembling those parts yourself. It is the wrong choice if you only need one offline transcription of a finished file, since faster-whisper alone does that with fewer moving parts, and it is also wrong if you need Python 3.13, which the repository states is not a release target. Before adopting it, run the microphone example on your target machine and confirm your chosen engine extra installs cleanly, then check docs/configuration.md for the parameters your application depends on.

Frequently asked questions

How do I install RealtimeSTT?

The README's default install is pip install "RealtimeSTT[faster-whisper]". On Linux you install python3-dev and portaudio19-dev first, and on macOS you install portaudio with Homebrew.

How do I use RealtimeSTT to transcribe from a microphone?

Construct AudioToTextRecorder inside an if __name__ == "__main__": guard and call recorder.text(), which waits for speech, stops after the detected utterance and returns the final transcript. For continuous dictation, pass a callback to text() so transcription completes while the loop keeps listening.

How does RealtimeSTT compare with Whisper?

RealtimeSTT uses Whisper-family engines rather than replacing them: the default general-purpose path is faster_whisper, and other engines are available through install extras. What RealtimeSTT adds around the model is voice activity detection, wake word activation, realtime interim text and a recorder loop.

Official sources

  1. Issues
  2. KoljaB/RealtimeSTT on GitHub
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/koljab-realtimestt.svg)](https://hysenlabs.com/projects/koljab-realtimestt)