PocketSphinx: a 1970s-vintage recognizer that still earns its keep
A small speech recognizer
At a glance
- What is it?
- PocketSphinx is Carnegie Mellon University's open source, large vocabulary, speaker-independent continuous speech recognizer written in C, now at 5.1.1 with official Python bindings. Its algorithms date back to the 1970s in some cases, kept alive by compactness and efficiency, a CMake build with no SphinxBase dependency, and a command line that emits line-delimited JSON.
- Who is it for?
- Use PocketSphinx when recognition has to run in a small footprint without a cloud service or a GPU era model, and when force alignment of audio to a known transcript is the actual task, since align mode with phone and state levels is a first-class feature. Skip it when raw transcription accuracy on difficult audio is the priority, the README itself warns the results may not be wonderful on a plain WAV.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 16 days ago.
- What is it written in?
- Mainly C, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Nineteen-seventies algorithms, kept for compactness
PocketSphinx is one of Carnegie Mellon University's open source large vocabulary, speaker-independent continuous speech recognition engines, and it is unusually candid about its age, acknowledging that the algorithms and models it implements are now quite old, dating back to the 1970s in some cases, while still being useful in many applications due to their compactness and efficiency. The current version is 5.1.1, and the version number itself carries a story, it is strangely large because there was a release people are using called 5prealpha, and the project switched to proper semantic versioning from there. Recent releases came in a burst during 2026, v5.1.1 on 2026-06-06, v5.1.0 on 2026-05-06, and the v5.1.0rc2 candidate on 2026-05-04, with the last push on 2026-09-14. The primary language is C, and a Python package wraps it as the official bindings.
SphinxBase is gone, and the audio library with it
Two historical burdens have been shed. There is no longer any dependency on SphinxBase, stated with a joke, there is no SphinxBase anymore, this is not the SphinxBase you are looking for, all your SphinxBase are belong to us. And the audio library, which never really built or worked correctly on any platform at all, has simply been removed rather than fixed. What remains is a CMake build that should give reasonable results across Linux and Windows, with the author admitting uncertainty about Mac OS X because he does not have one of those machines. This candor sets expectations correctly, the recognizer core is portable, the build is conventional, and anything involving audio capture is now explicitly your problem, which is why the examples directory offers separate live_portaudio.c, live_pulseaudio.c and live_win32.c entry points for each audio backend.
Two install paths: pip for Python, CMake for C
The Python module installs from the top level directory into a virtual environment:
python3 -m venv ~/ve_pocketsphinx
. ~/ve_pocketsphinx/bin/activate
pip install .The C library and bindings build with three commands, cmake -S . -B build, cmake --build build, and cmake --build build --target install, with -DCMAKE_INSTALL_PREFIX available in the first command when write access to /usr/local is missing. Example code carries optional dependencies that are not strictly necessary to build and install, and on Debian GNU/Linux and derivatives including Raspberry Pi OS and Ubuntu they come from one package command:
sudo apt install \
ffmpeg \
libasound2-dev \
libportaudio2 \
libportaudiocpp0 \
libpulse-dev \
libsox-fmt-all \
portaudio19-dev \
soxThat list is effectively a map of the audio ecosystems the examples touch, ALSA, PulseAudio, PortAudio, and sox for format conversion.
A command surface built around live, single and align
The pocketsphinx command-line program reads single-channel 16-bit PCM audio from standard input or from files, and attempts to recognize speech using the default acoustic and language model. It accepts a large number of options which you probably do not care about, in the README's own phrasing, a command which defaults to live, and one or more inputs except in align mode, or a dash to read from standard input. Five commands exist. help prints that long option list, and help with a command name such as pocketsphinx help align shows alignment-specific options. config dumps the configuration as JSON to standard output, loadable again through the -config option. live detects speech segments in each input and recognizes them. single recognizes each input as one utterance. align force-aligns audio to a known text. The quick start is one line, pocketsphinx single speech.wav, applied to a single-channel WAV file, with the caveat that the results may not be wonderful.
Line-delimited JSON with aggressively short field names
Recognition output is line-delimited JSON, a format the README defends as not the prettiest but surely better than XML. Each line is one JSON object whose fields have short names to keep the lines readable. b is the start time in seconds from the beginning of the stream. d is the duration in seconds. p is the estimated probability of the recognition result, a number between 0 and 1 representing the likelihood of the input according to the model. t is the full text of the result, and w is a list of segments, usually words, each of which carries its own b, d, p and t fields for start, end, probability and text. When -phone_align yes has been passed, a w field appears inside each word containing phone segmentations in the same format, so word and phone timing flow through one schema. Logging is quiet by default, only errors reach standard error, with -loglevel INFO available for more, and partial results are not printed, with a note not to hold your breath waiting for them.
Force alignment, and cleaning it up with jq
The align command force-aligns a single audio file to a word sequence. The first positional argument is the input, and all subsequent arguments are concatenated to make the text, a deliberate choice to avoid surprises when quoting is forgotten. The user is responsible for normalizing that text to remove punctuation, uppercase, centipedes, and so on. The basic form:
pocketsphinx align goforward.wav "go forward ten meters"Word-level alignment is the default, -phone_align yes adds phone alignments, and -state_align yes adds state-level alignments while automatically enabling phone alignment too. The output is described as not particularly readable, and jq is the suggested cleanup tool:
pocketsphinx align audio.wav $text | jq '.w[]|[.t,.b]'extracts word names and start times, while
pocketsphinx -phone_align yes align audio.wav $text | jq '.w[]|.w[]|[.t,.d]'extracts phone names and durations.
soxflags negotiates the input format for you
The cleverest small command is soxflags, which returns arguments to sox that create the appropriate input format for the recognizer. Because the sox command line is slightly quirky, those arguments must always come after the filename, or after -d when sox should read from the microphone. Live recognition from the mic becomes a two-stage pipe:
sox -d $(pocketsphinx soxflags) | pocketsphinx -and decoding from a compressed file works the same way:
sox audio.mp3 $(pocketsphinx soxflags) | pocketsphinx -The recognizer itself only consumes single-channel 16-bit PCM, so every other format routes through sox, and rather than documenting the exact invocation, the program emits it. It is format negotiation executed as command substitution, and it removes an entire class of wrong-flag bugs at the boundary between arbitrary audio and the recognizer's one accepted format.
Bindings, a GStreamer plugin and an Alpine image
The Python packaging tells the project's current story. pyproject.toml names the package pocketsphinx, the official Python bindings, with David Huggins-Daines as author, a single runtime dependency on sounddevice, and classifiers declaring Python 3.9 through 3.13, OS independent operation, BSD licensing, and development status Mature. It installs two extra console scripts, pocketsphinx_lm and pocketsphinx_to_textgrid. cibuildwheel configuration builds PyPy wheels plus cp311 and cp313, universal2 wheels for macOS where possible. Beyond Python, the repository carries a gst directory for GStreamer integration and a Dockerfile built on Alpine that installs sox, portaudio and ALSA utilities, compiles with CMake and Ninja, and drops the examples into /work. Documentation for the Python API lives at pocketsphinx.readthedocs.io and for the C API at the CMU Sphinx site, with Doxygen docs buildable via cmake --build build --target docs.
Editorial conclusion
Use PocketSphinx when recognition has to run in a small footprint without a cloud service or a GPU era model, and when force alignment of audio to a known transcript is the actual task, since align mode with phone and state levels is a first-class feature. Skip it when raw transcription accuracy on difficult audio is the priority, the README itself warns the results may not be wonderful on a plain WAV. Verify before adopting that your platform builds cleanly under CMake, Linux and Windows are the confident targets while macOS is untested by the author, and note that the audio library older versions carried has been removed entirely.
Frequently asked questions
How do I install PocketSphinx using pip?
From the top level directory of the repository, create a virtual environment, activate it, and run pip install . The shown sequence is python3 -m venv ~/ve_pocketsphinx, then activating it, then pip install .
How do you use pocketsphinx?
Run the pocketsphinx command with a mode such as live, single or align and one or more input files or a dash for standard input. It reads single-channel 16-bit PCM, so other formats are converted with sox using the arguments printed by pocketsphinx soxflags, and pocketsphinx help lists the full option set.
Is there a Python package that can convert speech-to-text?
Yes, the pocketsphinx package is the official Python binding for PocketSphinx, Carnegie Mellon's open source continuous speech recognition engine. It depends only on sounddevice and declares support for Python 3.9 through 3.13, with examples in C and Python in the repository's examples directory.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/cmusphinx-pocketsphinx)