VoiceMode is an MCP server plus a separate installer package, and it has been renamed twice
Natural voice conversations with Claude Code
At a glance
- What is it?
- Voice conversations for Claude Code and other MCP-capable agents, with local Whisper and Kokoro services as an offline path. The repository, the PyPI package, the CLI commands and the build file all carry different names for the same project.
- Who is it for?
- VoiceMode is a reasonable fit if your hands are busy and you already run an MCP-capable agent, and the local Whisper and Kokoro path removes the cloud dependency for anyone who cannot send audio to a hosted service. Two things to check first.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
One project carries five different names
The repository is voicemode, the package on the index is voice-mode, and the project description says it was formerly called voice-mcp. The command line surface continues the pattern with three more spellings: a voicemode command for configuration, a voice-mode-install command for the installer, and voicemode-mcp-launcher as the server entry point. The build file has not caught up, since its header comment still reads Voice MCP Makefile.
That is not cosmetic when you are automating an install. The plugin route and the Python route ask for different identifiers, and the server is registered under a short name that has to be chosen by whoever writes the configuration. Anyone writing documentation or a setup script for their own team will hit the mismatch between the repository name and the package name first.
The domain matches the current name rather than the old one, voicemode.dev, and the project's own identifier appears at the bottom of the readme as a namespaced name for the MCP server. The default branch is master, and the classifier list describes the package as beta while the readme says it works out of the box with no configuration required.
The dependency list keeps its change history in comments
The install requirements are ordinary until you read the annotations. The MCP framework is pinned to a 3.x range with a comment saying it was migrated in a tracked ticket and that the change fixes a vulnerability scored at the highest level of the CVSS scale, with the advisory identifier attached. The OpenAI client is required at version 2 or newer, with a comment recording that it used to be a 1.x floor and naming two tickets.
Two more comments explain non-obvious pins. The voice activity detection dependency is a maintained fork rather than the original package, chosen because it exposes import metadata and does not need the older packaging module. And the asynchronous HTTP client is there for the Whisper and Kokoro service installers, which is why a library that otherwise talks to audio hardware needs a web client at all.
There is also a conditional dependency pinned to Python 3.13 and above, which is the audio processing module that left the standard library in that release. That single line explains the 3.10 to 3.14 support range in the classifiers: the package runs on three versions before the module was removed and brings its own copy for the versions after it. Credentials go through the system keyring rather than a file, and process information is read through a system information library for the service supervision side.
Installing means installing a second package that installs the first
There are two documented routes and they use different integration models. The recommended one for Claude Code is a plugin from a marketplace:
# Add the VoiceMode marketplace
claude plugin marketplace add mbailey/voicemode
# Install VoiceMode plugin
claude plugin install voicemode@voicemode
## Install dependencies (CLI, Local Voice Services)
/voicemode:install
# Start talking!
/voicemode:converseNote that installation continues inside the agent, as a slash command that installs dependencies and local voice services. The second route goes through a package manager and registers an MCP server explicitly:
# Run the installer (sets up dependencies and local voice services)
uvx voice-mode-install
# Add to Claude Code
claude mcp add --scope user voicemode -- uvx --refresh --from voice-mode voicemode-mcp-launcherTwo details in that line matter. The scope is the user, so the registration applies to your account rather than to one project. And the server is launched through the package runner with a refresh flag, which tells it to re-resolve the package rather than reuse a cached one, so a start can pick up a newer published version than the one you tested. The installer package, the repository's installer directory, its shell script and the build targets for building and publishing an installer are all separate from the server package.
The MCP surface is two tools, and the allow list names both
The permissions section is short enough to be the clearest statement of the tool's surface. To use the server without prompts, two entries go into the allow list in the settings file:
{
"permissions": {
"allow": [
"mcp__voicemode__converse",
"mcp__voicemode__service"
]
}
}Those two names are the whole interface: one for conversation and one for service management, which is what the local Whisper and Kokoro installers attach to. That naming convention is worth recognising on sight, because every tool from this server appears with that prefix, and an allow list that includes both of these entries will stop prompting for the features the project considers core.
It is also a fairly wide grant for something that records audio. Converse is the tool that captures the microphone and plays responses, so allowing it permanently means an agent can start a conversation without asking, and the configuration guide is where the remaining options live. The troubleshooting table treats a missing microphone as a permissions problem at three different levels: the terminal or application, the operating system, and, for WSL2, the pulseaudio packages that the page calls out separately.
Offline means two local servers that imitate the OpenAI API
The privacy and offline story is implemented by substitution rather than by a separate code path. Two local services can be installed, one for speech to text based on Whisper and one for text to speech with multiple voices based on Kokoro. Both expose the same API shape as the hosted service, so the client switches between them without the agent noticing.
That design has a practical consequence. The offline path is not one flag, it is two local servers to install, configure and keep running, and the page links a separate setup document for each. The requirement list says only a computer with microphone and speakers, and the feature list calls local services optional, so a first run with no key set gets silence or an error rather than speech.
There is one setting here that deserves a decision rather than a glance. Saving audio for debugging is off by default and writes files under a dated directory:
export VOICEMODE_SAVE_AUDIO=true
# Files saved to ~/.voicemode/audio/YYYY/MM/Turning that on during a support session and leaving it on afterwards leaves microphone recordings in your home directory, organised by month. Nothing in the file mentions a cleanup command or a retention policy, so it is worth remembering that the directory exists.
Audio needs a per-platform package list, and macOS is the only one asking for node
Before any of that runs, the system needs audio libraries, and the four platform lists differ more than you would expect. Debian and Ubuntu get the longest list, including the ALSA development headers, the PortAudio development package, pulseaudio and its utilities, and the Python headers, plus ffmpeg and a compiler. Fedora and RHEL get the equivalents through dnf with the ALSA and PortAudio development packages. WSL2 users are told explicitly that the pulseaudio packages from the Debian list are what microphone access depends on.
macOS gets a single line, and it includes node:
brew install ffmpeg node portaudioNode appears on the macOS list and on no other, which suggests the plugin route needs a JavaScript runtime that the Python route does not. NixOS is handled by the flake rather than by a package list, either a development shell or a profile install from the repository's GitHub address, with a separate snippet for adding the package to a system-wide configuration.
Supported platforms are Linux, macOS, Windows natively or under WSL, and NixOS, with Python 3.10 through 3.14. From source, the route is a clone and an editable tool install, which is the shortest command sequence in the whole document.
The repository ships agent definitions, a server manifest and per-distro installer tests
The tree is broader than a Python package. There is a plugin directory for the Claude Code integration, a separate directory for agent configuration, a server manifest, an MCP configuration file, an agent instructions file, a documentation site built with mkdocs, an llms.txt file, a Makefile and a lock file for the package manager. The server namespace is declared in the readme itself as a namespaced name for the MCP server.
The Makefile is where the project's process shows. Its target list has four parallel families for the server, the package build, the installer and the release, each with its own test, publish and release variants, so publishing an installer is a separate action from publishing the package. The installer tests are per distribution, covering Ubuntu, Fedora and PyPI-hosted installs, and there is a virtual machine target set for the same three platforms with quick variants and a clean target. Dependency auditing has its own targets, including a critical-only variant and a JSON output variant.
The help output describes itself with the old project name, and it names the server package as voice-mode, which is the same mismatch as the rest of the repository. For a project whose pitch is voice, that is a small thing, but it is the sort of small thing that costs an afternoon when a script disagrees with a readme.
Editorial conclusion
VoiceMode is a reasonable fit if your hands are busy and you already run an MCP-capable agent, and the local Whisper and Kokoro path removes the cloud dependency for anyone who cannot send audio to a hosted service. Two things to check first. The privacy story depends on which services you point it at, because the local option is a pair of servers speaking the OpenAI API shape rather than a special offline mode, and one environment variable persists your microphone audio to disk for debugging. Second, the naming is inconsistent across the repository, the package index, the installer and the build file, so pin the PyPI package rather than the repository name when you script an install, and read the dependency list before you pin versions, since it records a security upgrade and an OpenAI major version jump in comments rather than in a changelog entry you would find by searching.
Frequently asked questions
What is VoiceMode and what does it add to Claude Code?
It adds voice conversation over MCP, so you can speak to the agent and hear replies, with local speech services as an alternative to hosted ones. Integration is either a Claude Code plugin from a marketplace or an MCP server registered with the user scope.
How do I install VoiceMode from the command line?
Install the package manager, then run the installer package, then register the MCP server with your user scope using the launcher entry point. The server is launched through the package runner with a refresh flag, so a start re-resolves the published package.
Can VoiceMode work without sending audio to a cloud service?
Yes, by installing two local services, one for speech to text based on Whisper and one for text to speech based on Kokoro. Both expose the same API shape as the hosted service, so the client switches between them, but they are separate installations with their own setup guides.
Which platforms and Python versions does VoiceMode support?
Linux, macOS, Windows either natively or under WSL, and NixOS, with Python 3.10 through 3.14. Each platform needs its own audio packages, and WSL2 users are told the pulseaudio packages are required for microphone access.
Which MCP tools does VoiceMode expose?
Two, named converse and service, which appear in the permission allow list as mcp__voicemode__converse and mcp__voicemode__service. Adding both to the settings file removes the prompts for the features the project treats as core.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mbailey-voicemode)