# VoiceMode is an MCP server plus a separate installer package, and it has been renamed twice

> Voice conversations for Claude Code and other MCP-capable agents, with local Whisper and Kokoro services as an offline path. The repository, the PyPI package, the CLI commands and the build file all carry different names for the same project.

**mbailey/voicemode** — Natural voice conversations with Claude Code

- Repository: https://github.com/mbailey/voicemode
- Website: https://voicemode.dev
- Stars: 1,387 · Forks: 196
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/mbailey-voicemode

## One project carries five different names

The repository is voicemode, the package on the index is voice-mode, and the project description says it was formerly called voice-mcp. The command line surface continues the pattern with three more spellings: a voicemode command for configuration, a voice-mode-install command for the installer, and voicemode-mcp-launcher as the server entry point. The build file has not caught up, since its header comment still reads Voice MCP Makefile.

That is not cosmetic when you are automating an install. The plugin route and the Python route ask for different identifiers, and the server is registered under a short name that has to be chosen by whoever writes the configuration. Anyone writing documentation or a setup script for their own team will hit the mismatch between the repository name and the package name first.

The domain matches the current name rather than the old one, voicemode.dev, and the project's own identifier appears at the bottom of the readme as a namespaced name for the MCP server. The default branch is master, and the classifier list describes the package as beta while the readme says it works out of the box with no configuration required.

## The dependency list keeps its change history in comments

The install requirements are ordinary until you read the annotations. The MCP framework is pinned to a 3.x range with a comment saying it was migrated in a tracked ticket and that the change fixes a vulnerability scored at the highest level of the CVSS scale, with the advisory identifier attached. The OpenAI client is required at version 2 or newer, with a comment recording that it used to be a 1.x floor and naming two tickets.

Two more comments explain non-obvious pins. The voice activity detection dependency is a maintained fork rather than the original package, chosen because it exposes import metadata and does not need the older packaging module. And the asynchronous HTTP client is there for the Whisper and Kokoro service installers, which is why a library that otherwise talks to audio hardware needs a web client at all.

There is also a conditional dependency pinned to Python 3.13 and above, which is the audio processing module that left the standard library in that release. That single line explains the 3.10 to 3.14 support range in the classifiers: the package runs on three versions before the module was removed and brings its own copy for the versions after it. Credentials go through the system keyring rather than a file, and process information is read through a system information library for the service supervision side.

## Installing means installing a second package that installs the first

There are two documented routes and they use different integration models. The recommended one for Claude Code is a plugin from a marketplace:

```bash
# Add the VoiceMode marketplace
claude plugin marketplace add mbailey/voicemode

# Install VoiceMode plugin
claude plugin install voicemode@voicemode

## Install dependencies (CLI, Local Voice Services)

/voicemode:install

# Start talking!
/voicemode:converse
```

Note that installation continues inside the agent, as a slash command that installs dependencies and local voice services. The second route goes through a package manager and registers an MCP server explicitly:

```bash
# Run the installer (sets up dependencies and local voice services)
uvx voice-mode-install

# Add to Claude Code
claude mcp add --scope user voicemode -- uvx --refresh --from voice-mode voicemode-mcp-launcher
```

Two details in that line matter. The scope is the user, so the registration applies to your account rather than to one project. And the server is launched through the package runner with a refresh flag, which tells it to re-resolve the package rather than reuse a cached one, so a start can pick up a newer published version than the one you tested. The installer package, the repository's installer directory, its shell script and the build targets for building and publishing an installer are all separate from the server package.

## The MCP surface is two tools, and the allow list names both

The permissions section is short enough to be the clearest statement of the tool's surface. To use the server without prompts, two entries go into the allow list in the settings file:

```json
{
  "permissions": {
    "allow": [
      "mcp__voicemode__converse",
      "mcp__voicemode__service"
    ]
  }
}
```

Those two names are the whole interface: one for conversation and one for service management, which is what the local Whisper and Kokoro installers attach to. That naming convention is worth recognising on sight, because every tool from this server appears with that prefix, and an allow list that includes both of these entries will stop prompting for the features the project considers core.

It is also a fairly wide grant for something that records audio. Converse is the tool that captures the microphone and plays responses, so allowing it permanently means an agent can start a conversation without asking, and the configuration guide is where the remaining options live. The troubleshooting table treats a missing microphone as a permissions problem at three different levels: the terminal or application, the operating system, and, for WSL2, the pulseaudio packages that the page calls out separately.

## Offline means two local servers that imitate the OpenAI API

The privacy and offline story is implemented by substitution rather than by a separate code path. Two local services can be installed, one for speech to text based on Whisper and one for text to speech with multiple voices based on Kokoro. Both expose the same API shape as the hosted service, so the client switches between them without the agent noticing.

That design has a practical consequence. The offline path is not one flag, it is two local servers to install, configure and keep running, and the page links a separate setup document for each. The requirement list says only a computer with microphone and speakers, and the feature list calls local services optional, so a first run with no key set gets silence or an error rather than speech.

There is one setting here that deserves a decision rather than a glance. Saving audio for debugging is off by default and writes files under a dated directory:

```bash
export VOICEMODE_SAVE_AUDIO=true
# Files saved to ~/.voicemode/audio/YYYY/MM/
```

Turning that on during a support session and leaving it on afterwards leaves microphone recordings in your home directory, organised by month. Nothing in the file mentions a cleanup command or a retention policy, so it is worth remembering that the directory exists.

## Audio needs a per-platform package list, and macOS is the only one asking for node

Before any of that runs, the system needs audio libraries, and the four platform lists differ more than you would expect. Debian and Ubuntu get the longest list, including the ALSA development headers, the PortAudio development package, pulseaudio and its utilities, and the Python headers, plus ffmpeg and a compiler. Fedora and RHEL get the equivalents through dnf with the ALSA and PortAudio development packages. WSL2 users are told explicitly that the pulseaudio packages from the Debian list are what microphone access depends on.

macOS gets a single line, and it includes node:

```bash
brew install ffmpeg node portaudio
```

Node appears on the macOS list and on no other, which suggests the plugin route needs a JavaScript runtime that the Python route does not. NixOS is handled by the flake rather than by a package list, either a development shell or a profile install from the repository's GitHub address, with a separate snippet for adding the package to a system-wide configuration.

Supported platforms are Linux, macOS, Windows natively or under WSL, and NixOS, with Python 3.10 through 3.14. From source, the route is a clone and an editable tool install, which is the shortest command sequence in the whole document.

## The repository ships agent definitions, a server manifest and per-distro installer tests

The tree is broader than a Python package. There is a plugin directory for the Claude Code integration, a separate directory for agent configuration, a server manifest, an MCP configuration file, an agent instructions file, a documentation site built with mkdocs, an llms.txt file, a Makefile and a lock file for the package manager. The server namespace is declared in the readme itself as a namespaced name for the MCP server.

The Makefile is where the project's process shows. Its target list has four parallel families for the server, the package build, the installer and the release, each with its own test, publish and release variants, so publishing an installer is a separate action from publishing the package. The installer tests are per distribution, covering Ubuntu, Fedora and PyPI-hosted installs, and there is a virtual machine target set for the same three platforms with quick variants and a clean target. Dependency auditing has its own targets, including a critical-only variant and a JSON output variant.

The help output describes itself with the old project name, and it names the server package as voice-mode, which is the same mismatch as the rest of the repository. For a project whose pitch is voice, that is a small thing, but it is the sort of small thing that costs an afternoon when a script disagrees with a readme.

## Conclusion

VoiceMode is a reasonable fit if your hands are busy and you already run an MCP-capable agent, and the local Whisper and Kokoro path removes the cloud dependency for anyone who cannot send audio to a hosted service. Two things to check first. The privacy story depends on which services you point it at, because the local option is a pair of servers speaking the OpenAI API shape rather than a special offline mode, and one environment variable persists your microphone audio to disk for debugging. Second, the naming is inconsistent across the repository, the package index, the installer and the build file, so pin the PyPI package rather than the repository name when you script an install, and read the dependency list before you pin versions, since it records a security upgrade and an OpenAI major version jump in comments rather than in a changelog entry you would find by searching.

## FAQ

### What is VoiceMode and what does it add to Claude Code?

It adds voice conversation over MCP, so you can speak to the agent and hear replies, with local speech services as an alternative to hosted ones. Integration is either a Claude Code plugin from a marketplace or an MCP server registered with the user scope.

### How do I install VoiceMode from the command line?

Install the package manager, then run the installer package, then register the MCP server with your user scope using the launcher entry point. The server is launched through the package runner with a refresh flag, so a start re-resolves the published package.

### Can VoiceMode work without sending audio to a cloud service?

Yes, by installing two local services, one for speech to text based on Whisper and one for text to speech based on Kokoro. Both expose the same API shape as the hosted service, so the client switches between them, but they are separate installations with their own setup guides.

### Which platforms and Python versions does VoiceMode support?

Linux, macOS, Windows either natively or under WSL, and NixOS, with Python 3.10 through 3.14. Each platform needs its own audio packages, and WSL2 users are told the pulseaudio packages are required for microphone access.

### Which MCP tools does VoiceMode expose?

Two, named converse and service, which appear in the permission allow list as mcp__voicemode__converse and mcp__voicemode__service. Adding both to the settings file removes the prompts for the features the project treats as core.

## Sources

- [License: MIT](https://github.com/mbailey/voicemode/blob/master/LICENSE)
- [mbailey/voicemode on GitHub](https://github.com/mbailey/voicemode)
- [Project website](https://voicemode.dev)
- [README](https://github.com/mbailey/voicemode/blob/master/README.md)
- [Releases](https://github.com/mbailey/voicemode/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/mbailey-voicemode
