# isair/jarvis: an offline voice assistant that runs on your own machine

> Jarvis is a local, voice-first AI assistant for macOS, Windows and Linux that answers to its name mid-sentence, keeps memory on disk, and loads tools over MCP. The macOS path is the one the author develops on; the other platforms are second in line.

**isair/jarvis** — A 100% private AI voice assistant that lives on your computer (works offline). Talk naturally as if Jarvis is a third person in the room, and get conversational responses. It remembers everything, knows location and time, can check the web, control Chrome, track nutrition, and more with support for unlimited MCPs / tools without context rot.

- Repository: https://github.com/isair/jarvis
- Stars: 1,878 · Forks: 370
- Language: Python
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/isair-jarvis

## The problem Jarvis targets: a voice assistant that does not phone home

Cloud assistants are cheap to run because the heavy models live on someone else's hardware. The trade is that your audio, your questions and your follow-ups leave the machine. Jarvis takes the opposite position. The README states that processing is 100% local, that there are no subscriptions, and that sensitive information is redacted before anything is written to disk. That last clause matters more than the privacy slogan: an assistant that keeps unlimited memory is also an assistant that accumulates a transcript of your life, and redaction is the only stated control over what lands in that store.

The audience is narrow and specific. You need a desktop machine with enough compute to run a local model, and you need to be comfortable installing Ollama or pointing Jarvis at an OpenAI-compatible server you already operate. In exchange you get a wake word that works anywhere in a sentence, a rolling context of the current discussion, and a memory that persists across sessions. The README frames the use case as a third person in the room: you can ask Jarvis what it thinks in the middle of a conversation with other people, and it answers from what was just said rather than from a fresh prompt.

## Wake word, rolling context and echo rejection in the pipeline

The README's transcript excerpts show the loop. Audio is transcribed, then an intent step labelled "Intent (wake word)" decides whether the utterance was directed at Jarvis, and rewrites it into a clean instruction. In the example, the heard phrase "What do you think Jarvis?" becomes the intent "what do you think about the weather for the picnic", which means the surrounding conversation is being held somewhere and folded into the request. Only after that does a tool stage run, marked in the log as "Tool: getWeather", followed by response generation.

Echo rejection is a separate stage with its own log line. After Jarvis speaks, it listens for a short follow-up window (the transcript shows 3s), transcribes what it hears, and compares it against its own output. When the two match, the line reads "Heard (echo)" and the assistant drops back to wake word mode. This is the part that decides whether a voice assistant is usable in a room with speakers, and the README is honest that it is imperfect: issue #24 notes that stop commands during speech are sometimes filtered as echo, which is the same mechanism failing in the direction that hurts.

Tool selection is described as smart enough that adding tools does not slow things down. The README claims support for unlimited MCPs and tools "without context rot", and requirements.txt pins mcp==1.13.1, so the integration is the standard Model Context Protocol rather than a bespoke plugin format. The claim about scale is a design assertion, not something the README backs with numbers. EVALS.md is where the project says it tracks accuracy, and that file is the right place to look before trusting the tool-selection behaviour on your own set of servers.

## Installing Jarvis on Windows, macOS or Linux

The README's quick install has two steps: install a model server, then download the build for your platform. Ollama is the default, but any OpenAI-compatible server works, and the README lists LM Studio, Jan, llama.cpp, vLLM, oMLX and LocalAI as examples. Get Ollama from the download page before anything else, because Jarvis expects a server to talk to.

```bash
# Ollama is the default provider; install it first
# https://ollama.com/download
```

The builds come from GitHub Releases. On Linux the README gives a tar command and then runs the binary from the extracted directory.

```bash
# Linux: extract and run
tar -xzf Jarvis-Linux-x64.tar.gz
./Jarvis/Jarvis
```

On Windows you extract Jarvis-Windows-x64.zip and run Jarvis.exe. On macOS you extract Jarvis-macOS-arm64.zip, move the app to Applications, then right-click and choose Open, which is the usual way to launch an unsigned build. According to the README, Jarvis starts listening automatically after launch, and a setup wizard walks through an initial check, model selection, Whisper configuration and dictation. The first real use is the one the README demonstrates: say "Jarvis" anywhere in a sentence and keep talking. Watch the log lines for "Heard", "Intent (wake word)" and "Tool" to see which stage your request reached.

If you prefer to run from source rather than a release build, requirements.txt is the dependency list, and it is heavier than it looks. It pins faster-whisper for speech recognition, Piper and Chatterbox for speech output, PyQt6 and PyQt6-WebEngine for the desktop shell, Playwright for browser control, faiss-cpu for vector search, and geoip2 for location. On Apple Silicon it pulls mlx-whisper instead of relying on CPU inference, and on Windows it pulls CUDA libraries for GPU-accelerated recognition.

## Where Jarvis falls short: voice only, macOS first, dictation broken on Tahoe

The README keeps a known-limitations list, and it is worth reading as a list of blockers rather than caveats. There is no text chat interface yet, tracked as issue #35, so every interaction goes through speech. There are no mobile apps, issue #17. Stop commands during speech are sometimes eaten by echo filtering, issue #24. And dictation, the offline replacement for cloud transcription tools, is unavailable on macOS 26 and later because of a pynput incompatibility, issue #172.

The platform split is the larger constraint. The README states that primary development happens on macOS and that Windows and Linux support may lag behind. The dependency file reflects that asymmetry: mlx-whisper is gated to darwin and arm64, while the CUDA packages are gated to Windows. A Linux user is running the least-trodden path, and the release archives exist for it, but the README does not promise parity.

Two more things the README does not document. There is no rollback procedure for the memory store, and no description of what the redaction filter catches or misses. If your reason for choosing Jarvis is that sensitive data never persists, that second gap is the one to close before you rely on it, because the claim is broad and the mechanism is unspecified.

## How Jarvis differs from a local Whisper plus a chat window

The obvious alternative is assembling the parts yourself: a local Whisper build for transcription, a model server such as Ollama or llama.cpp, and a chat client in front of it. That stack is more transparent, because you choose each component and can inspect every prompt. It is also stateless by default. You get no wake word, no rolling context of a live conversation, no echo rejection, and no memory that survives a restart unless you build one.

Jarvis is the integrated version of that stack with opinions attached. The wake-word intent step, the echo comparison, the persistent memory with redaction, and the MCP tool layer are the parts you would otherwise write. The cost is that you inherit the author's platform priorities and the known limitations above. If your goal is dictation into other applications, the README positions Jarvis against subscription transcription services on the grounds of being free, offline and private, with the caveat that this mode does not work on macOS 26 or later. If your goal is a scriptable assistant you can wire into your own pipelines, a plain Ollama server plus your own client gives you more control and no wake-word layer to fight.

## Licence, release cadence and the cost of keeping up

The repository's licence is reported as NOASSERTION, which means GitHub could not match the LICENSE file to a known template. Read that file directly before you plan to redistribute a build or ship it inside a product; nothing in the README clarifies the terms, and this is not a question to settle by inference.

Upgrade cost is shaped by the dependency pins. requirements.txt fixes faster-whisper at 1.0.3, mcp at 1.13.1, Pillow at 10.4.0, numpy below 2.0.0 and setuptools below 81, among others. Those pins mean an upgrade is not a matter of pulling the newest release and hoping; a moving dependency can break the speech or tool layers, and the pins are how the project avoids that. The release history shows v2.2.0, v2.2.1 and v2.2.2 within about ten days in late July and early August 2026, and the last push to the repository was on 2026-08-25. That is a project that ships often, which cuts both ways: fixes arrive quickly, and so does churn.

If you install from a release archive, keep the previous archive. The README documents no downgrade path, and the memory store is the part you would least want to lose to a bad upgrade.

## Conclusion

Adopt Jarvis if you want a voice assistant whose audio, memory and tool calls stay on your machine, and you accept that the author's primary platform is macOS. Do not adopt it if you need a text chat window, a mobile client, or dictation on macOS 26 or later, since the README flags a pynput incompatibility there. Before installing, check that Ollama or another OpenAI-compatible server is already running, and read EVALS.md rather than the description to judge accuracy.

## FAQ

### What does Jarvis stand for?

The README does not expand the name into an acronym or give an origin for it. The project is named Jarvis and the wake word is "Jarvis", which you can say anywhere in a sentence.

### Is Jarvis a true AI?

Jarvis is a local assistant that runs speech recognition, a language model and tool calls on your own computer, and the README states that processing is 100% local. The model it uses is whichever OpenAI-compatible server you point it at, with Ollama as the default.

### how to use jarvis ai

Install Ollama or another OpenAI-compatible server, download the archive for your platform from GitHub Releases, and launch the app. Jarvis starts listening automatically, so you say "Jarvis" anywhere in a sentence and keep talking; the README's transcript examples show the heard phrase, the wake-word intent and any tool call in the log.

### how to install jarvis on windows 11

Download Jarvis-Windows-x64.zip from GitHub Releases, extract it, and run Jarvis.exe. Windows builds also pull CUDA libraries for GPU-accelerated speech recognition, and the README notes that primary development happens on macOS, so Windows support may lag behind.

## Sources

- [isair/jarvis on GitHub](https://github.com/isair/jarvis)
- [Issues](https://github.com/isair/jarvis/issues)
- [README](https://github.com/isair/jarvis/blob/main/README.md)
- [Releases](https://github.com/isair/jarvis/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/isair-jarvis
