# izwi: local voice AI behind an OpenAI-shaped API

> A Rust runtime that runs speech recognition, diarization, alignment, text to speech, voice cloning and chat models on your own machine, then exposes them as OpenAI-compatible routes so existing clients work unchanged. The install is three artifacts, the model catalogue is broad, and the weights carry their own licences.

**izwi-ai/izwi** — Voice AI runtime. Local first transcription, speaker diarization, TTS, and voice cloning with an OpenAI compatible API.

- Repository: https://github.com/izwi-ai/izwi
- Website: https://docs.izwiai.com/
- Stars: 383 · Forks: 40
- Language: Rust
- License: MIT
- Published: 2026-09-18 · Updated: 2026-09-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/izwi-ai-izwi

## One runtime behind four surfaces

Izwi is described in four words that matter: a desktop app, a web interface, a CLI, and a local inference server.

They are four views of the same runtime rather than four products, and that is the design decision. The same local models are reachable from a person clicking in the app, from a browser, from a terminal, and from code over HTTP.

The API is the part that carries the argument. Alongside its own product routes, it exposes OpenAI-compatible /v1 endpoints for models, chat completions, audio speech, audio transcriptions, and preview support for Responses. A client written against a hosted provider can therefore be repointed at localhost without a rewrite, which is what makes it usable inside an existing pipeline rather than only as an app.

The capability list is broader than the name suggests. Beyond real-time voice conversation with local recognition, chat and synthesis models, it includes long-form Studio projects, voice cloning, voice design and saved voices, transcription with speaker diarization, forced alignment, realtime speech to text, and model download, load, unload and delete alongside history, exports and settings.

The privacy claim is specific rather than absolute: inference data stays local, and anonymous desktop analytics are disabled unless a user opts in, and they do not send prompts, transcripts, audio payloads, local paths or personal identifiers.

## The catalogue is wide, and the weights are not yours

The enabled catalogue is listed by family, and it covers four capability groups.

Text to speech includes Qwen3-TTS, Kokoro-82M, Voxtral TTS and VibeVoice. Recognition includes Parakeet, Whisper, Qwen3-ASR, Nemotron 3.5 ASR, VibeVoice ASR, LFM2.5 Audio and Voxtral Mini. Diarization and alignment are covered by Sortformer diarization and the Qwen3 ForcedAligner. Chat models are Qwen3, Qwen3.5, LFM2.5 and Gemma.

Diarization and forced alignment appearing as first-class families is the interesting part. Transcribing a meeting is easy; knowing who spoke and aligning text to audio is the work people end up doing themselves, and having both in the same runtime with the same API shape is what makes it usable for anything longer than a voice note.

Then the licensing paragraph, which is one sentence and is the thing to read carefully. Some model weights and bundled assets have their own licences or access terms, and the models guide is the place to check before redistribution or commercial use of a downloaded artefact.

That sentence covers a real gap: the runtime being MIT says nothing about the right to redistribute Kokoro's weights or a particular checkpoint, and anyone shipping a container image full of models needs to check each one.

## Three artifacts, and the runtime you get depends on the OS version

Installation is a download from the releases page, and there are three artifacts. On macOS you install the .dmg, drag Izwi.app into Applications and launch it. On Linux you install the .deb with a package command. On Windows you run the .exe installer.

The Linux command is the only one written out:

```bash
sudo dpkg -i izwi_*.deb
```

A glob rather than a version, which is convenient and also means you have to be in a directory holding exactly one such file.

Then the runtime support depends on which artifact you downloaded, and this is the paragraph people miss. A macOS Apple Silicon release build uses Metal on macOS 15 and above and falls back to CPU on macOS 12 to 14. Linux and Windows release builds are CPU only. CUDA is supported through the Docker CUDA profile, or through source builds on compatible NVIDIA hosts.

So the published Windows and Linux binaries have no GPU path at all, and the oldest macOS you can accelerate is two versions behind the newest. If your requirement is acceleration on a specific machine, that requirement is checked against this table rather than discovered when a model runs slowly.

## The CLI is serve, pull, then run

The quick start is three verbs, and the order matters because the models are not bundled.

Serving the web interface and the API:

```bash
izwi serve --mode web
```

Two other modes exist, a server-only mode and a desktop mode, so the same command covers all three.

Pulling a synthesis model and generating speech:

```bash
izwi pull Qwen3-TTS-12Hz-0.6B-Base
izwi tts "Hello from Izwi." --output hello.wav
```

And transcribing:

```bash
izwi pull Parakeet-TDT-0.6B-v3
izwi transcribe audio.wav --model Parakeet-TDT-0.6B-v3
```

The pull step appearing before every task is the point. Nothing is downloaded until you ask for it, which keeps the install small and lets you choose one model family instead of a bundle.

Once the server is up, three addresses matter. The application is at localhost on port 8080, the local API reference is served at the /docs path, and the raw OpenAPI document is at /openapi.json. Serving the specification itself means an existing client generator can produce a typed client without anyone maintaining a schema by hand.

## Eight crates, and the MLX dependency is deliberately switched off

The workspace has eight member crates, and their names describe the layering: an ASR toolkit, an agent crate, a core crate, hooks, voice activity detection, the server, the CLI and the desktop shell.

The dependency list is more informative than the crate names. Async is tokio with the full feature set plus tokio-stream, futures and async-stream. The web layer is axum with websocket and multipart features, tower-http for CORS, static files and tracing, and utoipa with scalar_api_reference for the served documentation.

Model loading is deliberately minimal: hf-hub with rustls, tokio and ureq, and no default features; safetensors for weights; tokenizers at a pinned minor.

Inference is Candle, at 0.11 across core, nn and transformers, with half-precision support. Audio is hound for WAV output, symphonia with default features off for decoding, rubato for resampling and rustfft for the transform.

One line explains the Apple Silicon story. The MLX bindings are present as a comment, with a note that the Rust MLX binding is still experimental and that FFI bindings will be used instead, and the git dependency is commented out. So the Metal path is real in the shipped build but is not going through that crate.

Building manually is scoped per binary:

```bash
cargo build --release -p izwi-cli
cargo build --release -p izwi-server
```

## Prefix caching is opt-in, and multi-token prediction is on

The environment file is annotated, and two features deserve reading.

Committed prefix caching is off by default and requires an explicit namespace salt plus a model family whose cache contract permits shared spans, with a page limit variable alongside. The requirement for a salt is the interesting part: a shared prefix cache without one lets two tenants read each other's cached spans, so the switch is not merely a performance knob.

The second is the opposite default. Qwen3.8 native multi-token prediction is enabled at depth one on every backend, with a draft-token count that starts at one, supported values from one to three, and a disable switch. The instruction attached to raising the value is precise: only after validating acceptance and throughput on the target model and device.

That is the right way to document a default that a user might want to change, because a draft token that is rejected often costs throughput rather than saving it, and the only way to know is to measure on your own hardware.

Two smaller things are visible in the same file. Legacy variable names for concurrency and timeout are kept as aliases with a comment preferring the canonical ones, and a Python unbuffered-output variable survives from an earlier implementation.

## Docker builds the interface and the runtime in separate stages

The container build has three stages and the order is the design.

The first stage uses a Node slim image to build the React interface, installing from the lockfile with scripts disabled and copying the source only after the dependencies are cached.

The second uses a Rust image to build the server binary in release mode, with the lock file honoured, and it builds only the server because the desktop shell is not needed in a container. A CUDA variant uses the NVIDIA CUDA developer image as a third stage, with the compute capability and the feature list passed in as build arguments.

The compose file makes the operational contract visible. The CPU service exposes one port, mounts named volumes for models and data, points the config and interface directories at their in-container paths, and sets both HuggingFace cache variables into the models volume so a downloaded model survives a rebuild. The readiness endpoint is checked with curl every thirty seconds with a sixty-second start period, and logging is capped at ten megabytes across three files.

The CUDA service is a second definition behind a profile, built from the same context with the CUDA target and its own image, activated by passing the profile flag rather than being on by default.

The service is also marked to never pull, which means an image built locally is the one that runs, and a missing local build fails loudly instead of silently pulling something else.

## MIT in the readme, Apache-2.0 in the manifest, beta-18 in the workspace

Three administrative details disagree with each other, and none of them is fatal.

The readme states that Izwi is licensed under the MIT licence, while the workspace manifest declares Apache-2.0. For a Rust project that matters, because the manifest is what cargo and any downstream consumer read.

The workspace version is 0.1.0-beta-18, while the most recent published release is beta-17 from 22 June 2026, preceded by beta-16 in June and beta-15 in May. So the branch is one prerelease ahead of what anyone can download, and all three releases are betas in a five-month window.

The rest of the repository is consistent with a Rust desktop project: a Cargo lock file, a dev Dockerfile, separate development and production compose files, a benchmarks directory, a data directory, a tasks directory, an agent instructions file, and two application icons, one of them named as a redesign.

The oddity is a small TypeScript file at the top level named vercel.ts, in a project whose primary language is Rust and whose runtime is a desktop app and local server. What it configures is not described anywhere in what is readable here.

## Conclusion

izwi is worth considering when you want voice models on your own hardware but do not want to rewrite the client. The OpenAI-compatible surface means a script that already calls a hosted API can point at localhost, and the model catalogue covers transcription, alignment, diarization and synthesis rather than one task. Two things to check before committing. Model weights and bundled assets have their own licences and access terms, so review the models guide before shipping anything commercially. And the runtime you can download is narrower than the one you can build: Linux and Windows artifacts are CPU only, with CUDA arriving through a Docker profile or a source build.

## FAQ

### What does izwi provide locally?

Real-time voice conversation with local recognition, chat and synthesis models, long-form Studio projects, voice cloning and voice design with saved voices, transcription with speaker diarization and forced alignment, and local model management with history and exports.

### Can existing OpenAI-compatible code talk to izwi?

Yes. It exposes OpenAI-compatible /v1 routes for models, chat completions, audio speech, audio transcriptions and preview Responses support, with the served reference at the docs path on port 8080 and the raw OpenAPI document at the openapi.json path.

### Does the izwi release build use a GPU on Windows or Linux?

No. Linux and Windows release builds are CPU only. A macOS Apple Silicon build uses Metal on macOS 15 and above and falls back to CPU on macOS 12 to 14, and CUDA is available through the Docker CUDA profile or a source build.

### Can I use izwi models commercially?

Check first. The project states that some model weights and bundled assets have their own licences or access terms, and points to the models guide before redistribution or commercial use of a downloaded artefact.

## Sources

- [izwi-ai/izwi on GitHub](https://github.com/izwi-ai/izwi)
- [License: MIT](https://github.com/izwi-ai/izwi/blob/main/LICENSE)
- [Project website](https://docs.izwiai.com/)
- [README](https://github.com/izwi-ai/izwi/blob/main/README.md)
- [Releases](https://github.com/izwi-ai/izwi/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/izwi-ai-izwi
