Moshi from Kyutai Labs: a full-duplex speech-text model with three inference stacks
Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audio codec.
At a glance
- What is it?
- Moshi is a speech-text foundation model and full-duplex spoken dialogue framework built on the Mimi streaming codec. It ships as PyTorch, MLX and Rust implementations, and the choice between them is the first real decision an adopter has to make.
- Who is it for?
- Adopt Moshi if you need a research-grade full-duplex speech loop and can work inside the PyTorch or MLX stack, or if you want the Rust and Candle path for a production service with an NVIDIA GPU. Do not adopt it if you need a documented fine-tuning workflow in this repository, since the README points to kyutai-labs/moshi-finetune for that, or if you expect a text-only model, because Moshi's output is audio with a parallel text stream.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 21 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Moshi solves that a turn-taking speech pipeline does not
Most spoken dialogue systems are built as a pipeline: speech recognition produces text, a language model answers, and speech synthesis speaks the answer. That structure forces turn-taking, because each stage waits for the previous one to finish. Moshi is built around the opposite assumption. The README describes it as a full-duplex framework, meaning the model handles both sides of the conversation at once rather than alternating. The repository models two audio streams, one for Moshi speaking and one for the user, and predicts text tokens for its own speech alongside them, which the README calls the inner monologue.
The audience is narrow but specific. This is for engineers and researchers who want to experiment with simultaneous listening and speaking, not for teams that need a drop-in transcription API. The theoretical latency figure the README gives is 160ms, split into 80ms of Mimi frame size and 80ms of acoustic delay, with a practical overall latency described as low as 200ms on an L4 GPU. Those numbers come from the project's own description and depend on hardware, so treat them as the design target rather than a guarantee for your machine.
The two-transformer architecture and what Mimi contributes
Moshi splits its modelling work across two transformers of very different sizes. A large Temporal Transformer with 7B parameters handles dependencies over time. A small Depth Transformer handles the dependencies between codebooks within a single time step. That division is what makes streaming feasible: the expensive model runs once per frame, and the cheap model fills in the codebook detail.
Mimi is the codec underneath, and its frame rate is the reason the architecture works at all. It processes 24 kHz audio down to a 12.5 Hz representation at 1.1 kbps in a fully streaming manner, with an 80ms latency equal to the frame size. The README compares this against SpeechTokenizer at 50 Hz and 4 kbps and SemantiCodec at 50 Hz and 1.3 kbps, both of which are non-streaming. A lower frame rate means fewer autoregressive steps for the Temporal Transformer, which is where the latency saving comes from. Mimi adds a transformer to both encoder and decoder, uses a distillation loss so the first codebook tokens match a WavLM self-supervised representation, and according to the README relies only on an adversarial training loss with feature matching. The design borrows from SoundStream and EnCodec, so anyone familiar with neural audio codecs will recognise the lineage.
Installing Moshi and running the first inference
The README states a minimum of Python 3.10, with 3.12 recommended, and notes that specific requirements live in the individual backend directories. The PyTorch and MLX clients are both on PyPI, so installation is a single command per backend. The MLX package is described as best with Python 3.12.
pip install -U moshi # moshi PyTorch, from PyPI
pip install -U moshi_mlx # moshi MLX, from PyPI, best with Python 3.12.The README also mentions bleeding-edge versions for both packages, but the text is truncated at that point, so the exact command for the development versions is not available here. After installing, you need a checkpoint. Models are published on Hugging Face in several formats, and the choice depends on the backend. For PyTorch, the bf16 checkpoints are kyutai/moshika-pytorch-bf16 and kyutai/moshiko-pytorch-bf16, with q8 variants marked experimental. For MLX, the same two voices are available in q4, q8 and bf16. For Rust and Candle, the README lists q8 and bf16 variants. Mimi is bundled inside each of those repositories and always uses the same checkpoint format, so you do not fetch it separately.
The repository also ships a docker-compose.yml that builds the PyTorch service and runs it behind a Cloudflare tunnel. The service exposes port 8998 over TCP, mounts a named volume at /root/.cache/huggingface for the model cache, and requests all NVIDIA GPUs through the deploy resources block. The model is selected with the HF_REPO environment variable, which defaults in the file to kyutai/moshiko-pytorch-bf16, with a commented-out line for kyutai/moshika-pytorch-bf16. A second service, tunnel, runs cloudflare/cloudflared:latest with TUNNEL_URL set to http://moshi:8998 and exposes port 43337 for metrics. If you bring this up, the tunnel service is what makes the local Moshi instance reachable from outside the host, so the compose file is aimed at a demo deployment rather than a private one.
Where Moshi is the wrong tool
The clearest limitation is fine-tuning. The README does not document a training or fine-tuning workflow inside this repository. It says explicitly that if you want to fine tune Moshi, you should head to the separate kyutai-labs/moshi-finetune project. Anyone who picks up Moshi expecting to adapt it to their own voice or domain from this codebase alone will find the path missing.
The second limitation is the model itself. The released checkpoints are two synthetic voices, Moshiko (male) and Moshika (female), plus Mimi. There is no documented procedure here for adding a third voice. If your product needs a specific speaker identity, the released voices are what you get.
The third is hardware. The compose file requests an NVIDIA GPU, and the latency figure in the README is tied to an L4. The MLX path exists for Apple silicon, so Mac and iPhone inference is a real option, but the PyTorch path is not aimed at CPU-only machines. Finally, the q8 PyTorch checkpoints are labelled experimental, which means if you want the smaller format on that backend you are accepting a less settled configuration.
How Moshi differs from an STT plus LLM plus TTS pipeline
The obvious alternative is the conventional three-stage pipeline built from a speech recognition model, a text language model, and a text-to-speech model. The difference is structural, not a matter of quality. In that pipeline, the recogniser must decide a turn is over before the language model sees anything, and the synthesiser cannot start until the language model has produced enough text. Each stage adds buffering, and the end-to-end delay is the sum of the stages.
Moshi removes the staging. It consumes the user's audio stream and produces its own audio stream continuously, so there is no explicit end-of-turn decision inside the model. The text it predicts is a byproduct of its own speech, not an input from a separate recogniser. That is what allows the 160ms theoretical latency figure to be meaningful in the first place. The cost is that you lose the modularity: you cannot swap in a better language model or a different voice without touching the whole stack, and the model's knowledge comes entirely from what Moshi was trained on rather than from a text model you can prompt or retrieve against.
A second comparison worth noting is against the non-streaming codecs the README names, SpeechTokenizer and SemantiCodec. Those operate at 50 Hz and higher bitrates and are not streaming, so they are not substitutes for Mimi in a real-time loop. If your use case is offline audio compression rather than live dialogue, Mimi's low frame rate is not an advantage.
Maintenance, licensing and the cost of staying current
The repository is not archived, and the last push was on 2026-09-09, which is recent enough that the codebase is being touched. The release history tells a more granular story: the only releases listed are rustymimi-0.2.2 on 2024-09-22 and rustymimi-0.2.1 on 2024-09-20. Those are the Rust Mimi Python bindings, and no release has been cut for them in roughly two years. The PyTorch and MLX packages are distributed through PyPI and the README points to bleeding-edge installation for both, which suggests the Python side is consumed from the repository rather than from tagged releases. If you depend on rustymimi, plan around a binding that has not seen a tagged release recently.
Licensing splits in two. The repository carries LICENSE-APACHE and LICENSE-MIT, so the code is dual-licensed. The models are different: the README states that all models are released under CC-BY 4.0. That means the weights carry an attribution requirement that the code does not, and anyone shipping a product built on Moshiko, Moshika or Mimi needs to handle that separately. The README gives no further guidance on commercial use, so this is a question for your own review rather than something the project answers.
Upgrade cost is tied to which stack you chose. The Rust path pins you to a binding with no recent release. The PyTorch path moves with the repository and with PyPI packages, so a `pip install -U moshi` can pull changes you have not reviewed. The MLX path is tied to Apple's framework and to Python 3.12 being the recommended interpreter.
Three stacks, one model: choosing between PyTorch, MLX and Rust
The repository deliberately maintains three separate inference stacks, and the README frames them by purpose rather than by capability. PyTorch, in the moshi/ directory, is for research and tinkering. MLX, in moshi_mlx/, is for on-device inference on iPhone and Mac. Rust, in rust/, is for production, and it contains the Mimi implementation in Rust with Python bindings published as rustymimi.
That framing is honest about the trade-off. The PyTorch stack is the one you modify when you are studying the architecture or trying a new decoding idea. The MLX stack exists because Apple silicon needs a different execution path, and the README recommends Python 3.12 for it. The Rust stack exists because a Python inference loop is not what you want serving traffic. There is no claim that one stack is more capable than another, only that they target different deployment surfaces. If you are prototyping, start with PyTorch and the bf16 checkpoint. If you are shipping to Mac or iPhone, the MLX checkpoints in q4, q8 or bf16 are the ones to evaluate. If you are serving from a GPU host, the Rust and Candle checkpoints plus the compose file are the path the repository lays out.
Editorial conclusion
Adopt Moshi if you need a research-grade full-duplex speech loop and can work inside the PyTorch or MLX stack, or if you want the Rust and Candle path for a production service with an NVIDIA GPU. Do not adopt it if you need a documented fine-tuning workflow in this repository, since the README points to kyutai-labs/moshi-finetune for that, or if you expect a text-only model, because Moshi's output is audio with a parallel text stream. Before committing, verify which Hugging Face checkpoint matches your backend and quantization, confirm that the model licence (CC-BY 4.0) is compatible with how you intend to distribute output, and check that the Python version you have meets the documented minimum of 3.10.
Frequently asked questions
How do I install Moshi?
The README gives two PyPI commands: pip install -U moshi for the PyTorch client and pip install -U moshi_mlx for the MLX client. You need at least Python 3.10, with 3.12 recommended, and the MLX package is described as best with Python 3.12.
How do I use Moshi?
After installing a client and picking a checkpoint from Hugging Face, you run one of the three inference stacks: the PyTorch code in moshi/ for research, the MLX code in moshi_mlx/ for on-device inference on iPhone and Mac, or the Rust code in rust/ for production. The README does not give a single end-to-end usage example beyond the install commands.
What does Moshi mean in this project's context?
In this repository, Moshi is the name of a speech-text foundation model and full-duplex spoken dialogue framework from Kyutai Labs, built on the Mimi streaming neural audio codec. The README does not discuss the Japanese greeting or any other meaning of the word.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/kyutai-labs-moshi)