MiniMind-O: A Complete Omni Model Training Pipeline at 0.1B Parameters
🎙️ A 0.1B Omni model trained from scratch, capable of listening, speaking, and seeing!
At a glance
- What is it?
- MiniMind-O is an educational omni model project that implements a complete text, audio, and vision training pipeline from scratch using native PyTorch. The main minimind-3o model has approximately 0.1 billion parameters and can be trained on a single RTX 3090 in about two hours using the mini dataset. Its purpose is to let anyone read and run the full end-to-end pipeline rather than call a finished API.
- Who is it for?
- MiniMind-O is the right starting point for researchers and practitioners who want to understand how an end-to-end omni model is assembled: how speech and vision encoders connect to a language backbone, how the Talker generates audio codes, and how barge-in and voice cloning are implemented. It is not a production inference system.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 8 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What MiniMind-O Is For and Who Should Use It
Several open-source omni models appeared in 2025 and 2026, including Mini-Omni2, Moshi, GLM-4-Voice, and Qwen3-Omni. Most of them publish weights and inference code, but the training pipeline is either absent or requires distributed infrastructure. MiniMind-O fills the gap for engineers who want to understand the full omni pipeline by training it themselves.
The project is the third in a series from the same author: MiniMind covered language models from scratch, MiniMind-V extended it to vision multimodal, and MiniMind-O adds audio to produce a model that can accept text, speech, and image input and produce both text output and streamed speech output.
All core algorithm code is implemented using native PyTorch without relying on high-level abstractions from external frameworks. The README states this explicitly as a design goal: the implementation should be readable from the first line without requiring familiarity with a large framework's internal conventions. The result is a codebase that sacrifices performance optimizations for clarity.
The Thinker-Talker Architecture
MiniMind-O splits inference into two paths that run over the same backbone.
The Thinker is the understanding and reasoning component. It processes text tokens, audio features (encoded by SenseVoice-Small), and vision features (encoded by SigLIP2) and produces text output at the semantic level. The speech and vision encoders are frozen; their output is projected into the MiniMind hidden space through two-layer MLP projectors.
The Talker handles speech synthesis. Instead of generating text and passing it to a separate TTS system, the Talker predicts multiple layers of Mimi audio codes simultaneously using MTP (Multi-Token Prediction). Mimi is the audio codec used in this project: it operates at 8 codebook layers, a 12.5 Hz frame rate, and 24 kHz audio output. The CAMPPlus speaker encoder enables voice cloning by conditioning the Talker on a reference audio sample.
This architecture avoids the cascade approach where ASR converts speech to text, the LLM processes the text, and TTS generates the response. Cascading introduces latency and loses prosody information at each conversion step. In MiniMind-O, speech features enter the model at the hidden state level and audio codes exit directly from the Talker, keeping the latency lower and preserving intonation signals through the chain.
VAD (voice activity detection) enables barge-in: the model can detect when the user starts speaking during a response and interrupt playback. The WebUI supports near-duplex interaction where the model listens while still generating audio.
Setting Up and Running Inference
Clone the repository and install dependencies:
git clone --depth 1 https://github.com/jingyaogong/minimind-o
pip install -r requirements.txt -i https://pypi.tuna.tsinghua.edu.cn/simpleDownload the required encoder models and the audio codec using modelscope:
modelscope download --model gongjy/SenseVoiceSmall --local_dir ./model/SenseVoiceSmall
modelscope download --model gongjy/siglip2-base-p32-256-ve --local_dir ./model/siglip2-base-p32-256-ve
modelscope download --model gongjy/mimi --local_dir ./model/mimi
modelscope download --model gongjy/campplus --local_dir ./model/campplusDownload the published minimind-3o weights:
modelscope download --model gongjy/minimind-3o-pytorch --local_dir ./outRun command-line inference:
python eval_omni.py --load_from model --weight sft_omniAlternatively, download from HuggingFace and use the transformers format:
git clone https://huggingface.co/jingyaogong/minimind-3o
python eval_omni.py --load_from minimind-3oThe README gives the author's reference hardware as an Intel i9-10980XE with 128 GB RAM and eight RTX 3090s running Ubuntu 20.04 with CUDA 12.2 and Python 3.10. This reflects training infrastructure, not the minimum inference requirement; the README notes that CPU-only inference is possible given the 0.1B parameter count.
Training on the Mini Dataset
The project ships two dataset sizes. The mini dataset is intended for verifying the full training pipeline quickly. On a single RTX 3090 with 24 GB of VRAM, the README states that running the complete SFT pipeline on the mini dataset takes approximately two hours. The full dataset corresponds to the published weights and covers Chinese speech and image tasks at greater scale.
Training follows a sequence of stages. The first stage trains the text-to-audio path (T2A), the second trains the audio projection layer (A2A mode), and later stages add vision. Each stage uses torchrun with specific flags for the data path, epochs, batch size, and which weights to load and save.
The trainer/ directory contains the training scripts and a train.sh convenience script. The full training command includes flags such as `--learning_rate`, `--data_path`, `--use_compile`, `--from_weight`, `--save_weight`, `--max_seq_len`, and `--use_moe`. DDP multi-GPU training is supported for the full dataset path.
The mini dataset can be downloaded from HuggingFace at the dataset page listed in the README. Only the `_mini` variant is needed for a first run through the pipeline.
Voice Cloning and the WebUI
MiniMind-O ships five built-in voice prompts and seven additional unseen voice prompts for testing voice cloning. Any arbitrary reference audio can be used as the conditioning input: the CAMPPlus speaker encoder converts the reference to a speaker embedding that conditions the Talker's output.
The WebUI has two modes. scripts/web_demo_omni.py is a Gradio application that accepts uploaded or recorded audio, resamples it to 16 kHz on the backend, and returns generated audio as a file. This is not real-time. The webui/web_demo.py is the real-time voice call interface that supports barge-in and near-duplex interaction. A telephone mode is also available through the WebUI for lower-latency audio.
To launch the Gradio demo, the transformers-format model directory must be copied into scripts/ before running the script, since it scans that directory for subdirectories containing weight files. The README notes this requirement explicitly and warns that the demo will error without the model files in place.
MiniMind-O Against Production Omni Models
Qwen3-Omni is an omni model published by Alibaba that accepts text, audio, and image input and produces text and audio output. It is designed for production-quality inference and is available in multiple parameter sizes intended for real-world deployment. Training Qwen3-Omni from scratch requires distributed infrastructure that goes well beyond a single consumer GPU.
MiniMind-O is not a competitor to Qwen3-Omni in terms of output quality. The 0.1B backbone produces limited response quality that the README does not present as production-ready. The value is the pipeline, not the output: a researcher reading MiniMind-O's code can trace every step from raw audio waveform through SenseVoice encoding, MLP projection, transformer inference, MTP audio code prediction, and Mimi decoding, all in plain PyTorch without a framework wrapper obscuring the logic.
The technical report at arxiv.org/abs/2605.03937 covers the architecture, training curves, CER and WER evaluation, voice cloning similarity scores, and cross-model comparisons.
Published Models and License
Two models were released on 2026-05-05: minimind-3o with approximately 0.1 billion parameters and minimind-3o-moe, a mixture-of-experts variant with approximately 0.3 billion total parameters and 0.1 billion active parameters. Both are available on HuggingFace and ModelScope.
The project is licensed under Apache 2.0. The repository is not archived. The last push was on 2026-09-22. The requirements.txt pins specific versions of the key dependencies, including transformers 4.57.6, gradio 5.49.1, funasr 1.3.1, and modelscope 1.37.0. PyTorch and torchaudio are listed in comments without version pins, indicating the user should install a CUDA-compatible version separately.
Editorial conclusion
MiniMind-O is the right starting point for researchers and practitioners who want to understand how an end-to-end omni model is assembled: how speech and vision encoders connect to a language backbone, how the Talker generates audio codes, and how barge-in and voice cloning are implemented. It is not a production inference system. The minimind-3o backbone is 0.1 billion parameters, which produces noticeably limited response quality compared to larger omni models designed for real-world use. Anyone who needs a deployable voice assistant should look at production-scale models. Anyone who wants to read, modify, and retrain a complete omni pipeline without a cluster should start here. The technical report at arxiv.org/abs/2605.03937 documents the architecture in full.
Frequently asked questions
What hardware does MiniMind-O require for training?
The mini dataset pipeline completes in about two hours on a single NVIDIA RTX 3090 with 24 GB of VRAM. The full dataset training uses the same single-GPU setup but takes longer. DDP multi-GPU training is supported for larger-scale runs.
What is the difference between MiniMind, MiniMind-V, and MiniMind-O?
MiniMind is a language model trained from scratch. MiniMind-V adds vision multimodal capability. MiniMind-O is the third project in the series, adding audio input and streaming speech output to produce a model that handles text, speech, and image together.
Can MiniMind-O run inference on a CPU?
The README notes that the 0.1B parameter count of minimind-3o allows CPU inference. It uses the same eval_omni.py script as GPU inference; CUDA is not required for the inference step, only for training.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/jingyaogong-minimind-o)