# Mochi 1: a 10 billion parameter text to video model you can run locally

> Genmo released the weights, the inference code and the VAE under Apache 2.0. The README asks for about 60GB of VRAM while pointing at ComfyUI for anything smaller, and those two numbers deserve to be read side by side.

**genmoai/mochi** — The best OSS video generation models, created by Genmo

- Repository: https://github.com/genmoai/mochi
- Stars: 3,743 · Forks: 493
- Language: Python
- License: Apache-2.0
- Published: 2026-10-08 · Updated: 2026-10-08 · Language: en
- Canonical page: https://hysenlabs.com/projects/genmoai-mochi

## Installing the inference harness with uv

The whole install is seven commands, and five of them are about Python packaging rather than about video:

```bash
git clone https://github.com/genmoai/mochi
cd mochi
pip install uv
uv venv .venv
source .venv/bin/activate
uv pip install setuptools
uv pip install -e . --no-build-isolation
```

The editable install is what makes `genmo` importable, and the README insists on `--no-build-isolation` because the build needs setuptools already present in the environment. There is an optional extra for flash attention, which the README lists as a separate command rather than folding into the main install.

Two things sit outside Python. FFMPEG has to be installed separately because the pipeline writes video files, and the model weights are not in the repository at all. A reader who expects a `git clone` to produce a working generator will be disappointed by exactly one thing: nothing runs until roughly ten billion parameters of weights are on disk.

The package metadata is minimal and a little generic. `pyproject.toml` names the distribution `genmo`, version `0.1.0`, with the description `Genmo models`, and pins several dependencies to exact versions rather than ranges, including `av==13.1.0`, `moviepy==1.0.3` and `pillow==9.5.0`. The Python floor is `>=3.10`.

## Getting weights from a script, Hugging Face, or a magnet link

Weights come down in three documented ways. The first is the bundled script:

```bash
python3 ./scripts/download_weights.py weights/
```

The script is described as fetching the model plus the VAE into a local directory, and the trailing argument is that directory. The second is the Hugging Face repository, which the README links for direct file browsing. The third is a magnet URI pointing at a public tracker, formatted as a magnet link rather than an HTTP URL, which is unusual for a model README and saves everyone from rehosting several gigabytes.

The repository tree shows the pieces these commands touch. `scripts/` holds the downloader, `demos/` holds the entry points, `src/` holds the library, `contrib/` is separate from the demos, and `assets/` is where media lives. There is a `uv.lock` next to `requirements.txt`, which means the pinned dependency set is committed in two places and can drift.

That drift is real, and it is worth naming. `pyproject.toml` lists `torchvision>=0.19.1` among its main dependencies, and also defines an optional extra named `torchvision` that pins `torchvision>=0.15.0` and `pyav>=13.1.0`. `requirements.txt` lists neither torchvision nor pyav. So the file you install from and the file you audit are not the same set of packages, and the extra named after a package already in the main list is redundant by construction.

## Generating from the CLI, a Gradio page, or the pipeline object

Two demo scripts cover the interactive cases, and both take the same two flags:

```bash
python3 ./demos/gradio_ui.py --model_dir weights/ --cpu_offload
```

```bash
python3 ./demos/cli.py --model_dir weights/ --cpu_offload
```

`--cpu_offload` is the interesting one. It is what makes a single-GPU run possible at all, at the cost of moving weights across the bus during generation. Fine-tuned LoRA weights in safetensors format are added with `--lora_path <path/to/my_mochi_lora.safetensors>` on either script.

For code, the README documents a composable API imported from `genmo.mochi_preview.pipelines`, assembled from a text encoder factory, a DiT factory and a decoder factory. The call itself sets the generation parameters, and they are worth reading as defaults rather than as suggestions:

```python
video = pipeline(
    height=480,
    width=848,
    num_frames=31,
    num_inference_steps=64,
    sigma_schedule=linear_quadratic_schedule(64, 0.025),
```

480 by 848 pixels and 31 frames at 64 steps is a short, low-resolution clip, which matches the README's own limitations section. A full example lives at `demos/api_example.py`, and the `pipelines` import path names three separate factories, which tells you the text encoder, the transformer and the VAE decoder are constructed and owned independently rather than hidden inside one class.

## AsymmDiT's deliberate split between text and visual capacity

The architecture is the part of this repository that generalizes beyond video. Mochi uses a 10 billion parameter diffusion transformer called AsymmDiT, trained from scratch, with 48 layers and 24 heads. The model spec table gives the visual stream a hidden dimension of 3072 and the text stream 1536, which puts roughly four times as much parameter capacity on the visual side.

The asymmetry is structural rather than cosmetic. Text and visual tokens are attended to jointly with multimodal self-attention, but each modality gets its own MLP layers, and the projections into and out of self-attention are non-square so the two streams can meet at the same attention operation with different widths. The stated benefit is lower inference memory, since the text branch carries far fewer parameters and processes only 256 text tokens against 44520 visual ones.

The second choice is the text encoder. The README is explicit that many recent diffusion models lean on several pretrained language models to represent a prompt, and that Mochi encodes prompts with a single T5-XXL. One encoder means one tokenizer, one embedding table and one failure mode to reason about, at the price of giving up whatever a second, smaller encoder was doing for you.

The video side uses AsymmVAE, a 362 million parameter asymmetric encoder-decoder with 64 encoder base channels, 128 decoder base channels and a 12 channel latent space. Compression is 8x8 spatial and 6x temporal, which the README describes as causally compressing video by 128x.

## Two memory numbers and a news log that stopped in 2024

The hardware section and the ComfyUI note contradict each other on the surface, and the honest reading is that they describe two different pieces of software. The README says this repository supports multi-GPU splitting across graphics cards and single-GPU operation, but requires approximately 60GB VRAM on a single GPU. It then says ComfyUI can optimize Mochi to run on less than 20GB VRAM, and that this implementation prioritizes flexibility over memory efficiency, with a recommendation to use at least one H100.

Both facts are true about their own subject. The repository is a flexible reference implementation with context-parallel support, and it is not memory-frugal. ComfyUI's Mochi wrapper, credited in the README's related work, is a different integration with a different budget. Practical judgment rule: judge the memory question by which code you actually plan to run, and treat the 60GB figure as a property of `demos/cli.py` and `demos/gradio_ui.py` rather than of the weights.

The second contradiction is temporal. The repository was last pushed on 2026-10-06, but the README's news log has exactly two entries, dated November 5, 2024 for ComfyUI support and November 26, 2024 for LoRA fine-tuning. Nothing in the README describes what changed in the intervening time. So the file is not stale enough to distrust and not current enough to rely on. Treat the news log as the end of the documented feature work, and the push date as evidence that something moved afterwards without being written down.

The README also calls the model the largest video generative model ever openly released and says it dramatically closes the gap between closed and open systems. Those are preview-checkpoint claims in a document that also describes itself as a living and evolving checkpoint under research preview. A claim about being the largest is only checkable against other releases, which the README does not enumerate.

## Fine-tuning on one GPU, and what the trainer does not fix

The README points at a LoRA trainer under `demos/fine_tuner/`, described as easy to use and as able to build fine-tunes on your own videos. The stated hardware floor is one H100, or one A100 with 80GB. That is a much friendlier number than the 60GB single-GPU inference figure suggests, and it is a sign the trainer has its own memory path rather than reusing the demo scripts unchanged.

What the trainer does not fix is content coverage. The limitations section says Mochi 1 is optimized for photorealistic styles and does not perform well with animated content, that extreme motion can produce minor warping and distortions, and that the community is expected to fine-tune the model toward other aesthetics. Read that as an admission that the base checkpoint is narrow, not as a promise that fine-tuning is cheap.

The safety section is unusually direct for a model card. It says the models reflect biases and preconceptions from their training data, that steps were taken to limit NSFW content, and that organizations should implement additional safety protocols before deploying the weights in commercial products. On licensing, the model is under Apache 2.0 with a `LICENSE` file in the repository root, which is a permissive grant covering both the code and the weights, though the README's framing of the release as a preview is a maturity signal rather than a legal one.

## Conclusion

Mochi 1 is worth reading as an architecture reference even if you never run it, because the AsymmDiT split between text and visual capacity is described precisely enough to imitate. Run it only if you have an H100 or are willing to route inference through ComfyUI, since the repository asks for roughly 60GB of VRAM and the README's own news log stops at November 2024. Start with scripts/download_weights.py and demos/cli.py, check the pinned pillow and av versions in pyproject.toml before you build an environment around them, and decide about the 480p and photorealistic limits before committing a pipeline to it.

## FAQ

### How much VRAM does Mochi 1 need to generate video?

The README states roughly 60GB of VRAM for single-GPU use with this repository, and recommends at least one H100. It also notes that ComfyUI can optimize Mochi to run on less than 20GB, which is a different integration rather than a different setting in this code.

### Can I fine-tune Mochi 1 with LoRA?

Yes. The README links a LoRA trainer under demos/fine_tuner and says it can be fine-tuned on one H100 or one A100 80GB GPU. Fine-tuned weights are loaded back into the demos with the `--lora_path` flag in safetensors format.

### What resolution and frame count does Mochi 1 produce?

The pipeline example in the README uses height 480, width 848 and 31 frames at 64 inference steps. The limitations section confirms the initial release generates 480p video and mentions minor warping in edge cases with extreme motion.

### How do I install Mochi 1 and get the weights?

Clone the repository, create a virtual environment with uv, and run an editable install with `--no-build-isolation`. FFMPEG is needed separately, and the weights are fetched with `python3 ./scripts/download_weights.py weights/`, or taken from the Hugging Face repository or the magnet link the README lists.

## Sources

- [genmoai/mochi on GitHub](https://github.com/genmoai/mochi)
- [Issues](https://github.com/genmoai/mochi/issues)
- [License: Apache-2.0](https://github.com/genmoai/mochi/blob/main/LICENSE)
- [README](https://github.com/genmoai/mochi/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/genmoai-mochi
