minWM ships two layouts and installs only one of them
A Minimal and Elegant Framework & Tutorial for Real-Time Interactive World Models
At a glance
- What is it?
- An Apache-2.0 framework and tutorial for turning a bidirectional text-to-video foundation model into an action-conditioned world model, with a four-stage distillation recipe, two backbones at very different sizes, and Claude skills that encode the failure modes its authors hit. Its runtime dependencies are deliberately kept out of pyproject.toml, so the documented install is five commands rather than one.
- Who is it for?
- minWM is a teaching repository that happens to be runnable, and that framing is stated at the top rather than hidden. It documents the four-stage path from a bidirectional diffusion model to a four-step autoregressive student, ships two reference integrations rather than one, and packages its hard-won failure modes as skills so a newcomer is not reverse-engineering someone else's repo.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 30 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 10, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The distillation is four stages and stage two forks
The research content of the repository is a single pipeline diagram, and it is worth reading closely because the fork in the middle is the interesting part.
Phase 1 Phase 2 — Distillation to Causal Few-Step
───────────────────── ────────────────────────────────────────────
Bidirectional SFT ──▶ Stage 1 Teacher Forcing AR Diffusion
Stage 2a Causal ODE (proposed in [Causal Forcing](https://arxiv.org/abs/2602.02214))
Stage 2b Causal CD (proposed in [Causal Forcing++](https://arxiv.org/abs/2605.15141))
Stage 3 Asymmetric DMD with Self Rollout
▼Phase 1 is bidirectional supervised fine-tuning, which starts from the foundation model rather than from scratch. Stage 1 then trains the student autoregressively with teacher forcing, so it sees ground-truth frames rather than its own. Stage 2 is where the path splits: a causal ODE method attributed to Causal Forcing, and a causal consistency-distillation method attributed to Causal Forcing++. Stage 3 is asymmetric distribution matching distillation with self rollout, which is what collapses the model to four inference steps.
The value of the framework is that all four stages are in one place with the same abstractions behind them, and the diagram's own framing is that a third backbone would be a wrapper-and-config exercise rather than new research.
Two backbones that differ in architecture and by six times the parameters
The multi-backbone claim rests on a table with two rows, and the two rows are not similar.
| Backbone | Architecture | Params | Training | Inference | | -------------------- | --------------------- | ------ | -------------- | ------------ | | Wan 2.1 | Cross-attention + DiT | 1.3 B | all 4 stages | 4-step DMD | | HunyuanVideo 1.5 | MMDiT | 8 B | all 4 stages | 4-step DMD |
One is a cross-attention plus diffusion transformer at 1.3 billion parameters, the other is an MMDiT at 8 billion, and both are run through all four training stages and both end at four-step DMD inference. That is the actual evidence behind the portability claim, so it is worth being clear that it is two data points, one of which is small enough to train on modest hardware and one of which is not.
Inference is where the backbones converge. The page lists 4-step DMD inference for HY Action2V, HY TI2V and Wan Action2V, with multi-GPU sequence parallelism for the inference path as well as the training path. The training description adds FSDP with sequence parallelism and both single-node and multi-node runs.
The same trainer, loss and dataset abstractions serve both lines, which is the structural claim the repository rests on, and the project says the point is that adding a backbone should be configuration work.
Two steps to install, because runtime deps are not declared
The install recipe is short enough to copy whole, and the last line explains why it is not one command.
conda create -n minwm python=3.12 -y
conda activate minwm
pip install -r requirements/base.txt
pip install flash-attn --no-build-isolation
pip install -e . # editable install: makes `import minwm` resolve, no PYTHONPATHThe packaging file says the reason in a comment: only the `minwm` package is installed, runtime dependencies stay pinned in `requirements/` and are not declared in the metadata, so install is two steps, the requirements file and then the editable install. The consequence for a user is that `pip install -e .` alone produces an importable package with none of its runtime dependencies, and `flash-attn` is a third step installed without build isolation.
Versions are looser than that. The conda command asks for Python 3.12 while the project requires 3.10 or newer, and the classifier list names only 3.10, so the environment the page tells you to create is newer than the one the metadata advertises.
The rest of setup is delegated to `INSTALL.md`, which also covers verification, developer setup, troubleshooting, and remote checkpoint storage. That last one matters for anyone training off a cluster: checkpoints can go to `s3://`, `oss://` or an S3-compatible store such as Baidu BOS, and the page says to install the matching fsspec backend.
The formatter still excludes three directories the tree no longer has
The top-level entries are `minwm/`, `configs/`, `demos/`, `tools/`, `scripts/`, `requirements/`, `docs/`, `docker/`, `licenses/`, `assets/` and `tests/`, beside `INSTALL.md`, `MIGRATION.md`, `pyproject.toml`, `mkdocs.yml` and `CLAUDE.md`.
Three directory names do not appear there, and they appear everywhere in the configuration. The packaging comment describes the legacy trees as `HY15/`, `Wan21/` and `shared/`, states they are not packaged, and says their run-scripts set `PYTHONPATH` themselves. The formatter settings repeat all three: the Black exclude pattern lists them and the isort skip list names them again, with a comment saying they are excluded until migrated.
So the tooling still carries the shape of the older layout while the layout itself has been consolidated into `minwm/`. Practically that means a formatter pass silently ignores three path patterns that no longer resolve, and anyone reading the config for guidance about the legacy code finds comments pointing at directories they cannot open.
The news section dates that transition. The 2026-09-05 entry announces a rebuilt version with optimized infrastructure, points at `MIGRATION.md` for anyone on the old layout, and records that the old layout is preserved at the tag `v0.1-legacy`. The last commit on the default branch is dated 2026-09-10, days after that release note.
The failure modes are the documentation
Three Claude skills sit at the center of the project's onboarding story, and one of them is a list of things that went wrong.
`debug-world-model` collects failure modes observed in the training pipeline: loss going to NaN, frame-to-frame jitter, camera drift, memory attenuation, and distillation collapse. The stated purpose is that an assistant diagnoses likely root causes from symptoms rather than guessing.
`integrate-new-backbone` is a step-by-step recipe for plugging a new video diffusion transformer into the framework, grounded in the two existing integrations, with the example given as looking at how one backbone does teacher forcing and doing the same for another model.
`onboarding-world-model` is aimed at people entering the field for the first time and has two halves. Foundations covers the minimum background to follow the pipeline, which is teacher forcing for autoregressive diffusion training plus causal forcing and causal forcing plus for distillation. Pitfalls holds the non-obvious mistakes from building the framework itself.
The audience is stated without hedging: graduate students, independent researchers, and junior labs that want into the world-model space without spending three months reverse-engineering existing repositories. The repository supports that claim structurally, since `.claude/` and `CLAUDE.md` are both at the root rather than tucked into a subdirectory.
Camera control is a run-length string with no legend
The inference section names the control surface and shows one example of it. Camera trajectories are driven through pose strings, written in a compact form such as `"a*4,w*8,s*7"`, or through JSON files.
The notation is run-length encoded, which is clear from the shape, and it compresses well enough to be worth the loss of clarity. What the visible text does not do is define the letters. There is no statement of what `a`, `w` and `s` correspond to, whether the trailing numbers are counts or indices, or how the string relates to the JSON form, which is offered as an alternative without a mapping between the two.
That is the one place in the documentation where a newcomer has to read code rather than text, and it sits in the sentence that describes what you can do with the finished model. The demos directory and the config files are where the notation has to be decoded.
For a framework whose selling point is that every stage is inspectable, a control interface with an undocumented vocabulary is the gap most worth closing before you build tooling on top of it.
Licensing is split across three paths and two CI systems
The housekeeping tells you how the project is run rather than what it does, and a few details are worth having.
Licensing takes three places: the Apache-2.0 declaration in the packaging metadata, a `NOTICE` file, and both `THIRD_PARTY_LICENSES.md` and a `licenses/` directory at the root. For a project that starts from existing foundation models and downloads third-party base weights, the HunyuanVideo base checkpoint being fetched from its own model card is exactly the case those paths exist to cover, so they are the first place to look rather than the license field.
Continuous integration is configured twice, with a `.github/` directory and a separate `.gitlab-ci.yml`, alongside `.flake8`, `.pre-commit-config.yaml` and a Black and isort configuration scoped to two directories. The formatter line length is 100 and the target version is Python 3.10, matching the classifier rather than the conda command.
Documentation has its own tool, `mkdocs.yml`, separate from the docs directory, and the repository also carries a `docker/` directory, `CONTRIBUTING.md`, and a `scripts/` tree. Version numbering is dynamic, read from `minwm.__version__` at build time rather than written in the metadata, so the released version lives in the package source.
Editorial conclusion
minWM is a teaching repository that happens to be runnable, and that framing is stated at the top rather than hidden. It documents the four-stage path from a bidirectional diffusion model to a four-step autoregressive student, ships two reference integrations rather than one, and packages its hard-won failure modes as skills so a newcomer is not reverse-engineering someone else's repo. Four things to know. The install is two steps on purpose, with runtime dependencies pinned under `requirements/` instead of declared in the packaging metadata, and `flash-attn` installed separately without build isolation. The Python versions do not line up neatly either, since the conda command asks for 3.12 while the metadata floor is 3.10 and the classifier list stops at 3.10. The formatter configuration still names three legacy directories that the current tree no longer shows, which tells you how recently the migration landed. And the camera-control notation is given as a pose string such as `a*4,w*8,s*7` without a legend, so plan to read the demo scripts rather than guess the letters.
Frequently asked questions
What does minWM provide, and what is it not?
It is a full-stack framework and tutorial rather than a specific model: the complete data, training and inference pipeline for turning a bidirectional text-to-video foundation model into an action-conditioned video world model, with example data, runnable scripts, Claude skills and onboarding material. The page states that framing explicitly rather than presenting a checkpoint of its own.
Which video backbones does minWM support?
HunyuanVideo 1.5 with an MMDiT architecture at 8 B parameters, and Wan 2.1 with cross-attention plus DiT at 1.3 B. Both lines run all 4 training stages and both end at 4-step DMD inference, sharing the same trainer, loss and dataset abstractions.
How do I install minWM and its dependencies?
Create a conda environment with Python 3.12, then pip install -r requirements/base.txt, then pip install flash-attn --no-build-isolation, then pip install -e . for the editable install. Runtime dependencies are deliberately not declared in pyproject.toml, so the editable install alone leaves the environment without them. INSTALL.md covers verification, developer setup and remote checkpoint storage on s3, oss or S3-compatible stores.
What do the Claude skills in minWM cover?
debug-world-model collects training failure modes including loss NaN, frame-to-frame jitter, camera drift, memory attenuation and distillation collapse. integrate-new-backbone is a recipe for adding a new video DiT, grounded in the two reference integrations. onboarding-world-model has a Foundations half on teacher forcing and causal forcing, and a Pitfalls half.
How is the camera trajectory controlled in minWM?
Through pose strings in a compact run-length form such as "a*4,w*8,s*7", or through JSON files. The visible text names both forms and shows the string example, but does not define what the letters stand for or how the two formats relate.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/shengshu-ai-minwm)