# SolarWM: open data and scalable training for long-horizon video world models

> SolarWM is a multi-backbone training framework that turns bidirectional video diffusion models into camera-controlled autoregressive world models. It ships an installable CLI, a documented three-stage recipe, and a data release split into small annotation packages and separate latent generations.

**Junchao-cs/SolarWM** — Open data and scalable training for long-horizon video world models.

- Repository: https://github.com/Junchao-cs/SolarWM
- Website: https://junchao-cs.github.io/SolarWM-Web/
- Stars: 752 · Forks: 39
- Language: Python
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/junchao-cs-solarwm

## What SolarWM is actually for

SolarWM targets a specific gap: training a video model that keeps generating coherently as a user interacts with it, rather than producing a fixed five-second clip. The README frames the project as "a fully open foundation for building interactive video world models from data preparation through scalable training and long-horizon inference." The intended user is a research or engineering team that already has a video diffusion backbone and wants a documented path to a causal, camera-conditioned model.

The claim that matters most is in the fourth bullet: after training only on 5-second sequences, the resulting causal models support real-time interaction with rollouts spanning minutes to hours, without long-sequence fine-tuning or attention-sink mechanisms. That is the design goal, stated by the authors, not a measured result you can verify from the repository alone. The README does not publish latency figures, hardware requirements for real-time interaction, or a definition of what counts as real time.

The second audience is data engineering. SolarWM converts 1.43 million canonical clips from 14 datasets into what it calls a unified, frame-aligned contract covering observations, metric camera geometry, captions, quality metadata, selection and provenance. If your problem is mixture design across heterogeneous sources, the decoupling of source processing from training-mixture design is the part worth reading first.

## The three-stage recipe and what each stage changes

The training progression is the core mechanism, and it is unusually explicit for a research repository. Stage0.5 learns full-clip bidirectional flow matching and establishes the base video, text and camera-conditioned representation. Stage1 combines teacher forcing with the AnyFlow loss in a single stage: clean history conditions noisy target chunks while the model learns both denoising and finite-step flow maps. Stage2 performs DMD via self-gradient forcing, training the causal student on its own autoregressive rollout with a frozen teacher and a trainable critic.

The interesting choice is the elimination of a separate ODE or consistency-distillation initialization before Stage2. Most distillation pipelines need that warm start; SolarWM folds it into Stage1 and states the removal as a design property. Whether that holds across backbones is the question a reader should test, because the README asserts a "shared route across heterogeneous video backbones" without showing per-backend loss curves.

Backend coverage is uneven and the table says so. Wan2.2-5B and MiniMax-H3 have all three stages marked available, with train, infer and preencode interfaces. Wan2.2-14B and LTX-2.5 have Stage0.5 only, with Stage1 and Stage2 listed as "Coming soon". The model family spans four 5B to 33B models across Wan2.2, LTX-2.5 and MiniMax-H3, and the framework claims to preserve each backbone's native representation and objective rather than forcing a common architecture.

## Installing SolarWM and running the environment probe

The README is direct about a constraint that shapes everything else: Wan, LTX and MiniMax-H3 require separate runtime environments. There is no single install that covers all backends. You activate the environment for the selected backbone first, then install the shared SolarWM source into it.

```bash
python -m pip install -e .
solarwm environment probe
```

The first command installs the package in editable mode. The second is the CLI entry point defined in pyproject.toml as solarwm = "solarwm.cli:main", and the README presents it as the way to check that the environment is usable. The README does not document what the probe prints or which checks it performs, so treat a clean exit as the minimum signal, not proof that a backend is correctly wired.

The base package is deliberately small. Core dependencies are numpy>=1.24 and PyYAML>=6.0, with Python >=3.10. Everything heavy lives in optional extras. If you are training on Wan2.2, the wan extra pins exact versions, including diffusers==0.38.0, flash-attn==2.8.3, transformers==5.12.1 and peft==0.20.0. A comment in pyproject.toml explains the diffusers pin: it "Matches the release-tested Wan runtime and its UniPC scheduler behavior."

The h3 extra pulls flash-attn==2.8.3, imageio==2.37.4 and imageio-ffmpeg==0.6.0. The ltx extra is the odd one out: its comment states that the official ltx_core and ltx_pipelines modules are installed from LTX-2, so the extra itself only adds peft and av. That means the LTX path cannot be reproduced from this repository alone.

For the first real run, the README points to per-backend guides rather than inline commands. The Wan2.2 TI2V-5B guide covers Stage0.5, Stage1 and Stage2 commands, and the MiniMax-H3 guide provides launch commands for all three stages plus checkpoint setup. Weights come from the SolarWM model collection on Hugging Face. Those guides are where the actual training invocations live; the top-level README does not duplicate them.

## Data access: annotations are public, payloads are not

This is the part most likely to surprise a new user. The public SolarWM-Data release contains release controls, licenses, recipe and test indexes, small format examples, and the SolarWM-Data-Annotation package. It does not include the full releases-v1/raw-wds/ or releases-v1/latent-wds/ payloads.

There are two routes. The first is to use preencoded latents: each latent generation is published in a separate repository, and the docs/latent-wds.md page lists available downloads plus generations that are still being uploaded. For a released recipe that uses preencoded data, downloading its matching latent generation is sufficient for training and does not require raw-WDS. The second route is to rebuild raw-WDS from annotations. SolarWM-Data-Annotation is annotation-only, with no videos, and contains camera trajectories, captions, metadata, source identities and reconstruction tools.

Raw-WDS is needed only for specific workflows: the full processed video corpus, online encoding, your own latent generation, or any index that points to raw data. The practical consequence is that the phrase "still being uploaded" in the documentation means a recipe can be published before its data is fully available. Check the latent-WDS list before you plan a training run around a specific recipe.

There is also an access form linked from the README badges, separate from the Hugging Face dataset page, which suggests some data paths are gated rather than open. The README does not explain which parts require the form.

## Where SolarWM is the wrong tool

If you want one repository that installs, trains and evaluates without external dependencies, SolarWM is not that. The LTX path explicitly depends on modules installed from LTX-2, and every backend needs its own runtime environment. Reproducing a multi-backend comparison means maintaining several environments with conflicting pins: the dev extra carries transformers>=4.46.2 while the wan extra pins transformers==5.12.1.

Backend maturity is the second constraint. Two of the four listed backends have Stage1 and Stage2 marked "Coming soon", so if your work depends on LTX-2.5 or Wan2.2-14B reaching the autoregressive stage, the code is not in the repository yet. The README states this plainly in the table, which is better than leaving it implicit, but it still narrows the usable surface to Wan2.2-5B and MiniMax-H3.

The long-horizon claim is also the least verifiable part. The README says rollouts span "minutes to hours" and that interaction is real time, but it publishes no benchmark table, no hardware description and no evaluation protocol. For a project whose headline is long-horizon interaction, the absence of a stated evaluation setup is a real gap. Treat the claim as a research direction you will have to measure yourself.

Finally, this is a training framework, not an inference product. There is no mention of a serving layer, quantization, batching strategy or deployment story. If you need to ship an interactive video model to users, SolarWM gives you the trained weights and inference code, not the surrounding system.

## How SolarWM differs from fine-tuning a single backbone

The obvious alternative is to take one video diffusion model, for example Wan2.2, and fine-tune it yourself toward causal or interactive generation. That route keeps one environment, one dependency set and one set of weights, and it avoids the abstraction layer entirely. If you only ever intend to work with one backbone, the multi-backend design in SolarWM is overhead you pay for nothing.

The difference in approach is what you get instead. SolarWM separates source processing from training-mixture design, so the data contract stays fixed while you swap mixtures. It defines a three-stage route with named losses (AnyFlow in Stage1, self-gradient forcing in Stage2) rather than leaving the distillation strategy to you. And it keeps each backbone's native representation and objective instead of normalizing everything to one architecture, which is the opposite trade-off from a framework that ports models into a shared codebase.

A second alternative is to use a general training harness and write the world-model stages yourself. That gives you control over the loss schedule but means reimplementing the teacher-forced initialization and the DMD critic loop. SolarWM's value proposition is that those stages are already written for four backends, with the caveat that only two have all three stages marked available.

## Licence, maintenance and what upgrading costs

SolarWM is Apache-2.0, declared both in pyproject.toml with license = "Apache-2.0" and license-files = ["LICENSE", "NOTICE"], and in the README badge. Apache-2.0 is permissive and includes a patent grant, but the repository also ships a NOTICE file, and Apache-2.0 requires that NOTICE contents be preserved in distributions. There is a second NOTICE.md scoped to the Wan2.2 runtime modeling package, listed in package-data, which suggests third-party code is vendored there. Anyone redistributing a build should read both files rather than assume a single licence covers every directory. This is not legal advice.

The last push to the default branch was on 2026-09-13. The repository is not archived. The release history in the README shows two dates: September 3, 2026 for the initial open-source drop of training and inference code, the dataset, the pipeline and several weight sets, and September 13, 2026 for SolarWM-H3 training and inference code, weights and preencoded training data. There are no versioned releases retrieved, so the upgrade unit is the git history and the pinned extras rather than tagged versions.

Upgrade cost is dominated by the pins. diffusers==0.38.0, transformers==5.12.1, flash-attn==2.8.3 and peft==0.20.0 are exact, and the diffusers comment ties the pin to UniPC scheduler behavior in the release-tested Wan runtime. Moving any of those forward is a compatibility question you answer by rerunning training, not by reading a changelog. The pyproject.toml also carries a uv-specific setting, tool.uv.extra-build-dependencies, that augments the isolated flash-attn build environment with the runtime's exact torch, which tells you flash-attn builds are expected to be fragile enough to need that help.

## Conclusion

Adopt SolarWM if you already run a Wan2.2, LTX-2.5 or MiniMax-H3 pipeline and want the three-stage route to a causal, camera-controlled model without writing your own distillation initialization. Do not adopt it if you need a single self-contained repository that trains end to end: SolarWM depends on separately installed backbone runtimes, and only Wan2.2-5B and MiniMax-H3 currently list Stage1 and Stage2 as available. Before committing, run solarwm environment probe inside the backbone environment, confirm your target backend's row in the stage table, and check whether the latent generation matching your recipe has finished uploading, because the full raw-wds and latent-wds payloads are not part of the public data release.

## FAQ

### What is SolarWM in simple terms?

It is an open framework for building interactive video world models, covering data preparation, scalable training and long-horizon inference. It supports four 5B to 33B models across the Wan2.2, LTX-2.5 and MiniMax-H3 backbones.

### How do I install SolarWM?

Activate the runtime environment for your chosen backbone, then run python -m pip install -e . followed by solarwm environment probe. Backend-specific extras such as .[wan] or .[h3] add the pinned training dependencies.

### Which backends support all three training stages?

Wan2.2-5B and MiniMax-H3 have Stage0.5, Stage1 and Stage2 marked available in the README table. Wan2.2-14B and LTX-2.5 show Stage0.5 only, with Stage1 and Stage2 listed as coming soon.

### Does the public SolarWM-Data release include the full video corpus?

No. It contains release controls, licenses, recipe and test indexes, small format examples and the annotation package, but not the full releases-v1/raw-wds/ or releases-v1/latent-wds/ payloads. Preencoded latent generations are published in separate repositories.

### What licence does SolarWM use?

Apache-2.0, declared in pyproject.toml along with license-files for LICENSE and NOTICE. The repository also ships a separate NOTICE.md inside the Wan2.2 runtime modeling package.

## Sources

- [Issues](https://github.com/Junchao-cs/SolarWM/issues)
- [Junchao-cs/SolarWM on GitHub](https://github.com/Junchao-cs/SolarWM)
- [License: Apache-2.0](https://github.com/Junchao-cs/SolarWM/blob/main/LICENSE)
- [Project website](https://junchao-cs.github.io/SolarWM-Web/)
- [README](https://github.com/Junchao-cs/SolarWM/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/junchao-cs-solarwm
