LeWorldModel: Training a World Model End-to-End from Raw Pixels
Official code base for LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
At a glance
- What is it?
- LeWorldModel (LeWM) is a Joint-Embedding Predictive Architecture that trains stably on a single GPU using only two loss terms. It is aimed at robotics and reinforcement learning researchers who need a compact world model without pretrained encoders or complex multi-objective tuning.
- Who is it for?
- LeWM is the right fit for robotics and model-based control researchers who want a trainable world model with a minimal parameter budget and a clearly defined loss function. Verify that Python 3.10, uv, and the stable-worldmodel package resolve before committing to a training run, and confirm your WandB account is active since the main config requires it.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 130 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The Problem LeWM Targets and Who It Is For
World models that learn entirely from raw pixels without auxiliary supervision have been difficult to stabilize. Existing Joint-Embedding Predictive Architectures (JEPAs) have addressed collapse through exponential moving averages, pretrained visual encoders, or complex multi-term objectives that introduce six or more tunable hyperparameters. LeWM is designed for researchers who want to train a world model end-to-end from scratch without those scaffolds.
The repository accompanies a 2026 arXiv preprint by Lucas Maes, Quentin Le Lidec, Damien Scieur, Yann LeCun, and Randall Balestriero. Its target domain is continuous control: the benchmark environments are pusht, cube, tworooms, and reacher, covering both 2D and 3D tasks. Researchers working on video synthesis, image generation, or natural language modeling will find no useful artifacts here.
Architecture: Two-Loss Training and Collapse Avoidance
LeWM reduces the training objective to two terms. The first is a next-embedding prediction loss: the model predicts the latent representation of the next frame given the current one. The second is a regularizer that enforces Gaussian-distributed latent embeddings, which the README states prevents representation collapse without requiring an exponential moving average of encoder weights or a pretrained vision model.
The result, according to the abstract, is a reduction from six tunable loss hyperparameters to one compared to the previously described end-to-end alternative. The core PyTorch implementation lives in jepa.py, with encoder and predictor components defined in module.py as the ARPredictor, Embedder, and MLP classes. Training is configured through Hydra config files in config/train/, with the primary config at config/train/lewm.yaml. Each dataset is specified by a Hydra override such as data=pusht, which resolves to a YAML file under config/train/data/.
The README's abstract also states that the latent space encodes meaningful physical structure through probing experiments, and that surprise evaluation detects physically implausible events.
Installing LeWM and Running a First Training
LeWM requires Python 3.10 and uses uv for environment management. The installation creates a virtual environment, activates it, and installs the stable-worldmodel package with the train and env extras:
uv venv --python=3.10
source .venv/bin/activate
uv pip install stable-worldmodel[train,env]Datasets are HDF5 archives hosted on HuggingFace in the quentinll/lewm collection. After downloading, decompress each archive:
tar --zstd -xvf archive.tar.zstPlace the extracted .h5 files under the STABLEWM_HOME directory, which defaults to ~/.stable-wm/. To use a different path:
export STABLEWM_HOME=/path/to/your/storageBefore training, open config/train/lewm.yaml and set the wandb.config.entity and wandb.config.project fields to your WandB account details. The README does not document an offline training mode, so a WandB account is required. Launch training with:
python train.py data=pushtCheckpoints are written to STABLEWM_HOME when training completes. The README does not document intermediate checkpoint saving during a run.
Running Evaluation and Loading Pretrained Checkpoints
Evaluation configs live under config/eval/. The policy field takes a path relative to STABLEWM_HOME without the _object.ckpt suffix. The README gives an explicit example of the correct invocation:
python eval.py --config-name=pusht.yaml policy=pusht/lewmPretrained checkpoints covering all training environments are available on HuggingFace: quentinll/lewm-pusht, quentinll/lewm-cube, quentinll/lewm-tworooms, and quentinll/lewm-reacher. HuggingFace checkpoints ship as a weights.pt state dict and a config.json, and require conversion to the _object.ckpt format before eval.py can use them. After conversion, loading uses the stable_worldmodel Python API:
import stable_worldmodel as swm
# Load the cost model (for MPC)
cost = swm.policy.AutoCostModel('pusht/lewm')AutoCostModel takes a run_name relative to STABLEWM_HOME and an optional cache_dir override. The returned module is in eval mode with weights accessible via .state_dict().
A full baseline comparison suite including PLDM, LeJEPA, IVL, IQL, GCBC, DINO-WM, and DINO-WM-noprop checkpoints for all four environments is available on Google Drive, as listed in the README.
Concrete Limitations and Where LeWM Does Not Apply
LeWM predicts future embeddings in latent space, not future pixel frames. Any research problem that requires pixel-level video prediction or reconstruction is outside its scope.
The repository has no GitHub releases. There is no stable version tag to pin in a downstream project, and the README does not describe a versioning or stability policy. Breaking changes in the stable-worldmodel or stable-pretraining dependencies could affect LeWM without warning. Both are external packages under the galilai-group organization, and their own release cadences are independent of this repository.
Training requires a WandB account with a valid entity and project configured before the first run. The README does not document how to disable WandB logging or run in offline mode. The model also assumes reasonably synchronized clocks across nodes if planning is distributed, as stated in the stable-worldmodel documentation.
DINO-WM: A Different Architecture for the Same Problem
The benchmark table in the README includes DINO-WM as a direct comparison. DINO-WM is a world model that uses DINO-pretrained visual features as a fixed encoder and learns a separate dynamics model on top of those features. The fundamental difference is that DINO-WM starts from pretrained visual representations, while LeWM trains its encoder jointly with the prediction objective using only the two-loss formulation.
That architectural choice has a trade-off: LeWM avoids the dependency on a large pretrained vision model and is lighter in total parameters, but it must learn stable visual representations and temporal dynamics at the same time. In environments where DINO-pretrained features transfer well, DINO-WM may converge more reliably without tuning the Gaussian regularizer. LeWM's advantage is in settings where no suitable pretrained encoder exists, or where a low parameter count and a simple dependency graph are more important than leveraging prior visual knowledge.
Maintenance Status and License
The last push to the repository was on 2026-05-26. The repository is not archived. The project is licensed under the MIT License, which permits use in both research and commercial contexts without a patent grant.
The codebase delegates environment management, planning, and evaluation to the stable-worldmodel package and training to stable-pretraining. Both packages are listed as dependencies in the README but their license terms and release cadences are managed independently of this repository. Any deployment outside academic replication should verify those dependency licenses and check for breaking changes before building on them.
Editorial conclusion
LeWM is the right fit for robotics and model-based control researchers who want a trainable world model with a minimal parameter budget and a clearly defined loss function. Verify that Python 3.10, uv, and the stable-worldmodel package resolve before committing to a training run, and confirm your WandB account is active since the main config requires it. Pixel-level video prediction is out of scope, and the absence of GitHub release tags means there is no version to pin.
Frequently asked questions
How does LeWorldModel prevent representation collapse without an exponential moving average?
LeWM uses a Gaussian regularizer as its second loss term, which constrains the encoder to produce embeddings following a Gaussian distribution. The README states this is sufficient to prevent collapse when combined with the next-embedding prediction loss, eliminating the need for an EMA of encoder weights.
What hardware does LeWorldModel require for training?
According to the README, LeWM has approximately 15 million parameters and is designed to train on a single GPU in a few hours. The README does not specify a minimum GPU model or VRAM requirement.
Where are pretrained LeWorldModel checkpoints available?
Pretrained checkpoints for the pusht, cube, tworooms, and reacher environments are hosted on HuggingFace under the quentinll/lewm collection. A full baseline suite including DINO-WM and other comparison methods is available on Google Drive, as linked in the README.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lucas-maes-le-wm)