DW05 announces three expert heads and ships two, and its sample record contradicts its own field list
An Open-Source World Model for Action-Conditioned Embodied Intelligence.
At a glance
- What is it?
- OpenDW is the repository for DW05, a multimodal world model that predicts future video, generates actions and estimates state value from robot observations, released with weights, inference code and training code. The interesting material is in the mismatches: the value head is promised for a later release, the manifest pins numpy and boto3 exactly while claiming only Python 3.10 and 3.11, and the worked JSONL example carries a success flag the field list never asks for and omits the action array the field list requires.
- Who is it for?
- Read this as a research release with an honest gap in it. The video and action paths are the ones to try, and the sample layout, the five dataset environment variables and the norm stats step are documented closely enough to build a small RobotWin-style split of your own.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 89 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A repository named opendw that installs as dexbotic-dw05
Three names are in play and they do not match. The repository is opendw, the distribution is dexbotic-dw05 at version 0.1.0, and the project calls itself DW05 throughout its own documentation. The install instruction reflects the middle name rather than the repository name:
cd /path/to/dexbotic-dw05
pip install -e .Anyone who cloned into a folder called opendw has to correct that path by hand. What ends up in the wheel is narrower than the repository: the package discovery rule includes only names beginning with dexbotic and allows namespace packages, so the playground directory holding the experiment entry point, the script directory, the hardware directory and the assets directory are all outside the distribution and only exist in a clone. The runtime notes sit in a single Markdown file at the root rather than in a docs tree, and the repository has no tagged releases at all.
Three expert heads, and the value one is not shipped
The architecture paragraph is short and load bearing. Inputs are language, image and video, robot type, state and action. The framework is described as MoT, in which a Wan backbone is followed by three expert heads, one each for video, action and value. The introduction lists the same three capabilities as what the model unifies: future video prediction, action generation and state value estimation. Then a single note undercuts the third one, stating that the Value Expert will be updated in a later release. So one third of the announced design is a placeholder today, and a project built on state value estimation has nothing to load. Two more names in that area are unexplained by the manifest: the runtime variables set both a DW05 model root and a DiffSynth model root to the same value, and no dependency carries either the backbone name or a diffusion synthesis package.
Two checkpoints, and the release note names fewer things than the intro
Weights are pulled from a model hub rather than from the repository, and two entries are offered: a base model pretrained on a multi source dataset, and a second one described as a Robotwin 2.0 fine tuned model for evaluation. The table labels that second row SFT-RBT while the address it points at is named for Robotwin, so the label and the repository name use different shorthands for the same checkpoint.
huggingface-cli download Dexmal/DW05-Base --local-dir ./checkpoints/DW05-Base
huggingface-cli download Dexmal/DW05-Robotwin --local-dir ./checkpoints/DW05-RobotwinThe update log holds a single entry, the initial public release in July 2026, and it mentions weights and inference code. The introduction claims weights, inference code and training code. The last commit in the repository is dated the ninth of July 2026, so nothing has landed since that entry. A second model hub also appears in the dependency list, alongside the first.
The worked record carries a flag the field list never asks for
The data format is specified twice, once as a list of what a sample must provide and once as a complete line, and the two disagree. The worked line is a single JSON object per frame, with one entry per camera:
{"type":["action","wm"],"images_1":[{"type":"video","url":"videos/front/episode_000000.mp4","frame_idx":0}],"images_2":[{"type":"video","url":"videos/left/episode_000000.mp4","frame_idx":0}],"images_3":[{"type":"video","url":"videos/right/episode_000000.mp4","frame_idx":0}],"robot":{"prompt":"pick up the red block","state":[0.0,0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8,0.9,1.0,1.1,1.2,1.3],"subtask":"reach the red block"},"worldmodel":{"caption":"A video recorded from a robot's point of view executing the following instruction: pick up the red block"},"robot_task_success":1}The field list requires action trajectories, and the type list above it is allowed to contain action, yet the example contains no action array at all. The example carries a task success flag that the list never mentions. The three camera keys map to front, left and right in the example and can be remapped through a configuration key, and an image only sample replaces each video item with an image item pointing at a frame file. The same fields repeat for every frame of an episode, with the frame index, the image paths and the robot state changing as the line count grows.
Norm statistics stop after 500 batches of 128
Action normalization statistics are computed by the same experiment entry point with a dedicated flag, and the example pins the cost of the pass:
python playground/example_dw_exp.py --compute-norm-stats \
data_config.annotations=/path/to/annotations \
data_config.data_path_prefix=/path/to/media/root \
data_config.text_embedding_cache_dir=/path/to/text_embeddings \
norm_stats_config.norm_save_path=./runs/dw05/norm_stats \
norm_stats_config.batch_size=128 \
norm_stats_config.max_batches=500A batch size of 128 with a ceiling of 500 batches means the statistics are drawn from at most 64,000 samples, which is a fraction of a large Robotwin style dataset and worth knowing before you treat the resulting means as global. The output lands at a fixed path under the run directory, and it can then be supplied three ways: as a configuration value, as an environment variable, or as a runtime argument, the last one qualified as available where it exists rather than everywhere. For computing action statistics alone, three fields are said to be enough.
Five dataset variables, one recipe, and an override path
Dataset wiring is done with five environment variables, one each for the annotations directory, the text embedding cache, the normalization statistics file, a prefix for the media root and a prefix for the index root. Three more variables cover runtime, and they include a setting that disables tokenizer parallelism alongside the two model roots. Paths can be local or remote, with remote reads going through a file library that handles object storage, and a mirror variable takes one or more local roots joined by the platform path separator, which is how a remote layout gets mapped onto local disk.
export DW05_MODEL_BASE_PATH=/path/to/local/model/root
export DIFFSYNTH_MODEL_BASE_PATH=$DW05_MODEL_BASE_PATH
export TOKENIZERS_PARALLELISM=falseA built in recipe named after the benchmark is registered inside the data source module, and it can be bypassed entirely by passing configuration overrides on the command line to the example experiment, which ships with a task called smoke for exactly that purpose.
Exact pins on numpy and boto3, classifiers that stop at 3.11
The manifest is more precise in some places than others. Three dependencies are pinned to exact versions, numpy, boto3 and botocore, while the machine learning stack uses lower bounds: torch at 2.6 or newer, transformers at 4.49 or newer and deepspeed at 0.18 or newer. Deepspeed being a hard requirement rather than an extra means it is installed even on a single machine with no distributed training in mind. Two video decoding libraries are listed side by side, two augmentation and imaging libraries are listed, and a serving trio of fastapi, uvicorn and a multipart parser is present for exposing the model. Python 3.10 and above is required, yet the classifier list advertises only 3.10 and 3.11, so a 3.12 interpreter installs cleanly without being claimed. The optional attention extra installs flash attention and xformers with no version constraints at all, and the text asks you to install them only when they fit your stack.
Editorial conclusion
Read this as a research release with an honest gap in it. The video and action paths are the ones to try, and the sample layout, the five dataset environment variables and the norm stats step are documented closely enough to build a small RobotWin-style split of your own. Two things to check before you rely on it: the value head is explicitly deferred, so anything depending on state value estimation has to wait, and the runtime expects both a model root and a second variable pointing at the same root under a different library's name, which the dependency list does not explain. The last commit in the repository is dated the ninth of July 2026, and there are no tagged releases, so pin to a checkpoint rather than to a version.
Frequently asked questions
What is DW05 in the opendw repository?
DW05 is a multimodal world model for embodied decision making, meant to understand how actions shape future outcomes in robot environments. It unifies future video prediction, action generation and state value estimation in one framework, and the repository releases weights, inference code and training code under Apache-2.0.
Is the DW05 value head available today?
No. The architecture uses a backbone followed by three expert heads for video, action and value, and a note in the README states that the Value Expert will be updated in a later release.
Where do I get the DW05 model weights?
From a model hub rather than from the repository, which has no releases. Two entries are offered: a base model pretrained on a multi source dataset, and a Robotwin 2.0 fine tuned model for evaluation, downloaded with the huggingface-cli download command.
What does a DW05 training sample have to contain?
Camera observations referenced by three configurable image keys, a type list holding action, wm, or both, a robot prompt at robot.prompt for action samples, a world model caption at worldmodel.caption for video samples, robot state, optional proprioception, action trajectories, cached text embeddings or a configured fallback, and optional normalization statistics in a named format.
How is the DW05 dataset pointed at a local directory?
Through five environment variables, one each for the annotations directory, the text embedding cache, the normalization statistics file, the media path prefix and the index path prefix. A built in recipe named robotwin_baseline can be used instead, or bypassed by passing configuration overrides to the example experiment entry point.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/dexmal-opendw)