OpenWAM, whose architecture and benchmark tables do not render
Official repository for "OpenWAM: An Open, Modular Exploration Towards Systematic World–Action Model Pretraining".
At a glance
- What is it?
- A research stack for World-Action Models, released as a codebase in September 2026 alongside a paper and a set of hosted checkpoints. The most specific part of the front page, covering six architecture variants and ten benchmark suites, sits inside an HTML comment and never appears when the page is rendered. The dependency list below it is unusually disciplined: twenty-nine entries, every one with an upper bound.
- Who is it for?
- OpenWAM is worth reading if you are working on robot policies and want the design space laid out rather than one architecture handed to you, because the variant table and the benchmark status list are exactly what most papers leave out. Three things to know before you install anything.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 10, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The architecture and benchmark tables are inside an HTML comment
The most specific content on the front page does not render. The Support Status heading and everything beneath it, which is the architecture table and the benchmark table, sits between an opening comment marker and a closing one, along with the repository layout listing that shows the package tree. On a rendered page the text jumps straight from the three-part description of the project to the news list. That is an editorial state rather than a rendering quirk, and it matters to anyone reading the page on a host that strips comments, which is most of them. The commented content is worth having: six architecture variants spread across three families, and ten benchmark rows of which eight are marked supported, one is marked external, and two are planned, with both planned ones carrying the same note about external environment setup.
Three architecture families, selected by one config key
Inside the comment, the model choices are three families, two of which carry two variants each. The single system puts video, action, and state tokens into one sequence through a shared DiT, and its second variant replaces the feed-forward layers on the bridge layers with mixture-of-experts ones. The dual system keeps a separate action DiT and video DiT and fuses them either per layer through a single mixed self-attention, or as a two-stage run in which the video DiT completes, its features are handed over, and the action DiT runs once against them. That second form has its own switch, which decides whether action gradients flow back into the video DiT or are blocked and the action side trains on detached video features. A third family adds a frozen language-and-vision understanding expert into the sequence. Selection is one key, `architecture.variant`, inside a per-framework YAML file.
Ten benchmark rows, eight supported, and two that need an outside simulator
The benchmark table is a status list rather than a leaderboard, and no scores appear anywhere in it. The supported rows cover a fifty-task manipulation suite, a benchmark and a perturbation suite layered over it, a household suite with a named state and action dimension contract, a humanoid tabletop suite, a primitive-task suite spread across six evaluation tracks, and a generalist manipulation suite. One row is marked external rather than supported, and it is the interesting one: training happens inside the project against both simulation and real data, while evaluation is delegated to a separate policy lab. Two rows are marked planned and both carry the same note that they require external environment setup. So the summary is that most suites are wired up, one is deliberately delegated to someone else, and the two that are not are the two that need a simulator you have to stand up yourself.
Twenty-nine dependencies and not one of them is open-ended
The dependency list is unusual in one specific respect. Every entry carries an upper bound, and several are tight: torch is held below 3, torchvision below 1, safetensors below 1, einops below 1, and the deepspeed entry is the narrowest of the set at above 0.18.5 and below 0.19. The rest follow the same shape, with numpy and pandas capped well short of their next majors and pyarrow given the widest span. The page recommends one exact build inside that envelope, and the command is three packages from the CUDA 12.8 wheel index:
pip install torch==2.7.1 torchvision==0.22.1 torchaudio==2.7.1 --index-url https://download.pytorch.org/whl/cu128That recommendation sits against a manifest floor of 2.5. One console script gets installed, and the development extra adds four entries, one of which is pinned to an exact version while the other three carry ranges. A C compiler is required on the path, because the deepspeed entry ships as a source distribution and compiles during install.
make test runs a subset and two gates are opt-in
The Makefile separates what runs by default from what does not, and the default is the narrower one. The plain test target deselects anything marked as a GPU test and additionally ignores one specific smoke file, so it is a strict subset of the full target, which is just the whole tests directory with no markers applied. That ignored file is the third target on its own. A compile target runs bytecode compilation across the package, the benchmarks, the scripts, the tests, and the docker directory, and a check target chains that compile step to the default test run. Lint and format both point at the same five directories through ruff, with format also applying fixes afterwards. The help target prints its own list and then delegates to a docker help target pulled in from a separate makefile, so the container recipes live in their own file rather than in the main one.
The container air-gaps itself and refuses to invent missing directories
The compose defaults are conservative in three ways and unconditional in one. Hub offline mode and the transformers offline flag both default on, and the experiment tracker is set to offline, so a freshly started container reaches for nothing on the network, which is exactly why the asset downloaders have to be run first. Both bind mounts are marked not to create the host path, so a missing output directory fails loudly instead of quietly appearing owned by root. The process runs as a numeric user defaulting to 1000 on both sides, and an init process is enabled to reap children. The unconditional part is the platform, pinned to linux/amd64 with no override, so an ARM host runs this under emulation or not at all. The image is also marked never to be pulled, so the build has to have happened on the machine first.
The service listens on every interface while the host publishes it on loopback
The serving service is where the compose design pays off. The command inside the container binds to 0.0.0.0 on port 8848, names a checkpoint directory, and pins one GPU by device index, with the GPU reserved through the container runtime rather than merely requested. The published mapping is the opposite: its host half defaults to 127.0.0.1. The bind address, the port, the GPU index, the shared memory size, the image name, and the uid and gid are all overridable by environment variable, and the defaults are the conservative one in each case. Shared memory defaults to sixteen gigabytes, which is the other number a robotics workload will feel. The checkpoint bind defaults to a path inside the released assets tree, so the default run assumes you have already downloaded the released checkpoint rather than trained one.
Editorial conclusion
OpenWAM is worth reading if you are working on robot policies and want the design space laid out rather than one architecture handed to you, because the variant table and the benchmark status list are exactly what most papers leave out. Three things to know before you install anything. The page is five weeks old: the codebase, the hosted checkpoints, and a paper all landed in the first week of September 2026, the project is at version 0.1.0, and no release has been tagged. The most valuable table on the page is commented out, so read the raw markdown rather than the rendered page. And the dependency set is capped on every entry, which is a good sign for reproducibility and a real constraint when a new release of anything lands: you will be resolving against ceilings the maintainers chose. Plan for a CUDA 12.8 PyTorch build, a C compiler on the path, and a large checkpoint download before anything runs.
Frequently asked questions
What is OpenWAM and what is in it?
It is an open research stack for World-Action Models, made of three parts: a modular infrastructure for composing and comparing model, representation, training, inference, deployment, and evaluation choices; controlled studies that derive principles for inheriting world knowledge and coupling world and action learning; and an open pretrained model trained on 518.5M frames, about 6400 hours, of egocentric human and robot data.
What does installing OpenWAM require?
Python 3.10 or newer, a CUDA 12.8 PyTorch build, and a C compiler on the path, because the deepspeed entry compiles during install. The page recommends PyTorch 2.7.1 with torchvision 0.22.1 and torchaudio 2.7.1 from the CUDA 12.8 wheel index, then an editable install of the project. A container route exists that bundles the Python environment, and still needs an NVIDIA driver and the NVIDIA Container Toolkit on the host.
Which model architectures does OpenWAM support?
Three families with variants. The single system puts video, action, and state tokens in one sequence through a shared DiT, with a mixture-of-experts variant. The dual system fuses a separate action DiT and video DiT either per layer through mixed self-attention or as a two-stage run, with a switch for whether action gradients flow back into the video side. A third family adds a frozen understanding expert to the sequence. Selection is by architecture.variant in a per-framework YAML file.
Which benchmarks does OpenWAM support?
Eight are marked supported, including a fifty-task manipulation suite, a benchmark and a perturbation suite layered over it, a household suite, a humanoid tabletop suite, a primitive-task suite across six tracks, and a generalist manipulation suite. One is marked external, where training happens inside the project and evaluation is delegated to a separate policy lab. Two are planned and both note that they require external environment setup.
Where does OpenWAM download model weights from?
From downloader scripts under scripts/download_assets, which are interactive by default but every menu step also has a flag so they can run unattended. Component downloaders store assets under assets/ and update the matching YAML path in the configuration, while the released-checkpoint downloader leaves each checkpoint's self-contained config unchanged. Five video backbones are listed as supported, and the container defaults to running in offline mode, so downloads have to happen before it starts.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/openwam-official-openwam)