opendm pins its deep learning stack exactly, floats everything else, and publishes its own reproduction gap
An Open-World Foundation Model for General-Purpose Embodied Intelligence.
At a glance
- What is it?
- An open vision-language-action model for robot control, released mid 2026 with a benchmark table against three competitors. The parts worth reading are the footnote admitting the linked recipe does not reproduce the headline number, the uneven competitor coverage across six benchmark groups, and the dependency split where the model core is pinned to exact versions and the web layer is not.
- Who is it for?
- Read this as a research release with unusually good disclosure practice, not as a product. The benchmark footnote and the per-benchmark guide links do more to make the results checkable than most robotics releases manage, and the fact that a physical robot modification guide is published suggests someone intends the numbers to be reproducible on real hardware.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The benchmark footnote admits the recipe does not reproduce the number
Immediately after a twelve-row benchmark table there is a note, and it is the most valuable sentence in the README.
It says the headline figures for one benchmark are for a specific released, fine-tuned generalist model, and that the training guide linked from the same row is a different thing: a single-task reference whose settings do not reproduce those results. The sentence is cut off in the visible text, but the claim is unambiguous.
So the number people will quote and the recipe a reader can follow are different artefacts. The leaderboard entry measures a model you can download; the guide is a starting point for training your own.
Almost no project publishes that distinction, and it changes what the benchmark means. It is evidence about a checkpoint, not about a method, and anyone planning a fine-tune from the guide should not expect to land on those numbers.
It also explains the rest of the table's structure. Each benchmark name is a link to its own training and evaluation guide, which means the page is trying to make every row checkable rather than presenting one authoritative table. The caveat is the honest edge of that effort.
One competitor is scored on two of six benchmark groups
The comparison table is wide, and reading it as a shape rather than as numbers tells you more than any single cell.
There are six benchmark groups, split into simulated tasks and real-world tasks, and three competitor models. Every row is won by the model in question, which the news section states as having swept four leaderboards.
But the coverage is uneven. The oldest competitor has figures on two of the six groups and a dash everywhere else, meaning it was either not evaluated or not available for those tasks. Two of the six groups are real-world, and one of those has only one competitor with a number.
The margins are also wildly uneven. On one benchmark the lead is two percentage points. On another it is nearly twenty. So "first on four leaderboards" is true and uninformative on its own; what matters is that the win is narrow where the task is easy and wide where it is hard, which is the pattern you would hope for but rarely get to see.
The real-world rows are the interesting ones for adoption, and they are also the sparsest, with one real-world score and one success rate against a single competitor.
The model core is pinned exactly and the plumbing floats
The dependency list is organised by comment into named groups, and the pinning convention is not uniform. It is deliberate in a way that is worth spelling out.
The deep learning core is pinned to exact versions: the framework and its vision companion, the transformer library, the accelerator library, the parameter-efficient tuning library, the quantisation library, the tokenizer, the vision backbone library, the dataset library, the array library, the diffusion library, and the validation library. A comment above the first entry notes that it must be installed first and from a specific wheel index, which tells you the resolver order matters and that the accelerator build is chosen deliberately.
Everything else floats. The web framework and its server, the two HTTP clients, the visualisation library and its client, the terminal formatting, the experiment tracker, the progress bars, the cloud storage client, the video decoders and the augmentation library all carry no version constraint at all.
The line is exactly where you would draw it: reproducibility where the model's behaviour is decided, flexibility where the plumbing does not affect it.
Two entries sit oddly inside the pinned group, though. A web microframework is pinned to a specific version with no explanation, and a filesystem abstraction is what explains why two AWS libraries are exact-pinned next to it.
The container is a GPU workstation, not a deployment image
The build file is a good example of an artefact that has outgrown its original purpose, or was written for a different one.
It starts from a developer-flavoured CUDA base image with the full toolkit rather than a runtime image, and it installs a remote shell server, a terminal editor, a terminal multiplexer and a large-file extension for version control. It also installs the InfiniBand libraries.
That combination is a developer workstation with a GPU. It is the right thing for reproducing a training run, and the wrong thing to hand to a scheduler.
The other half of the file is regional routing. Both the operating system package sources and the language package channels are rewritten to a single regional mirror, and the operating system rewrite handles two different file formats for the same configuration, which is a detail that dates the base image and hints at the platform it was authored on.
The language runtime is installed by downloading an installer by a pinned address, accepting its terms of service non-interactively, and prepending it to the path.
None of that is wrong for a research container. It is worth knowing that you are not being given a production image.
There is a guide for modifying two physical robots
One news entry has no competition from the rest of the list, and it is the one that tells you what kind of project this is.
It announces a published guide for modifying two specific physical robot platforms, documenting camera changes and a robot-name mapping that the algorithm expects to find.
Almost every robotics repository stops at simulation. Benchmark tables are simulation results, and the gap between a simulation score and a physical robot is exactly where a policy quietly stops working: the camera is mounted differently, the arm is named differently, the coordinate frame is off by a sign.
Publishing the hardware modification is a statement that the authors intend the numbers to be reproducible on the real platform, and that they know the software alone will not reproduce them.
The rest of the news follows a similar pattern of shipping artefacts rather than claims. A fine-tuned checkpoint for a specific bimanual manipulation task on named hardware. A generalist checkpoint from another simulator. A pick-and-place fine-tune with a documented training workflow. Each is a model you can download rather than a number you have to trust.
There are three version axes and not one tag
The repository has no releases. None. And it contains three different things that a reader might reasonably mistake for a version number.
The first is the model itself, referred to throughout as a point five, and described as the next generation of an earlier one called point zero.
The second is the package, whose manifest declares version zero point one point zero. So the package you install and the model you download are not the same version series, and neither carries a tag.
The third is the checkpoints, which are published as their own named models: one for a bimanual task on specific hardware, one generalist for another simulator, one fine-tuned for a pick-and-place task on a third platform.
For a repository whose entire purpose is that someone else reproduce a result, that is the sharpest operational gap here. There is nothing to pin. A user cannot write a requirements file that says which of these they mean without also saying which checkpoint name they downloaded.
The news entries do carry dates and links to guides, so the information exists; it is simply not expressed as a version you can depend on.
The policy went into somebody else's repository
The newest news entry describes work that landed outside this repository, which is worth understanding before you clone anything.
It says the policy was open-sourced in a separately run research lab, and the entry enumerates what went with it: a data conversion step for one benchmark, a memory-focused fine-tune, and a simulation evaluation of the released model.
Earlier entries reference that same external repository twice: once as the home of a fine-tuned model released here, and once as the place where a pull request handles evaluation integration.
So the split is roughly this: the model, the training and inference scripts, the dataset examples and the evaluation workflows live here, while the policy implementation and part of the evaluation harness live there.
That is a normal shape for a research group that maintains shared infrastructure, and it has a practical consequence. Reproducing the released checkpoints means following the model repository and then following the lab, and the two are versioned separately if at all.
For anyone auditing what they would actually be running, the policy code is the part you cannot get from a single download here.
Editorial conclusion
Read this as a research release with unusually good disclosure practice, not as a product. The benchmark footnote and the per-benchmark guide links do more to make the results checkable than most robotics releases manage, and the fact that a physical robot modification guide is published suggests someone intends the numbers to be reproducible on real hardware. Two cautions before you plan around it. The repository has no releases and three different version axes, so pin a specific checkpoint rather than a branch. And the reproduction caveat applies to the benchmark everyone will quote, so treat that headline as a claim about a model you can download rather than about a recipe you can rerun.
Frequently asked questions
what is open dm
It is an open-world foundation model for embodied intelligence: a vision-language-action model for open-world robot control, building on an earlier generation with systematic upgrades for open-ended instructions, long-horizon tasks, dynamic disturbances and multi-embodiment control. The repository ships model weights, training and inference scripts, dataset registration examples and evaluation workflows.
Does OpenDM publish releases?
No, the repository has no releases at all. The package version in the manifest is 0.1.0 while the model is referred to as a point five, and separately fine-tuned checkpoints are published as their own named models, so there are three version axes and nothing to pin.
Can I reproduce OpenDM's RoboDojo-Sim benchmark numbers?
Not from the linked guide, and the page says so directly. The leaderboard figures are for a specific released fine-tuned generalist model, while the linked training guide is a single-task reference whose settings do not reproduce them.
What does OpenDM compare itself against?
Three models, across six benchmark groups split into simulated and real-world tasks. One of the three has figures on only two of those groups and dashes elsewhere, and the winning margins range from about two percentage points on one benchmark to nearly twenty on another.
What hardware does OpenDM target?
The container is built from a CUDA developer base image with a deep learning runtime image, and the manifest installs the core framework first from a specific wheel index. The repository also publishes a guide for modifying two specific physical robot platforms, documenting camera changes and the robot-name mapping the algorithm expects.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/dexmal-opendm)