MIT in the metadata, Apache in the classifier, and transformers installed last
Dexbotic: Open-Source Vision-Language-Action Toolbox
At a glance
- What is it?
- dexbotic is a PyTorch toolbox for vision-language-action models covering pretraining, fine-tuning, inference and evaluation. Its own packaging disagrees with its licence metadata, ships seven linting tools as runtime dependencies, and builds on a CUDA 11.8 image whose every package index points at the same mirror.
- Who is it for?
- dexbotic is worth adopting if you are already training vision-language-action policies and want one environment rather than five, since the container pins the toolkit versions and the project ships benchmark scripts for each model it claims to support. Four things to check before you rely on it.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 19 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The licence metadata and the packaging classifier name different licences
Two licence signals exist in this repository and they do not agree. The licence carried with the project is MIT, and the README links a LICENSE file. The packaging configuration declares a classifier of License :: OSI Approved :: Apache Software License, which is a different licence family. Nothing else in the repository settles it, and there is no NOTICE file or dual-licence statement explaining a choice.
That matters more than it does for a small library, because what is being licensed is a training environment: pinned versions of torch, DeepSpeed and the video decoding stack, plus pretrained models, benchmark recipes and hardware integration guides. If you intend to redistribute an image built from this repository, the two sentences in the metadata and the classifier are the only statements available, so read the LICENSE file directly rather than inferring from either. The project description also differs between the two places it appears, one calling it an open-source vision-language-action toolbox and the packaging calling it training and serving a large language action model.
Seven linting tools are runtime dependencies, alongside three web frameworks
The dependency list is organised by comment blocks, which makes its shape easy to read, and one of those blocks should not be there. Code formatting and linting is listed under runtime dependencies rather than as an optional group: autopep8, pycodestyle, black, flake8, isort, mypy and the type stubs for requests. Installing the package therefore pulls in seven development tools, and there is no extras group visible in the configuration to separate them.
The web layer is equally generous. fastapi and uvicorn are both present, and so is flask, which means three serving approaches are declared rather than one. The data side is heavier: decord and av for video decoding, albumentations for augmentation, diffusers, datasets, and a fixed numpy. Cloud access is pinned exactly, with botocore and boto3 at fixed versions, which is what a training pipeline reading objects from storage needs. And one dependency carries a comment explaining an upstream incompatibility rather than a preference.
kernels is capped from below because of a transformers constructor change
One pin in that list exists because of a specific break, and the comment above it says so in two lines. The note records that transformers 5.3 instantiates repositories without a revision or version, and that kernels at or above 0.15 rejects that constructor form during import. The fix chosen is an upper bound on a different package, kernels held below 0.15, which stops the import from failing.
Two details follow from that comment. The declared dependency for transformers is a floor of 4.57.6, while the comment describes behaviour specific to 5.3, so the manifest allows a range the workaround was written against. And the constraint is one-directional: nothing stops a resolution that picks a transformers older than the one the comment assumes. It is a narrow, well-documented pin, and it is the kind of entry that tells you the maintainers track upstream breakage rather than guessing at versions.
The base image is CUDA 11.8 and every index points at one mirror
The Dockerfile starts from an NVIDIA CUDA 11.8 image with cuDNN 8 on Ubuntu 20.04, in the devel variant so it can build. It then rewrites the operating system's package sources, replacing both the archive and the security hosts with a Tsinghua mirror, sets the timezone to Asia/Shanghai, and installs build tools plus a handful of command line programs.
The same mirror appears twice more. Conda is configured to read its main and free channels from Tsinghua instead of the default host, and pip's global index URL is set to that mirror's package page. Miniconda is fetched from the Anaconda repository, and the terms of service for two conda channels are accepted non-interactively during the build, which is a step that only exists because the package host now gates downloads that way.
So the image is reproducible in the sense that every source is named, and opinionated in the sense that all of them are one organisation's mirrors. Anyone behind a network that cannot reach them, or anyone who wants an auditable list of upstream sources, has to rewrite the Dockerfile before building.
transformers is installed after the package, and torch is pinned tighter than the manifest
The dependency install order in the image is deliberate and it matters. A conda environment named dexbotic is created with Python 3.10, then torch, torchvision and xformers are installed at fixed versions from the CUDA 11.8 wheel index, then the project itself is installed in editable mode, and then transformers is installed at a fixed version afterwards. That last step overrides whatever the editable install resolved, so the effective version of the library the project is known to work with comes from the Dockerfile rather than from the manifest.
The gap between the two sources is easy to miss. The manifest asks for torch at 2.6.0 or newer and transformers at 4.57.6 or newer, while the image installs torch 2.6.2 and transformers 4.57.6 exactly. xformers is installed by the image and does not appear in the manifest at all. A developer who sets up the environment by hand from the manifest alone can therefore land on a combination the maintainers have not pinned, which is the usual reason a project of this shape insists on the container.
The documented mount points at the package folder, then installs from it as a project root
The quick start clones the repository, starts the published image with all GPUs passed through and the host network, and mounts a directory from the current path into the container.
# 1. Clone the repository
git clone https://github.com/dexmal/dexbotic.git
# 2. Start Docker container
docker run -it --rm --gpus all --network host \
-v $(pwd)/dexbotic:/dexbotic \
dexmal/dexbotic \
bash
# 3. Activate environment and install dependencies
cd /dexbotic
conda activate dexbotic
pip install -e .The mount source is the dexbotic subdirectory of the clone, not the clone root, and the destination is /dexbotic. The next steps are to change into that destination, activate the conda environment and run an editable install of the project.
Those two things do not line up. An editable install needs a project root, which is where pyproject.toml lives, and in the repository that file is beside the dexbotic directory rather than inside it. Following the commands exactly as printed means installing from a directory that holds the package folder and the Dockerfile but not the packaging metadata, so the install step has nothing to read. Fixing it is a one-word change to the mount source, and it is the kind of detail worth finding out before an afternoon of GPU time is spent on it.
One performance figure appears twice, for two different models
The project summary contains exactly one speed claim, an approximately fivefold model-inference speedup, and it is stated twice. The first time describes a realtime inference guide for DM0 backed by Triton, published in June 2026. The second describes an upgrade to DM05 inference on a high-performance backend, published in September 2026. Two different models, two different backends, the same rounded multiplier.
Neither entry gives the hardware, the batch size, the sequence length or the measurement method, so the number cannot be checked from the repository, and a fivefold figure on inference is the kind of claim that depends on all four. What the news list does supply is a shape worth reading on its own. Twenty-seven entries span eleven months, from the first release in October 2025 to history-aware inference for DM05 in September 2026, and the pattern is consistent: a new model or backend arrives with a guide in docs, a script in the playground, and for the newest models a LIBERO recipe for full fine-tuning and one for LoRA.
Model names change spelling between the summary and the news entries
The same policy appears under several spellings, which makes searching the repository harder than it should be. The summary lists pi-0 by its Greek letter alongside CogACT, OFT and MemVLA. The news entries refer to PI0 and PI05 in the July LoRA entries, to Pi05 in the March co-training entry, and to Pi0.5 in the December support entry. The newest model is written DM05 in the news and DM0.5 in the address of its own technical report, with DM0 as the earlier version.
Alongside that, the summary describes its own architecture in three words that turn out to be literal: layered configuration, factory registration and entry dispatch, with the claim that you change behaviour by editing experimental scripts. The scripts live in the playground directory and are named after the model they run, so the mechanism is visible in the file names. Hardware arrives the same way. The summary names UR5, Franka and ALOHA with a unified training data format, and the hardware directory carries per-vendor guides for a Unitree G1 SONIC setup, a DOS-W1 integration, XLeRobot and an SO-101 arm.
Editorial conclusion
dexbotic is worth adopting if you are already training vision-language-action policies and want one environment rather than five, since the container pins the toolkit versions and the project ships benchmark scripts for each model it claims to support. Four things to check before you rely on it. Read the LICENSE file, because the metadata and the packaging classifier name two different licences. Expect your package index to be redirected, since the Dockerfile rewrites the apt sources, the conda channels and the pip index to one mirror, which matters if you are behind a firewall or want an auditable source list. Note that the image pins torch and transformers to exact versions the manifest does not, so an editable install outside the container will not match it. And keep the licence question separate from the technical one: this is a fast moving research toolbox with a new model release every few weeks and two release tags in total.
Frequently asked questions
What is dexbotic used for?
It is a PyTorch toolbox for vision-language-action models, covering pretraining, fine-tuning, inference and evaluation. It ships ready-made environment configurations for mainstream policies so a model can be reproduced or fine-tuned with a short setup, and the summary names pi-0, CogACT, OFT and MemVLA among the policies it supports.
Does dexbotic need GPUs to run?
Yes. The quick start passes all GPUs through to the container, and the stated system requirements are Ubuntu 20.04 or 22.04 with an RTX 4090, A100 or H100 recommended, eight GPUs for training and one for deployment. Blackwell cards such as the B100 or RTX 5090 need a separate image tag.
Which robot hardware does dexbotic support?
The summary names UR5, Franka and ALOHA, with a unified training data format and deployment scripts. The hardware directory adds per-vendor guides for a Unitree G1 SONIC setup, DOS-W1, XLeRobot and an SO-101 arm, each shipped alongside an example script.
Which licence does dexbotic use?
The two sources in the repository disagree and neither settles it. The project metadata and the README badge say MIT, while the packaging classifier declares the Apache Software License. Read the LICENSE file at the repository root before relying on either, particularly if you intend to redistribute a built image.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/dexmal-dexbotic)