AgiBot-World's build points at a licence file that is missing
[IROS 2025 Best Paper Award Finalist & IEEE TRO 2026] The Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
At a glance
- What is it?
- Training code and download instructions for a robot manipulation dataset and a foundation model, where the data, the checkpoints and the task catalogue all live on other services. Three things in the metadata disagree with each other about licensing, the upstream framework dependency is pinned to an unreleased commit, and the capability table is one task demonstrated on three different robot arms.
- Who is it for?
- Use this if you are doing bimanual manipulation research and want both a million-scale trajectory dataset and a pretrained policy in one place, because the value is that they ship together and are pinned to the same framework commit. Two things to settle before you build on it.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 129 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The build points at a licence file that is not in the tree
Three artefacts answer the question of what you may do with this code and they do not agree. The packaging metadata declares the licence as a file reference, pointing at a file named LICENSE. The top level of the repository has twelve entries and no such file: a lint configuration, a workflows directory, a gitignore, two style configurations, a pre-commit configuration, a contributing guide, the readme, an assets folder, an evaluation directory, the model directory, this manifest, and a scripts directory. And the licence field the repository reports to tooling comes back empty, so a package index or a scanner sees a project with no terms at all. The header row of the readme does link something, but it is a creative commons badge, which is where the fourth answer comes from and where the real problem is: a non-commercial share-alike deed is an instrument for a dataset, and using it as the project's licence would forbid the commercial use the model is offered for.
The badge licence is for the data and the code claims the same deed
The readme is careful about what it is offering and the licences are not. The artefacts are three: a dataset, a foundation model, and this repository of training code. The dataset is the large one, a little over a million trajectories totalling tens of terabytes, and it is the thing a non-commercial share-alike deed fits, because the researchers publishing it want academic use to spread and industry use to come to them. The model checkpoints sit on a model hub where that service's own terms apply. The code sits in this repository with a manifest that points at a missing file. So a company that reads the badge concludes the code is non-commercial, a company that reads the manifest cannot determine the code's terms at all, and the only route to a real answer is the authors. None of that makes the project unusual; large research releases routinely ship with ambiguous code licensing. It does mean the licence line is the first thing to put in writing before anyone builds on it.
The framework dependency is pinned to a commit, not a release
This project is built on someone else's robotics library, and the readme names two things about it in one line: a dataset version and a commit. The manifest is more precise, and it pins that library to a full forty character commit hash on the upstream repository's own Git host rather than to a released version. So if the upstream library changes the data format, publishes a fix, or breaks compatibility, this project neither moves nor breaks; it stays exactly where it is until someone deliberately updates the hash. That is the right call for reproducibility and it is also a real dependency on a third party's mainline, because the pinned commit is not something you can install from an index and not something with a release note attached to it. The readme states the pinned commit in the installation section as well, so a reader who has already built can see what they have. Two axes, a data version and a code commit, named separately because they move independently.
Thirty four dependencies with the four that matter pinned exactly
The runtime dependency list has thirty four entries and most of them carry no version at all, which is normal for a research repository and wrong for anything you intend to rebuild. Four carry an exact pin and those are the ones worth noticing: the deep learning framework and its vision counterpart at a matching pair of versions, the text library, and the dataset library. The array library is capped below its second major release, which is the constraint that version of the framework requires and the one that will break first on a fresh machine. The rest is a wide and reasonable spread: a model hub client, an accelerator library, parameter-efficient tuning, a web framework, a video decoder, two image libraries, two config frameworks, an experiment tracker, a formatting library, and a kernel loader that turns out to matter more than it looks. Python 3.10 is the only version claimed, and the project's own status is declared as alpha in the same manifest.
The attention kernel downloads a binary or compiles one
The installation section is the strongest documentation in this file, and it exists because the usual failure mode here is a build that dies twenty minutes in. The attention kernel is loaded through a library that is already in the dependency list, and that library fetches pre-built binaries on first use, so no local compilation is needed on a covered machine. The readme claims the fetched binaries are bit-exact with an upstream source build, which is a strong claim and a checkable one. Where a machine is not covered, the code falls back to a source build by itself, and that fallback has its own documented failure mode: running out of memory while compiling. The remedy is a named environment variable that limits parallel compilation jobs, with a concrete example. The environment is otherwise a conda environment on Python 3.10, with a stated tested accelerator toolkit version. No other part of this readme is as practically careful. The whole setup is four commands:
git clone https://github.com/OpenDriveLab/AgiBot-World.git
cd AgiBot-World
conda create -n go1 python=3.10 -y
pip install -e .The capability table is one task on three robot arms
The key features list gives the scale honestly: a little over a million trajectories collected from a hundred robots, a hundred or more real-life scenarios replicated one to one across five target domains, and the hardware involved, which includes tactile vision sensors, a six degree of freedom dexterous hand, and mobile robots with two arms. Then the capability table below it has three headings, contact-rich manipulation, long-horizon planning and multi-robot collaboration, and three demonstrations under it. All three demonstrations are the same task, folding a shirt, on three different platforms. That is a reasonable choice for a table, since it isolates the variable, but it means the three-column layout reads as breadth of capability when it is really one capability measured across hardware. The dataset is where the breadth claim lives, and the smaller curated subset is where you would look first: roughly a ninety two thousand trajectories, under a tenth of the full set by trajectory count and under a fifth by volume.
The news list stops eight months before the last push
Five dated entries, from a sample dataset at the end of 2024 through a full dataset, a technical report, and the model being open-sourced in September 2025. Then nothing, while the default branch received a push on 29 May 2026. So eight months of work on the branch appears in no news entry. That matters here more than it would in a normal project because the description string itself carries awards dated 2026 and the surrounding traffic includes queries about a 2026 dataset and a 2026 challenge, so there is activity the readme is not tracking. The other thing to notice is the roadmap section, a checklist of five items with every box ticked, covering the sample set, the full set, the model with four sub-items, and a challenge. A completed checklist in a project of this size is a leftover, and a reader looking for direction has nothing there.
Editorial conclusion
Use this if you are doing bimanual manipulation research and want both a million-scale trajectory dataset and a pretrained policy in one place, because the value is that they ship together and are pinned to the same framework commit. Two things to settle before you build on it. Licensing, first, because the packaging metadata declares a licence file that is not in the repository and the badge in the header points at a non-commercial share-alike deed, which is the right instrument for a dataset and the wrong one for software; if you intend to use this in a product, get the terms in writing from the authors rather than inferring them from a badge. Second, the dependency set, which is large and mostly unpinned, with the framework itself pinned to an unreleased commit of a third-party project and four libraries pinned to exact versions from a 2024 generation. The installation notes are unusually good on the part that usually goes wrong, the compiled attention kernel, so follow those rather than improvising.
Frequently asked questions
What does the AgiBot-World repository contain?
Training and evaluation code plus download instructions, not the data. The dataset of just over one million trajectories, the curated subset, the foundation model checkpoints, the variant without the planner, and the task catalogue are all hosted elsewhere, on a data portal and a model hub.
What licence is AgiBot-World released under?
The file does not settle it. The packaging metadata declares a licence file that is not present in the repository and the repository reports no identifiable licence, while the badge in the readme header points at a non-commercial share-alike deed, which suits the dataset rather than the code. The authors also describe the code's status as alpha.
What does AgiBot-World depend on?
A third-party robotics library pinned to a full commit hash on that project's Git host rather than to a release, with a named dataset version alongside the commit. Thirty four runtime dependencies in total, with the deep learning framework, its vision library, the text library and the dataset library pinned to exact versions and the array library capped below its second major release.
How do I install AgiBot-World?
Clone the repository, create a conda environment on Python 3.10 and activate it, then install the project in editable mode. The attention kernel is fetched as a pre-built binary through a declared dependency library, with an automatic fallback to a source build, an exact version pin for that fallback, and a named environment variable to limit compilation jobs if the build runs out of memory. The tested accelerator toolkit version is stated in the readme.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/opendrivelab-agibot-world)