# RoboTwin: the only release tag is the word release, and the deps are not in the repository

> A bimanual robotic manipulation benchmark whose current version lives on the default branch rather than in a version tag, whose evaluation harness is a pinned git submodule in a third repository, and whose own page tells you twice to prefer external documentation. Over a hundred thousand pre-collected trajectories ship with it, and the top level contains no dependency manifest of any kind.

**RoboTwin-Platform/RoboTwin** — [ICML 2026] RoboTwin 2.0 Offical Code Repo

- Repository: https://github.com/RoboTwin-Platform/RoboTwin
- Website: https://robotwin-platform.github.io
- Stars: 2,938 · Forks: 505
- Language: Python
- License: MIT
- Published: 2026-09-24 · Updated: 2026-09-24 · Language: en
- Canonical page: https://hysenlabs.com/projects/robotwin-platform-robotwin

## The only release tag is the word release

There is exactly one release, and its tag is not a version number.

It is the literal string `release`, titled Stable Version, published 2026-02-26. That is the entire release history.

So there is no artefact you can pin to by version. Any tool that asks the repository for a described version gets back the word release, and two different states of the code are indistinguishable by tag.

The version is asserted everywhere except in the one place it would be verifiable. The page header says Latest Version: RoboTwin 2.0. The repository description names version 2.0 as well, and misspells the word official in the process, which is the text search engines and package indexes will mirror. The overview states that the default branch is the 2.0 branch.

The timing compounds it. The lone tag is dated 2026-02-26 and the last push to the default branch is 2026-09-24, so roughly seven months of commits sit past the only reference point in the repository. The update log's final entry is 2026/08/03, so the changelog has not kept pace with the branch either.

What you get from a clone is whatever main is. That is a workable answer for a research project and a poor one for anything that needs to name the version it ran.

## Ten branches, and one sentence says which is current

The default branch is 2.0, and the page says so in a single line. Behind it sit nine more.

Listed as legacy or special-purpose are an arena branch named after a different simulator, a branch for support in an external reinforcement-learning framework, a branch named for a 2026 benchmark, the 1.0 branch alongside an early-version branch, a branch named for a GPT-related effort, a 2025 challenge cup branch, and two rounds of a 2025 conference challenge.

That is the current benchmark plus at least three prior challenge variants, two framework integrations and two historical versions, all in one repository, distinguished only by branch name.

Two of the branches exist because other people contributed them. The reinforcement-learning support branch is credited in the update log to the team that contributed it, and the arena branch is named for a simulator the project does not own. That is healthy for adoption and it is exactly why version identification matters here.

The practical consequence for anyone trying to reproduce a number: a clone of the default branch gives one specific configuration, and a leaderboard entry may depend on a different branch. With no version tags to fall back on, the branch name is the only handle available.

The paper list has the same shape. Four identifiers cover the early version, the 1.0 release, the 2.0 release and a challenge report, and the page is careful to record which venues gave which distinctions.

## The evaluation harness is a pinned submodule in a third repository

The most consequential line in the installation section is about a dependency rather than about robots.

The overview says 2.0 and a second benchmark share their deployment through one policy-serving and evaluation stack, living in its own repository, supporting single-task evaluation, multi-GPU scheduling, and a split deployment with a remote policy server and a local simulator.

That stack is embedded as a git submodule, and three commands manage it. A fresh checkout clones recursively. An existing checkout initialises the version the project has pinned. And then there is the third:

```bash
bash scripts/update_xpolicylab.sh
# optional: stage the pin / reinstall the editable package
bash scripts/update_xpolicylab.sh --stage --install
```

The page describes that script as pulling the latest commit on the submodule's configured main branch and refreshing RoboTwin's pin.

So two dependency postures are available and the unpinned one is one flag away. The pinned path is reproducible. The update path deliberately moves you to the tip of another project's main branch.

The optional reinstall of an editable package explains how the stack is consumed: not vendored source, but a package installed in development mode from inside the submodule. Which means the evaluation code you run is whichever commit the submodule points at.

For a benchmark whose entire value is comparability across submissions, that is the line to read twice.

## There is no dependency manifest anywhere in the repository

The top level has fourteen entries and not one of them declares a Python dependency. There is no requirements file, no project file, no setup script and no environment file.

The installation section says where the instructions are instead, on the documentation site, and adds that installation takes about twenty minutes.

So the twenty minutes is spent following a document that is not in the repository, and the only commands this page gives are three git and shell invocations concerning the submodule. A fork cannot be set up from the repository's own contents.

That is a deliberate division rather than an oversight, and the page is explicit about it twice: the overview says to prefer the documentation for full guides because the README is a quick start, and the usage section says to prefer the documentation over this README whenever anything conflicts. Two independent statements that the page in front of you is not authoritative.

The one executable entry point in the tree is a shell script for collecting data. So the data-collection path has a single-line entry and the training path has none here.

Two directories sit side by side for the environment, one holding configurations and one holding implementations. The split is conventional and worth naming because it means a task's parameters and its behaviour are edited in two separate places.

## The changelog admits two evaluation-code bugs and then goes quiet

The update log is the most informative part of the page, and it runs from late 2024 to August 2026 with seventeen entries.

Two of them are corrections to evaluation code rather than to the environment. One fixes the evaluation code for a policy method in July 2025. Another fixes deployment code for a different method in August 2025. For a benchmark these are a different class of bug from a crash, because an error in the harness changes scores rather than stopping runs, and any result published before the fix is suspect.

A third entry, from July 2025, fixes a wrist bug and asks users to redownload an embodiment asset. So that correction required replacing data people had already fetched.

The shape of the log is worth noting too. Three of the seventeen entries are paper acceptances and awards rather than software changes. There is a five-month gap between the March 2026 entry and the August 2026 one, and nothing after August despite pushes into September. And one entry from July 2025 promises a paper update the following week; the next paper-related entry is five weeks later.

Set against the fact that there is one unversioned tag, the picture is of a project whose identity is carried by its branch and its paper rather than by its release artefacts.

It is also more forthcoming about measurement instability than most benchmark pages are.

## Two ways to get data, and the page recommends one of them

The data section offers a choice, and states a preference.

Over 100,000 pre-collected trajectories are released as an open-source dataset, hosted on a model hub. The alternative is to collect your own. The page recommends downloading, calling it the default path because it is ready to train on immediately, and says to collect yourself only when you need custom task configurations, domain randomisation or different embodiment setups.

So the benchmark ships both the environment and the data generated inside it, which is unusual and genuinely useful. It also means the dataset, not the code, may be the thing most users actually depend on, and it lives on a model hub under a named individual's account rather than in the repository.

Immediately after that comes a hard interoperability rule, stated as an absolute: always decode through one named function, and always encode through another. The sentence breaks off partway through its second clause.

That instruction is easy to skim and expensive to get wrong. Image data that must pass through one specific codec on the way in and a different one on the way out will not fail loudly if you mix it with ordinary image handling. The page offers no rationale and no error message for the mistake, which is the opposite of what you would want from a rule phrased that strongly.

A second section of the page is simply a link to the task documentation followed by an empty paragraph element, so whatever belonged there renders as nothing.

## Two sites, a typo in the description, and a bare video link

The presentation details are loose, and they cluster in one direction: this is a changelog and a link hub rather than documentation.

The header title links to a benchmark site on one domain, while every other link on the page, and the repository's own homepage metadata, point at a different domain. Two sites, one project, and nothing on the page explaining the split.

The repository description misspells the word official in its opening, and that string is what appears wherever the repository is indexed or listed.

Between the title and the version line sits a bare video URL with no label, no alt text and no sentence around it, hosted on a user-images path rather than anywhere inside the repository.

None of this changes what the project is or does. Together it is the visible cost of a page maintained as a running log, which is consistent with its own instruction to prefer the external documentation when anything conflicts.

The paper trail is thorough by comparison, and it is where the care went. Four identifiers, a conference acceptance, a highlight distinction, a workshop best paper, an outstanding poster with a stated ranking, a challenge technical report, and links to two Chinese-language technology publications covering that report.

## Conclusion

RoboTwin suits a research group that wants a dual-arm manipulation benchmark with published data and a leaderboard rather than building an environment from scratch, and the willingness to follow documentation off-repository is part of the deal. Three things to settle before you depend on it. Record which commit you used, since the branch name is the only version handle available and nine other branches carry prior versions and challenge variants. Leave the evaluation submodule at its pinned commit, because the provided script will otherwise move it to the tip of another project's main branch and quietly change your measurement stack. And treat the published trajectories as the fast path, while reading the codec requirement carefully, since mixing encoded and ordinary image handling fails quietly rather than loudly.

## FAQ

### What is RoboTwin?

RoboTwin is a bimanual robotic manipulation platform. The current version is RoboTwin 2.0, published as an ICML 2026 paper, with earlier work published as a CVPR 2025 highlight and an ECCV 2024 workshop best paper. A leaderboard and documentation are hosted alongside the project webpage.

### How do I install RoboTwin 2.0?

The README carries no dependency manifest and no install command beyond the git clone. Installation instructions are on the project documentation site and are said to take about 20 minutes. The evaluation stack is a submodule, so a fresh checkout must clone recursively.

### Does RoboTwin provide pre-collected training data?

Yes. Over 100,000 pre-collected trajectories are released as an open-source dataset, and the page recommends downloading them as the default path since they are ready to train on immediately. Collecting data yourself is only needed for custom task configurations or different embodiments.

### Which RoboTwin version is on the default branch?

The README states that the default branch main is RoboTwin 2.0. Nine other branches cover earlier versions, challenge rounds and framework integrations, so the branch name is the only version handle, because the repository's single release tag is the literal word release with no version number.

### What evaluation stack does RoboTwin 2.0 use?

It shares a policy-serving and evaluation stack with another benchmark, hosted in its own repository and embedded as a git submodule, supporting single-task evaluation, multi-GPU scheduling and a remote policy server with a local simulator. A provided script can move the pin to the latest commit on the submodule's main branch.

## Sources

- [License: MIT](https://github.com/RoboTwin-Platform/RoboTwin/blob/main/LICENSE)
- [Project website](https://robotwin-platform.github.io)
- [README](https://github.com/RoboTwin-Platform/RoboTwin/blob/main/README.md)
- [Releases](https://github.com/RoboTwin-Platform/RoboTwin/releases)
- [RoboTwin-Platform/RoboTwin on GitHub](https://github.com/RoboTwin-Platform/RoboTwin)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/robotwin-platform-robotwin
