Model or dataset
starVLA/starVLA avatar
starVLA/starVLA

StarVLA: the default branch is the unstable one, and the package is behind the tag

StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing

3,767 stars497 forksPythonNOASSERTION

At a glance

What is it?
StarVLA is a research platform for training and evaluating vision-language-action models on robots, structured so components can be swapped independently. Its news section is honest about what is stable, its dependency files are where the interesting detail lives, and its one performance number comes with its own disclaimer.
Who is it for?
StarVLA suits a robotics or embodied learning group that wants one codebase spanning simulation benchmarks and real robot data, and that will read the example directories rather than a single tutorial. Four things to check before you build on it.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 12 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The default branch is the development branch

Look at the branch notice near the top of the news section and there is an inversion of the usual convention. The repository's default branch, the one a clone gives you, is named with a development suffix, and the notice explains that this is where new features are merged and that it may be temporarily unstable. For verified results it tells you to use the stable branch, which is the same name without the suffix. So the path of least resistance lands you on the tree the project itself flags as the wrong one for reproducing results. The notice argues this is low risk because the design keeps components loosely coupled and that switching between branches is painless, and it invites pull requests against the development branch. That is a defensible trade for a research group that develops in public, and it is still the first thing to check when a result will not reproduce.

Three version numbers, and the package is behind the tag

Three numbers describe this project's maturity and they do not agree. The package manifest declares one point zero point one, and its development status classifier says alpha. The two releases in the list are named one point two and one point six, with the point six release described as standardising the framework on two named model families and the earlier one as standardising the benchmark suite across five named environments. So the published artefacts are ahead of the package version, which means installing the package from an index gives you something older than what the release notes describe. The tag naming is also inconsistent between the two releases: one carries a version-letter prefix and the other does not. And one of the two release descriptions contains a typo in its first word, which is the sort of thing you notice when a release title has been written once and never revisited.

The package declares no dependencies and the framework is commented out

The manifest's dependency list is present and empty. Everything the project actually needs lives in a separate requirements file at the root, and that file's most important line is a comment. The deep learning framework is not installed by it; instead there is a commented block of pinned framework builds with a specific CUDA version in one of them, alongside a note that a particular attention package may break older checkpoints and cause a performance regression on them. So the single most consequential dependency is left to the reader, pinned only in a comment, with a warning attached. The optional dependency groups are equally thin: one for development tooling and one for a specific cloud training vendor's client. There is no extra for the machine learning stack, no extra per accelerator, and no extra per robot.

Two forks of one video library, two parquet engines, three websocket libraries

Read the requirements file as a list of responsibilities rather than as packages and the overlap becomes clear. Video decoding is covered twice, once by a pinned video library and once by a pinned fork of that same library with a slightly higher version. Parquet reading is covered twice, by two entirely different engines, each pinned to a specific release. Websockets are covered three times: a client library pinned to a version, a second entry with no version at all, and a third with no version at all. Two of those websocket entries are there for one integration, since the zero-message-queue library carries a comment naming the file that needs it: a policy server for one particular third-party model family. And the tokenizer library appears twice, once pinned in the live list and once in the commented block at a different major version, so the version you run depends on which part of the file the reader looks at.

The one performance number comes with its own disclaimer

The most recent news item is the only place in this page with a measured result, and it is handled carefully. A fifty-task local snapshot of one simulated kitchen benchmark is reported for three model variants, and the best of them is given as a percentage with the raw count underneath, so a reader can see it is a fraction of a finite task set rather than a headline figure. Then the sentence that matters: these results are provided for research reference and are not official leaderboard submissions. Nobody else would add that sentence. So the number is real and attributable, and the framing tells you exactly what it is not, which means the only honest way to use it is as evidence that the integration works end to end rather than as a comparison against anything. Everything else on the page about performance is directional language.

A coming soon item dated five months before the page

The news section is dated entries in reverse order, and one of them is a promise. Dated April, it announces a unified multi-benchmark co-training example that combines several named simulation environments, and says it is coming soon with an invitation to stay tuned. The most recent entry is dated September. So a promised feature has been announced as forthcoming for five months while eight later entries have shipped around it, including two new model families, two new benchmarks and a reinforcement learning post-training integration. That is a normal thing for a research group to have in flight, and the entry is not updated to say so. It is worth knowing because a unified co-training example is exactly the kind of thing you would plan an experiment around, and the page gives you no way to tell whether it is nearly done or quietly abandoned.

The convention for your own scripts is implemented in two places

There is a tip in the news section that is really a repository convention. Any directory whose name is a particular two-syllable word is git-ignored, and the tip's point is that you can drop your own training scripts in one without them showing up as untracked files or polluting a pull request; it gives a concrete path as an example. That convention is not only a line of the ignore configuration. The same directory name appears in the package discovery configuration's exclusion list, alongside exclusions for the test, cache, playground, results, checkpoint, script, evaluation and asset directories. That exclusion list also carries a comment explaining why the playground directory is excluded: it contains symbolic links to dataset and model directories. So the packaging is deliberately refusing to walk into your data, which is a sensible thing to discover from the manifest and a sensible thing to have written.

The page asks for citation updates and ships a static citation file

Two mechanisms handle how you cite this project and they are not synchronised. There is a banner at the top of the page announcing a citation update: the technical report has moved to a preprint server, and anyone who cited an earlier version is asked to update their entry in a camera-ready copy or a future revision. There is a citation file at the repository root, which is the machine-readable form of the same thing and is what automated tooling reads. And each of the three listed research outputs, the framework itself, the evaluation work and the continued pre-training work, links to a citation anchor on the same page. So the project is asking people to refresh their citations through three separate routes while the machine-readable one is a single static file that a banner can invalidate. Given that the page also reports the framework is alpha, treat any version-specific citation as provisional.

Editorial conclusion

StarVLA suits a robotics or embodied learning group that wants one codebase spanning simulation benchmarks and real robot data, and that will read the example directories rather than a single tutorial. Four things to check before you build on it. The default branch is the development one, so a clone gives you the unstable tree and you need the stable branch explicitly for verified results. The package declares no dependencies and the framework itself is commented out of the requirements file, so installing means assembling an environment yourself. Three version numbers are in play and the package sits behind the released tag. And the one performance number in the news section is explicitly labelled as a local research result rather than a leaderboard submission, so do not quote it as one.

Frequently asked questions

What is StarVLA used for?

Training and evaluating vision-language-action models for robot manipulation. Each functional component, named on the page as model, data, trainer, config and evaluation, is separated so it can be replaced independently, which the page says enables plug-and-play use, rapid prototyping and debugging components one at a time. There is a guide for integrating your own robot or dataset, and example directories for simulation benchmarks and for real robots.

How do I install StarVLA?

There is no single install path, and that is the honest answer. The package manifest declares an empty dependency list, and the requirements file at the root omits the deep learning framework, leaving it in a commented block with a specific CUDA build pinned and a warning that an attention package may break older checkpoints. Optional dependency groups exist only for development tooling and for one cloud training vendor. Docker environments are credited to a partner organisation rather than shipped here.

Which branch should I use for StarVLA?

The stable one, and you have to ask for it by name. The repository's default branch is the development branch, which the page's own branch notice says is where new features are merged and may be temporarily unstable. For verified results the notice directs you to the stable branch. The argument for switching being painless is that the design keeps components loosely coupled, and the project invites pull requests against the development branch.

What benchmarks and hardware does StarVLA support?

The news section names simulation environments and a real-robot set, with example directories for each, and a release description that lists the benchmark suite the earlier release standardised across. On hardware, the page announces training with one model family on a non-NVIDIA accelerator, in addition to the general-purpose path, and separately announces a world-model-for-action integration that uses pretrained video generation models as backbones. Two external systems are credited or integrated for large-scale evaluation.

What results does StarVLA report?

One measured number, on a fifty-task local snapshot of one simulated kitchen benchmark across three model variants, with the best at just over forty percent and the raw task count given alongside. The page states that these are provided for research reference and are not official leaderboard submissions. There is no other performance figure anywhere in the documentation, so the rest of the claims about capability are directional rather than measured.

Official sources

  1. Issues
  2. Project website
  3. README
  4. Releases
  5. starVLA/starVLA on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/starvla-starvla.svg)](https://hysenlabs.com/projects/starvla-starvla)