Open-source project
physical-superintelligence-lab/Psi0 avatar
physical-superintelligence-lab/Psi0

Psi0 keeps its dependencies in the image and your checkout outside it

[RSS'26] Welcome to Psi-Zero, a Humanoid VLA towards Universal Humanoid Intelligence.

2,862 stars106 forksPythonNOASSERTION

At a glance

What is it?
Psi0 is an open vision-language-action model for dexterous humanoid loco-manipulation, built as a Qwen3-VL-2B backbone above a diffusion action expert with an RL tracking controller underneath. The engineering surface is where the detail is: a wall of exact version pins, a lerobot dependency taken from a fork at a fixed commit, a container that carries no code at all, and a server that binds every interface on the host network.
Who is it for?
Psi0 is for a lab that has a humanoid, teleoperation data and a GPU cluster, not for a first contact with humanoid policies: the quick start ends at a version string, and everything after that is fine-tuning scripts, deployment guides and a container recipe. Three things to check before you commit a week to it.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Two trained components above an RL controller

The architecture is described in three named layers. At the top is a vision-language backbone, labelled System-2, based on Qwen's Qwen3-VL-2B-Instruct, which extracts vision-language features from the observations and the instruction. Below it sits a multimodal diffusion transformer action expert, labelled System-1, which the page says is inspired by Stable Diffusion 3 and is flow-based, of roughly 500 million parameters, and which predicts future whole-body action chunks so that visual, linguistic and action representations can be fused. At the bottom is System-0, an RL-based tracking controller that executes the predicted lower-body action commands and is responsible for stable and precise physical control.

The two upper components are trained end to end together. The third is a conventional control layer beneath them.

The data story is in the same paragraph as the architecture. The model first learns task semantics and visual representation from large-scale human egocentric video, and is then post-trained on a much smaller amount of real-world teleoperated robot data in order to learn the general dynamics of the embodiment. The claimed result is that a new long-horizon loco-manipulation skill can be acquired by fine-tuning from as few as 80 trajectories, with the stated key finding being that scaling the right data in the right way is what matters.

The dependency wall is exact pins, and lerobot comes from a fork

The project manifest pins the interpreter to one minor series: requires-python is set to exactly 3.11, matching the virtual environment the quick start creates with a 3.11 flag.

From there almost every dependency is an exact version. The core list includes albumentations at 1.4.18, datasets at 3.6.0, dm-tree at 0.1.8, einops at 0.8.1, h5py at 3.14.0, imageio at 2.34.2, numpydantic at 1.6.7, plotly at 6.2.0, pydantic at 2.10.6 and tqdm at 4.67.1, with numpy held below 2.0.0 and opencv-python left unpinned. The training group is the same story: torch at 2.7.0, torchvision at 0.22.0, torchcodec at 0.4.0, transformers at 4.57.1, deepspeed at 0.17.1, triton at 3.3.0, accelerate at 1.7.0 and peft at 0.17.1.

The entry that stands out is the last one in that group. The lerobot dependency is not a released version at all: it is a git reference, github.com/songlin/lerobot.git at the commit 09929d8057b044b53aecaf5c6d7eb71f99e8beb9. Songlin Wei is the first name in the contributor list, so the shared robot data stack this project depends on is taken from a contributor's fork at a fixed commit. That is reproducible and it is also a dependency on one person's branch.

The documented clone command needs an SSH key

The installation section starts by cloning the project and changing into it:

bash
git clone [email protected]:physical-superintelligence-lab/Psi0.git
cd Psi0

That address is the SSH form, so the command as written needs a GitHub key configured before it will run. A reader who arrived without one has to translate it to an HTTPS URL themselves.

The environment is then built with uv rather than pip. The instructions are to install uv if it is not already there, using a shell command that pipes a remote script into sh, and then to create the virtual environment and sync three dependency groups:

code
uv venv .venv-psi --python 3.11
source .venv-psi/bin/activate
GIT_LFS_SKIP_SMUDGE=1 uv sync \
  --group serve \
  --group viz \
  --group psi \
  --index-strategy unsafe-best-match \
  --active
uv pip install flash_attn==2.7.4.post1 --no-build-isolation

Four details in that block are worth pausing on. The large file storage skip flag means git submodules are not smudged during the sync, and the repository does carry submodule configuration. The index strategy is named unsafe-best-match. flash-attn is deliberately outside uv sync, installed separately at a fixed version with build isolation disabled, which is the usual requirement for that package. And the three groups are the runtime split: serve, viz and psi.

The container holds dependencies, not code

The Docker path has one non-obvious step in it. You pull the prebuilt image and then retag it locally:

bash
docker pull ghcr.io/physical-superintelligence-lab/psi0:latest
docker tag ghcr.io/physical-superintelligence-lab/psi0:latest psi:train

The page explains why the retag is not optional: the compose file refers to the image as psi:${PSI_TAG:-train}, and the tag variable overrides only the tag, not the registry path. The same paragraph advises using a dated tag such as 260830 to pin an exact build, which sits awkwardly next to a pull command that ends in the floating latest tag.

What is inside the image is stated just as plainly: dependencies only. Python 3.11, torch 2.7 for CUDA 12 and flash-attn, installed into a virtual environment inside the container, on a plain CUDA base image. Your checkout stays outside the image and is bind-mounted, so the project is installed in editable form and source edits take effect on a container restart rather than a rebuild.

The compose file's own comments repeat the rule from the other side: one service owns the image and builds it, the serving services consume it and never rebuild, and a third gives an interactive shell in the same environment. That design is why editing policy code does not mean rebuilding a multi-gpu image.

The server runs on the host network and binds all interfaces

The compose file sets network mode to host for the shared runtime, and its comment says why: the server binds port 8014 on all interfaces on the host directly, so a client on the same machine reaches it at the loopback address, including a simulation component that is itself host-networked.

That is a deliberate design decision with a consequence worth naming. Host networking removes the network namespace, and binding all interfaces rather than loopback means anything that can reach the machine on that port can reach the server. The comment frames the exposure as a convenience for co-located evaluation, and nothing in the compose file adds authentication in front of it.

The rest of the runtime block is conventional and mostly overridable. The image reference is the tag variable again, the runtime is the NVIDIA one, the working directory is a workspace path, an env file supplies values, and the environment sets the visible devices from a variable defaulting to zero, the driver capabilities, and unbuffered Python output. Volumes bind the source directory and the runs directory, where checkpoints land, into the container. The README adds that every argument can be overridden from the shell or the env file, naming the port, the run, the checkpoint step, the action execution horizon, an RTC flag and the GPU selection, and that running the serve service with a help argument lists the server's own options.

The sample environment targets one compute capability

The sample env file is more informative than most, because it carries a commented table mapping CUDA compute capability to specific hardware, from 7.5 covering a T4, an RTX 2080 and a Quadro RTX, through 8.0 for an A100 and an A30, 8.6 for an A40, an A10, an RTX 3090 and an A6000, 8.9 for an L40S, an L4 and an RTX 4090, 9.0 for an H100 and an H200, and 10.0 or 12.0 for a B200 and an RTX 5090.

Against that table the sample sets one value: the torch CUDA architecture list is set to 8.0 with PTX. That is a floor with forward compatibility rather than a set covering the range above it. A T4 at 7.5 is below the list, and the newest parts at 10.0 and 12.0 are not in it either, so a build from this sample targets one generation and relies on PTX for anything newer.

Two other sample values are debugging or hygiene settings left switched on. CUDA launch blocking is true, which serialises kernel launches and slows execution down, and the deepspeed log level is warning. There is also a commented proxy block and a commented mirror endpoint override for the model hub, which tells you what to change when the default host is unreachable.

Seven baselines and three environment systems

The contents list seven baselines under their own headings: GR00T N1.6, OpenPi at version 0.5, InternVLA-M1, H-RDT, EgoVLA, Diffusion Policy and ACT. That is a comparison set spanning published large policies, an earlier architecture and a plain behaviour cloning baseline.

The repository then carries three environment systems at once. There is a pyproject with a uv workspace and a lock file, a flake with a Nix directory beside it, and a docker compose file with an env sample. The compose comments also mention an enroot and cluster recipe in the training scripts directory, which is the container tooling the same scripts use outside Docker.

Underneath those are the directories a user actually navigates: a baselines directory, a real directory holding the real-world deployment guide, scripts with training recipes, source, tests, examples, assets, and a third_party directory that pairs with the submodule configuration. The examples tree includes notebooks for a data configuration and for the vision backbone, directories for deployment, the model itself, a quick start, simulation and training, and two markdown guides, one on adding a new embodiment and one on the checkpoint released for the SONIC integration.

One licence caveat belongs here rather than in a footnote. A licence file sits at the repository root, while the licence attached to the repository is not asserted, so the two do not agree on their face and the file is the one to read.

The table of contents keeps three disabled entries

The first three lines of the contents list are commented out: installation, pre- and post-training, and data pre-processing. What remains starts at fine-tuning on the Unitree G1 humanoid robot and runs through installation, data collection, fine-tuning, open-loop evaluation, deployment and the SONIC section.

Those three disabled entries correspond to real material. There is a reproduce section for pre-training and post-training, there is a real-world deployment guide under the real directory, and the news list records that all nine real-world tasks were open sourced with the data available for download so a user can skip collection and go straight to fine-tuning. A deprecated block for an earlier integration with AMO also sits in a collapsed section.

The news list itself is dated and specific: checkpoints trained for SONIC with two training recipes and a release node on 2026-09-13, a whole-body training recipe and Docker support on 2026-08-30, a DreamZero baseline for the simulator on 2026-07-14, the SONIC integration on 2026-06-13, and a best paper award at the second 3D-LLM and VLA workshop at CVPR 2026 on 2026-06-03. There are no GitHub releases, so those dates are the only release record.

Editorial conclusion

Psi0 is for a lab that has a humanoid, teleoperation data and a GPU cluster, not for a first contact with humanoid policies: the quick start ends at a version string, and everything after that is fine-tuning scripts, deployment guides and a container recipe. Three things to check before you commit a week to it. Read the licence yourself, because the repository root carries a licence file while the licence attached to the repository is not asserted. Check that your GPU is in the arch list the sample environment ships, since the sample targets one compute capability with forward compatibility rather than the range its own table lists. And decide how you will expose the serving container, because it runs on the host network and binds all interfaces on its port.

Frequently asked questions

What is Psi0?

It is an open vision-language-action model for dexterous humanoid loco-manipulation. It learns task semantics and visual representation from large-scale human egocentric video, then post-trains on a smaller amount of real-world teleoperated robot data, and is built as a Qwen3-VL-2B-Instruct backbone above a flow-based diffusion action expert of about 500M parameters, with an RL tracking controller below.

How do I install Psi0?

Clone the repository, install uv, then create a Python 3.11 virtual environment named .venv-psi, sync the serve, viz and psi groups with uv sync, and install flash_attn at 2.7.4.post1 separately with build isolation disabled. The check is a one-line python import that prints the psi version.

Can I run Psi0 in Docker?

Yes, by pulling the prebuilt image from the project's GitHub Container Registry and retagging it to psi:train, because the compose file refers to psi:${PSI_TAG:-train}. The image carries dependencies only, so your checkout is bind-mounted and source edits need a container restart rather than a rebuild.

Does Psi0 depend on a fork of lerobot?

Its dependency list pins lerobot to a git commit on a fork rather than to a released version, at commit 09929d8057b044b53aecaf5c6d7eb71f99e8beb9 of the repository belonging to the first listed contributor. Nearly every other dependency is pinned to an exact released version.

Which baselines does Psi0 document?

Seven: GR00T N1.6, OpenPi at version 0.5, InternVLA-M1, H-RDT, EgoVLA, Diffusion Policy and ACT. Each has its own heading in the contents, alongside sections for fine-tuning on the Unitree G1, simulation, checkpoints and troubleshooting.

What licence does Psi0 use?

The repository root carries a licence file, while the licence attached to the repository itself is not asserted, so the two do not agree on their face and the file in the repository is the one to read.

Official sources

  1. Issues
  2. physical-superintelligence-lab/Psi0 on GitHub
  3. Project website
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/physical-superintelligence-lab-psi0.svg)](https://hysenlabs.com/projects/physical-superintelligence-lab-psi0)