Library / SDK
lucidrains/pi-zero-pytorch avatar
lucidrains/pi-zero-pytorch

pi-zero-pytorch: the π₀ robot policy architecture in a pip install

Implementation of π₀, the robotic foundation model architecture proposed by Physical Intelligence

587 stars29 forksPythonMIT

At a glance

What is it?
lucidrains' pi-zero-pytorch reimplements the π₀ vision-language-action model as a small PyTorch package with flow matching and joint attention. It is a training and research artifact, not a robot deployment kit, and the official openpi repository now exists alongside it.
Who is it for?
Adopt pi-zero-pytorch if you want to read, modify or pretrain the π₀ architecture in plain PyTorch and you already have a PaliGemma-scale VLM and a robot dataset to point it at. Do not adopt it if you need a working robot policy next week: there are no published checkpoints, no inference server for real hardware, and the README points at the official openpi repository for that.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 28 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What pi-zero-pytorch actually is, and who it is for

π₀ is a robotic foundation model architecture from Physical Intelligence: a vision-language model that emits robot actions instead of text. pi-zero-pytorch is Phil Wang's independent PyTorch implementation of that architecture. The README describes it as a simplified Transfusion with influence from Stable Diffusion 3, specifically flow matching for policy generation in place of diffusion, plus the separation of parameters that mmDIT calls joint attention. The model is built on top of a pretrained vision language model, PaliGemma 2B.

The audience is narrow and the README says so directly. Its appreciation section ends with an open invitation to "maybe a phd student who wants to contribute to the latest SOTA architecture". This is a library for people who want to read the attention plumbing, swap components, and train a policy from scratch or from a VLM checkpoint. It is not a stack you install on a workstation and drive a manipulator with the same afternoon. There is no model card, no downloadable weights, and no rollout script against real hardware in the repository layout.

The README also carries a one-line update that matters more than its length suggests: the official repository, openpi, has been open sourced. That changes the calculus for anyone choosing a dependency, and it is worth reading as an honest pointer rather than a footnote.

Flow matching, joint attention and the PaliGemma backbone

The architecture is a mixture-of-experts transformer over several token sets. Vision tokens come from the VLM's image encoder, command tokens from tokenized language, and joint state and action tokens from the robot. Einops pack and unpack are used throughout to manage those sets, and the README credits Flex Attention for allowing an easy mixture of autoregressive and bidirectional attention in one pass. That mixture is the point: language is generated causally, while action chunks attend bidirectionally over the horizon they are predicting.

Generation uses flow matching rather than diffusion. In the README example the model is called with trajectory_length = 32 and returns a tensor shaped (1, 32, 6), so the policy emits a whole action chunk in one shot rather than a single step. Training is a single forward pass that returns a loss and an auxiliary value, then loss.backward(). The action input during training is shaped (1, 32, 6), matching the sampling shape, which is the consistency you want: the same tensor layout appears whether you are scoring a demonstration or generating a trajectory.

The parameter separation from mmDIT means the vision-language pathway and the action expert do not simply share one weight matrix. The repository also cites work on value residual learning, register tokens, guidance scale artifacts, recurrent memory transformers, Hyper-Connections, and several 2025 reinforcement learning papers in its citation block. That bibliography is the clearest signal of what the code is doing: it is assembling recent tricks from the diffusion and RL literature into one policy model.

Installing pi-zero-pytorch and running a first forward pass

The README gives a single install line. Python 3.10 or newer is required by pyproject.toml, and torch must be at least 2.5.

bash
$ pip install pi-zero-pytorch

After that, the README's usage example constructs the model with a 512-dimensional backbone, 6 action dimensions, a 12-dimensional joint state and a 20,000-token vocabulary, then runs a forward and backward pass on random tensors.

python
import torch
from pi_zero_pytorch import π0

model = π0(
    dim = 512,
    dim_action_input = 6,
    dim_joint_state = 12,
    num_tokens = 20_000
)

vision = torch.randn(1, 1024, 512)
commands = torch.randint(0, 20_000, (1, 1024))
joint_state = torch.randn(1, 12)
actions = torch.randn(1, 32, 6)

loss, _ = model(vision, commands, joint_state, actions)
loss.backward()

What you should see is a scalar loss and a populated gradient graph. Nothing in this example touches a robot or a real image; it is a shape check. Once that runs, sampling is the same call without the action argument, passing trajectory_length instead.

python
sampled_actions = model(vision, commands, joint_state, trajectory_length = 32) # (1, 32, 6)

That returns a (1, 32, 6) tensor. If you want to contribute rather than consume, the contribution path is also documented: install the test extra, add cases to tests/test_pi_zero.py, and run pytest.

bash
$ pip install '.[test]'
$ pytest tests/

The repository root also contains several verify_ scripts and a test_pi_zero_vlm.py, which suggests the VLM path has its own smoke tests, though the README does not document what they assert.

EFPO wraps the model for online learning, with a mock environment

The second documented entry point is EFPO, which the README describes as the wrapper you use for online learning. The example imports a mock environment from pi_zero_pytorch.mock_env, constructed as Env((256, 256), 2, 32, 1024, 12), and passes the model to EFPO.

python
from pi_zero_pytorch import π0, EFPO
from pi_zero_pytorch.mock_env import Env

mock_env = Env((256, 256), 2, 32, 1024, 12)
epo = EFPO(model)

memories = epo.gather_experience_from_env(mock_env, steps = 10)
epo.learn_agent(memories, batch_size = 2)

The two method names describe the loop: gather_experience_from_env collects memories for a fixed number of steps, and learn_agent consumes them at a batch size. The README says plainly that you will want to supply your own environment. The mock is a shape harness, not a simulator you would benchmark against.

The dependency list backs this up. It includes evolutionary-policy-optimization, x-evolution, memmap-replay-buffer, gymnasium with box2d, moviepy, and a fastapi plus uvicorn plus jinja2 trio. Those are the ingredients of a training loop with a replay buffer and some kind of web visualization, not of a runtime. If you are looking for a policy server, the fastapi dependency is suggestive but the README does not document an endpoint, a port or a launch command, so treat that as unverified.

No checkpoints, no real-robot path, and a moving dependency graph

The first limitation is the one that decides most evaluations: there are no pretrained weights in this package. The README's own flow goes from random tensors to loss.backward() to sampled_actions, with a comment reading "after much training". Training π₀ from scratch means a PaliGemma 2B backbone, a robot demonstration dataset, and compute on a scale the README never discusses. If your plan was to download something and run it, this repository does not offer that.

The second is that the action and state interfaces are raw tensors. You are responsible for normalizing joint states, chunking actions into the 32-step horizon, and matching dim_action_input and dim_joint_state to your hardware. Nothing in the README describes a dataset format, an observation schema, or how images are resized before they reach the vision encoder. The dependency list includes transformers and safetensors, which hints at a Hugging Face loading path, but the README does not document one.

The third is churn. pyproject.toml declares version 0.5.15 while the most recent release listed is 0.5.9, and the dependency list pins lower bounds on more than twenty packages, several of them also authored by the same maintainer (vit-pytorch, x-transformers, x-einops-utils, x-mlps-pytorch, ema-pytorch). A version bump in any of those can change behavior here. The last push to the repository was on 2026-09-02, and the repository is not archived, but the release cadence in early 2026 shows three releases in two days, which is the opposite of a frozen API. Pin your versions.

Finally, the README itself points at openpi. If your goal is a working π₀ policy rather than a readable implementation, that pointer is the most useful sentence on the page.

How pi-zero-pytorch differs from openpi and from a general VLA library

The obvious alternative is openpi, the official Physical Intelligence repository that the README links to. The difference is not quality, it is scope. openpi is the reference implementation from the group that proposed π₀, and it is where checkpoints, evaluation harnesses and hardware integration would live. pi-zero-pytorch is a single-author reimplementation whose stated value is simplification: fewer moving parts, readable attention, and a package you can pip install and modify. If you need to reproduce published numbers, go to openpi. If you need to understand or alter the architecture, the small package is the easier read.

A second comparison is against general imitation learning libraries. Those typically train a policy on a fixed dataset with a fixed observation format. pi-zero-pytorch instead assumes a pretrained vision-language model as the backbone and adds an action expert on top, with flow matching for generation. That is a heavier starting point and a different failure profile: your results depend on the VLM you plug in, and the README names PaliGemma 2B specifically. A library that trains from pixels alone has no such dependency and no such ceiling.

Licence, maintenance and upgrade cost

The project is MIT licensed, declared both in pyproject.toml and in the LICENSE file at the repository root. MIT is permissive: you can use, modify and redistribute the code, including commercially, provided the copyright notice and licence text are retained. That covers this implementation only. It does not grant you rights to PaliGemma or to any weights you load through transformers, and those carry their own terms. Nothing here is legal advice; check the licences of the backbone and any dataset you train on separately.

On maintenance, the facts are these: the repository is not archived, the last push was on 2026-09-02, and the latest published release listed is 0.5.9 from 2026-01-22. The gap between the version in pyproject.toml (0.5.15) and the newest listed release suggests releases are not the primary distribution channel or that the listing is incomplete. The README credits two outside contributors for code review and bug fixes, so there is some external review, but the project is effectively one person's work.

Upgrade cost is mostly dependency risk. With more than twenty lower-bounded dependencies, several maintained by the same author, a fresh install six months from now may resolve to a combination the code was not written against. Pinning to a known-good set is the practical mitigation, and the test extra plus pytest tests/ is the check that a pinned set still works.

Editorial conclusion

Adopt pi-zero-pytorch if you want to read, modify or pretrain the π₀ architecture in plain PyTorch and you already have a PaliGemma-scale VLM and a robot dataset to point it at. Do not adopt it if you need a working robot policy next week: there are no published checkpoints, no inference server for real hardware, and the README points at the official openpi repository for that. Before committing, verify two things yourself: that your torch build is at least 2.5 and your Python at least 3.10, because pyproject.toml requires both, and that the action and joint-state dimensions you pass to π0 match your actual robot, since the README example uses dim_action_input = 6 and dim_joint_state = 12.

Frequently asked questions

What is pi-zero-pytorch?

It is a PyTorch implementation of π₀, the robotic foundation model architecture proposed by Physical Intelligence. The README describes it as a simplified Transfusion with flow matching for policy generation and separated parameters for joint attention, built on top of a pretrained PaliGemma 2B vision language model.

Is pi-zero-pytorch the same as the official openpi release?

No. The README notes that the official repository, openpi, has been open sourced, and links to it separately. pi-zero-pytorch is an independent implementation by Phil Wang, distributed on its own as a pip package.

Does pi-zero-pytorch ship pretrained model weights?

The README does not document any downloadable checkpoints. Its usage example goes from random tensors to a loss and backward pass, then to sampled actions with the comment "after much training", which implies you train the model yourself.

What Python and PyTorch versions does pi-zero-pytorch require?

pyproject.toml sets requires-python to ">= 3.10" and lists torch>=2.5 among the dependencies. The README's install line is a single pip install pi-zero-pytorch.

Can pi-zero-pytorch be used for online learning on a robot?

The README documents wrapping the model with the EFPO class, gathering experience from an environment and calling learn_agent on the collected memories. It states that you will want to supply your own environment, and the example uses the bundled mock environment.

Official sources

  1. Issues
  2. License: MIT
  3. lucidrains/pi-zero-pytorch on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/lucidrains-pi-zero-pytorch.svg)](https://hysenlabs.com/projects/lucidrains-pi-zero-pytorch)