Open-source project
PhyAgentOS/PhyAgentOS-core avatar
PhyAgentOS/PhyAgentOS-core

PhyAgentOS: a verifier-first runtime for embodied agents

PhyAgentOS is a Recursive Self-Improving (RSI) physical agent operating system that enables agents to recursively self-improve through agentic workflows.

2,635 stars128 forksPythonMIT

At a glance

What is it?
PhyAgentOS wraps robot actions in one versioned Forge Gateway contract, captures evidence around each bound Action, and lets a task-level verifier decide whether the goal was actually met. It is a Python 3.11+ framework for engineers who need auditable execution, not another planner.
Who is it for?
Adopt PhyAgentOS if you already have a robot or simulator stack and the missing piece is an auditable execution and verification boundary: the Forge Gateway contract, evidence records, and Planner-owned recovery are the parts you would otherwise build yourself. Do not adopt it if you want a ready-made manipulation policy, if you cannot run Python 3.11+ and the Docker image, or if your team will not maintain the SQLite task store.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap PhyAgentOS targets: execution facts versus task verdicts

Most embodied agent stacks are planners with a thin actuator layer bolted on. The planner emits a tool call, something executes it, and success is inferred from the absence of an exception. PhyAgentOS takes the opposite position. The README describes a loop in which the Agent plans high-level Tool calls, the Forge Tool API reports what the Gateway executed, an observation collector captures before/after evidence, and a task-level verifier decides whether the user-visible goal was achieved. Those are four distinct responsibilities with four distinct failure modes, and the project's claim is that collapsing them is what makes embodied agents hard to debug.

The intended user is an engineer who already has a working robot or simulator and is tired of guessing why a run failed. PhyAgentOS is not a policy. It is the harness around one. The changelog shows this focus hardening over time: v0.2.0 introduced the Forge execution architecture, immutable execution and evidence contracts, crash-safe SQLite orchestration, and removed the legacy Runtime execution chain entirely. That removal is the clearest statement of intent in the repository. The project does not keep two ways to execute an action.

One execution boundary: how the Forge Gateway and verifier fit together

The architecture rests on a single rule stated in the README's feature table: robot actions enter through one versioned Forge Gateway contract, and the Agent never reaches into a policy, simulator, Dora node, or hardware SDK. Everything else follows from that constraint. Because the Agent cannot call hardware directly, the Gateway becomes the only place where an action becomes real, and therefore the only place where evidence needs to be captured.

Evidence is captured around bound Actions and stored with source, sequence, time, size, digest, and retention metadata. The verifier then receives the goal, criteria, constraints, execution facts, evidence, lineage history, and optional Skill-scoped advisories. Two design choices stand out. First, the verifier is action-agnostic: there is no per-action verification switch, so adding a new Skill does not mean writing a new success check. Second, advisories cannot replace criteria or evidence, which means a Skill can hint at what matters without being able to declare victory.

Recovery is owned by the Planner, not the executor. A recovery verdict appends a bounded PlanRevision to the same task, and unknown effects are reconciled rather than retried blindly. The word "bounded" matters here: the README does not describe an unbounded retry loop, and a reviewer should treat the revision bound as a configuration detail to confirm in the code before trusting long-running tasks.

Installing PhyAgentOS and running the CLI for the first time

The repository ships a Dockerfile and a docker-compose.yml, and the Dockerfile is the most reliable description of the intended install path. It builds from ghcr.io/astral-sh/uv:python3.12-bookworm-slim, installs Python dependencies with uv, builds the WhatsApp bridge with npm ci and npm run build, then runs a release smoke check. That smoke check is the console entry point:

dockerfile
RUN paos --version

If the image builds, the paos command exists. The compose file defines two services from the same build context. The gateway service has no host port mapping, and the comment in the file explains why: it is a message-bus service that makes outbound connections only and does not bind an inbound port. The CLI service is behind the cli profile and runs the agent command with a TTY:

yaml
services:
  phyagentos-gateway:
    container_name: phyagentos-gateway
    command: ["gateway"]
    restart: unless-stopped
  phyagentos-cli:
    profiles:
      - cli
    command: ["agent"]
    stdin_open: true
    tty: true

Both services mount ~/.PhyAgentOS into /root/.PhyAgentOS, so configuration and state live on the host rather than in the container. For a local install, pyproject.toml declares requires-python >=3.11 and the package name PhyAgentOS-ai. The README points readers to docs/README.md for documentation; it does not reproduce a step-by-step pip install, so treat the Docker path as the documented route and read the docs directory before wiring a bare-metal environment.

Where PhyAgentOS is the wrong tool

The dependency list in pyproject.toml tells you what this project is not. It includes litellm, openai, tiktoken, python-telegram-bot, slack-sdk, lark-oapi, dingtalk-stream, mcp, and websockets. That is a chat-and-channel framework's dependency set, and the Dockerfile even builds a Node.js 20 WhatsApp bridge. If your problem is low-level control latency, a policy inference loop, or anything that must run inside a real-time control budget, this is not the layer to reach for. The Gateway is a message bus, and message buses add hops.

The second limitation is operational. Task aggregation is persisted through SQLite transactions covering AgentTask, PlanRevision, Query records, and Gateway invocation references. That is a sensible choice for a single-node harness and a poor one for a fleet: the README does not describe a distributed store, so multi-robot deployments would need to confirm how state is shared. Third, the verifier's quality is bounded by the evidence. If your sensors cannot produce validated images or robot state around an action, the verifier is reasoning over execution facts alone, and the README's own framing puts evidence before verdict. Finally, the repository is young. The changelog runs from 2026-04-29 to the v1.0.0 release on 2026-09-05, and the last push was on 2026-09-10. That is a short history for a system that asks to sit between your planner and your hardware.

PhyAgentOS compared with a plain tool-calling agent loop

The obvious alternative is a general agent framework driving a tool-calling loop, where each robot command is a function the model can invoke. The difference is where the truth lives. In a plain loop, the tool returns a string, the model reads it, and the next step proceeds on the model's interpretation of that string. PhyAgentOS inserts three things between the call and the next step: a versioned Gateway contract that defines what execution means, an evidence record with digest and retention metadata, and a verifier that receives criteria and constraints separately from the execution facts.

That extra machinery costs you. You must express your hardware as Query, Action, and Session tool lifecycles, and the changelog shows this interface has been reworked repeatedly, from the v0.2.2 unification of Forge execution on the Query/Action Tool API to the v0.2.3 addition of independent Skill installation and immutable AgentTask bindings. If your task is a one-off demo, a plain loop will get you there faster. If your task is a repeated physical workflow where a silent failure is expensive, the evidence-and-verdict split is the part you would otherwise have to invent, and inventing it badly is how embodied agent projects stall.

Licence, maintenance and the cost of upgrading

PhyAgentOS-core is MIT licensed, and pyproject.toml declares license = {text = "MIT"} with the OSI classifier. For most adopters that means permissive use with attribution and no copyleft obligation on your own code. It does not settle the licence of the models, simulators, or robot SDKs you connect through the Gateway, and the README does not address that boundary, so review it separately. This is not legal advice.

On maintenance, the last push was on 2026-09-10, and the repository is not archived. The changelog shows a fast cadence through 2026, with breaking structural changes as recently as v0.2.0, which removed the legacy Runtime execution chain. Upgrading across that boundary is not a dependency bump; it is a rewrite of anything that called the old path. The v0.2.3 notes describe Skills that can be installed and managed independently and activated into immutable AgentTask bindings, which suggests the intended upgrade story is version-scoped: experience and Skills carry version identity. Budget for reading CHANGELOG.md before every upgrade, and pin your version, because the project is still defining its own interfaces.

Editorial conclusion

Adopt PhyAgentOS if you already have a robot or simulator stack and the missing piece is an auditable execution and verification boundary: the Forge Gateway contract, evidence records, and Planner-owned recovery are the parts you would otherwise build yourself. Do not adopt it if you want a ready-made manipulation policy, if you cannot run Python 3.11+ and the Docker image, or if your team will not maintain the SQLite task store. Before committing, verify the paos console entry point exists after install, confirm the gateway's outbound-only design fits your network, and read the Forge Tool API lifecycle docs to check that your hardware adapter can be expressed as Query, Action, and Session calls.

Frequently asked questions

What is PhyAgentOS-core and who is it for?

It is an agent framework for embodied tasks in Python, where the Agent plans high-level Tool calls, the Forge Tool API reports what the Gateway executed, an observation collector captures before/after evidence, and a task-level verifier decides whether the goal was achieved. It targets engineers who already have a robot or simulator and need an auditable execution boundary rather than a policy.

How do I install PhyAgentOS?

The repository provides a Dockerfile and a docker-compose.yml. The Dockerfile builds from a uv Python 3.12 base image, installs dependencies, builds the WhatsApp bridge, and runs paos --version as a release smoke check; pyproject.toml declares requires-python >=3.11. The README points to docs/README.md rather than listing a bare-metal install.

Does the PhyAgentOS gateway expose a port I can connect to?

No. The docker-compose.yml comment states the gateway is a message-bus service that makes outbound connections only and does not bind an inbound port, and the service definition has no host port mapping. The CLI service runs under the cli profile with the agent command and a TTY.

Official sources

  1. License: MIT
  2. PhyAgentOS/PhyAgentOS-core on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/phyagentos-phyagentos-core.svg)](https://hysenlabs.com/projects/phyagentos-phyagentos-core)