google/ax: a declarative orchestrator for agent workloads on Kubernetes
Google's open agentic orchestration runtime
At a glance
- What is it?
- AX turns an agentic task into a Workspace and Task manifest, sandboxes it on Agent Substrate, and gives you kubectl-shaped verbs to watch, suspend and resume it. It is pre-1.0, and the README says breaking changes are coming.
- Who is it for?
- Adopt google/ax if you already run Kubernetes, have Agent Substrate installed, and want agent tasks expressed as YAML with suspend and resume as first-class verbs. Do not adopt it if you need a stable manifest schema, because the README states major breaking changes are likely before a stable release.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The workload AX was written for
Agent processes do not fit either of the two shapes a cluster usually handles. A microservice is long-lived and mostly stateless; a batch job runs to completion and exits. An agent accumulates state across turns, needs isolation because it may execute model-generated code, calls out to model APIs and tool servers, and can loop without producing anything. The README names that last property directly: an agent can burn money in a loop if nobody is watching.
AX addresses it with three declarative primitives under the ax.io/v1alpha1 API group. A Task is the sandboxed unit of work with CPU and memory limits. A Workspace pre-wires Git repositories, MCP servers and skill packages so an agent starts with its environment already in place. A Model configures which LLM the platform itself uses, with credentials pulled from a Kubernetes secret. The audience is platform teams who already think in manifests and want agent runs to be inspectable objects rather than processes on a VM.
How AX schedules a task: Substrate, Redis, and a gRPC control plane
AX is not self-contained. It runs on top of Agent Substrate, an external project that provides sandboxed execution, and it expects Substrate's Control API at api.ate-system.svc.cluster.local:443 in the ate-system namespace. Every task is scheduled as a sandboxed actor on Substrate, which means the sandbox boundary, not AX, is what contains the agent's code.
The control plane is a Go binary (cmd/ax-server) that the Makefile builds alongside the CLI. Dependencies in go.mod show the shape of it: google.golang.org/grpc for the client and server protocol, github.com/redis/go-redis/v9 for state, and the substrate and env modules from agent-substrate. The CLI is a gRPC client, and the README describes its verb set as deliberately kubectl-shaped: apply, get, describe, watch, delete, plus agent-specific suspend, resume and ssh.
That split matters when something breaks. If a task never leaves a pending phase, the fault may be in Substrate or in the Redis deployment rather than in your manifest. The README does not document rollback behaviour for a partially applied multi-document file, so treat apply as the point of no return for a given revision.
Installing the ax CLI and deploying the control plane
The prerequisites are a Kubernetes cluster with Agent Substrate installed, Go, kubectl, ko and a container registry your cluster can pull from. Substrate installs separately, following its own README, and lands in the ate-system namespace. Confirm it is reachable before doing anything else:
kubectl get svc api -n ate-systemIf that returns the api service, the Control API AX depends on is present. Next, install the CLI. The README uses go install, which places the binary in $(go env GOPATH)/bin, so that directory needs to be on your PATH:
go install github.com/google/ax/cmd/ax@latestThe Makefile offers an alternative that builds from a checkout with -trimpath and stripped symbols, running go install ./cmd/ax into the same GOPATH location.
With Substrate up, deploy the control plane. This target deploys Redis first, then builds and deploys the control plane images with ko, all into the ax-system namespace. AX_IMAGE_REPO defaults to gcr.io/ax-substrate/ate-images, so point it at a registry your cluster can actually pull from:
make deploy AX_IMAGE_REPO=<your-registry>Now apply a first task. The repository ships examples/task.yaml, which the README describes as a Task, Workspace and Model in one file:
ax apply -f examples/task.yaml
ax get tasks
ax watch task task123The get output has columns for NAME, ATESPACE, PHASE, ACTOR, WORKER-IP and AGE, so you should see a phase transition rather than a single static row. To look inside the running sandbox, the task needs spec.debug: true, and then the CLI gives you a shell or a one-off command:
ax ssh task123 -- ls -la /workspace
ax suspend task task123
ax resume task task123suspend checkpoints actor state and pauses; resume picks it up. There is also a demo.sh script in the repository root that applies a custom workspace, waits for readiness, runs commands over ax ssh, and suspends the task, which is the fastest way to see the whole lifecycle without writing YAML first.
Suspend and resume are the interesting part, and the least documented
Most orchestration systems can start work and kill work. Checkpointing an agent mid-run and resuming it later is a harder promise, because the agent's useful state lives in the sandbox, not in the control plane. The README lists ax suspend as checkpointing actor state and pausing, and ax resume as picking up where it left off, and it repeats the pair in the primitive table as the answer to pausing an idle agent.
What the README does not say is what happens to state that cannot be checkpointed: open network connections to an MCP server, a partially written file, an in-flight model call. The Sandbox guide is listed as covering what the runner does on boot and what a command can rely on, which is where that answer would live, but the README itself only asserts the capability. If your agent holds external sessions, test suspend and resume against your own workload before designing a cost-saving policy around it.
The same caution applies to the debug flag. ax ssh requires spec.debug: true on the task, so a production manifest without it gives you no interactive path into a misbehaving sandbox. That is a reasonable default for isolation and an inconvenient one for incident response, and it is a per-task decision you have to make in advance.
Where AX is the wrong tool
AX assumes Kubernetes and assumes Agent Substrate. If you have neither, the install path is not a single binary you can run on a laptop; the README's own prerequisites list a cluster, Substrate, ko and a pullable registry. For a single agent on a developer machine, that stack is overhead with no payoff.
The project is also explicitly unstable. A warning at the top of the README states that AX and several of its features are in heavy development, that core concepts, protocols and specifications are being refined, and that major breaking changes are likely before a stable release. The version numbers agree: v0.3.1 is the latest release. A team that needs a frozen manifest schema for a long-lived internal platform should not build on ax.io/v1alpha1 yet.
Finally, AX orchestrates agents; it does not give you an agent framework. There is no prompt library, no tool-calling abstraction and no evaluation harness in the README. It sits below that layer, wiring workspaces and sandboxes and models, and expects something else to decide what the agent actually does.
How AX differs from a general workflow engine
Airflow, Argo Workflows and similar engines model a DAG of steps: each node starts, does a bounded piece of work, and finishes. Retries and dependencies are the core abstractions, and a step that never terminates is an anomaly to be timed out.
AX inverts that. A Task is a single long-running sandboxed actor, not a node in a graph, and the interesting operations are on the actor's lifecycle rather than on its dependencies. The README's own framing is that agents are neither stateless microservices nor run-to-completion batch jobs, which is why the primitive set includes suspend and resume instead of retry policies. If your problem is a pipeline of deterministic steps, a workflow engine is the better fit and AX adds a sandbox you do not need. If your problem is many concurrent agents that each need an environment, a model binding and an isolation boundary, the workflow engine has no natural place to put those.
Licence, releases, and what an upgrade costs you
AX is Apache-2.0, the same licence as Kubernetes, and the Makefile headers carry the standard Apache boilerplate. That permits commercial use and modification, and it includes a patent grant. It does not tell you anything about the stability of the API group, which is a separate question governed by the README's development warning. This is not legal advice; check the LICENSE file and your own counsel for your situation.
The release cadence visible in the repository is fast: v0.2.3, v0.3.0 and v0.3.1 all landed within roughly six weeks, and the last push to main was on 2026-09-27. That is a project moving quickly enough that pinning matters. `go install github.com/google/ax/cmd/ax@latest` follows whatever is newest, so a CLI upgrade and a control plane upgrade can drift apart if you deploy the server from a different revision. The Makefile builds both binaries from the same checkout, which is the safer pairing.
Upgrade cost has two parts. The CLI and control plane are Go binaries you rebuild, so the mechanical work is small. The manifests are the expensive part: because the API is v1alpha1 and breaking changes are expected, a version bump can invalidate YAML you have already committed. Keep your Task, Workspace and Model files in version control next to the AX revision they were written against, and read the release notes for each tag before moving.
Editorial conclusion
Adopt google/ax if you already run Kubernetes, have Agent Substrate installed, and want agent tasks expressed as YAML with suspend and resume as first-class verbs. Do not adopt it if you need a stable manifest schema, because the README states major breaking changes are likely before a stable release. Verify first that kubectl get svc api -n ate-system returns a Control API, that a registry reachable by your cluster backs AX_IMAGE_REPO, and that spec.debug is set on any task you intend to reach with ax ssh.
Frequently asked questions
What is google/ax?
It is a declarative orchestrator for autonomous agent workloads that runs on a Kubernetes cluster, built on top of Agent Substrate for sandboxed execution. Tasks, Workspaces and Models are expressed as ax.io/v1alpha1 manifests and applied with the ax CLI.
How do I install the google/ax CLI?
The README uses go install github.com/google/ax/cmd/ax@latest, which places the ax binary in $(go env GOPATH)/bin. The Makefile also provides a make install target that builds from a checkout.
Does google/ax require Agent Substrate?
Yes. AX schedules every task as a sandboxed actor on Agent Substrate, and it expects Substrate's Control API at api.ate-system.svc.cluster.local:443 in the ate-system namespace. The README says to verify it with kubectl get svc api -n ate-system before deploying AX.
Why can I not open a shell into my google/ax task?
ax ssh requires the task to set spec.debug: true, as shown in the README's example manifest. Without that field the sandbox has no interactive entry point.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/google-ax)