OpenClaw Kubernetes Operator: one CRD, self-configure allowlists and a distroless manager binary
Kubernetes operator for deploying and managing OpenClaw AI agent instances with production-grade security, observability, and lifecycle management.
At a glance
- What is it?
- The OpenClaw operator turns one OpenClawInstance manifest into a hardened stack of 9 or more Kubernetes objects, with an allowlisted self-configure path, S3 snapshots and opt-in registry auto-rollback. It is also a v0.40 project whose install section is not on the repository page.
- Who is it for?
- Adopt this operator if you want OpenClaw running inside a cluster you already secure and monitor, and stay on the managed hosting from Paperclip Inc. if you would rather not own a StatefulSet, a NetworkPolicy and a default-deny posture for a workload that holds messaging tokens.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
One OpenClawInstance expands into 9 or more Kubernetes objects
OpenClaw is an AI agent platform that acts on your behalf across Telegram, Discord, WhatsApp and Signal, handling an inbox, a calendar, a smart home and more through 50+ integrations. Paperclip Inc. sells managed hosting for it. This operator is the other option: the same agent on infrastructure you control.
The entry point is one custom resource in the `openclaw.rocks/v1alpha1` group:
apiVersion: openclaw.rocks/v1alpha1
kind: OpenClawInstance
metadata:
name: my-agent
spec:
envFrom:
- secretRef:
name: openclaw-api-keys
storage:
persistence:
enabled: true
size: 10GiThat manifest reconciles into 9 or more resources: a StatefulSet, a Service, RBAC, a NetworkPolicy, a PVC, a PodDisruptionBudget, an Ingress and others. The argument is that the workload is not the hard part of putting an agent in a cluster. Network isolation, secret handling, persistent storage, health monitoring and config rollouts are.
No provider is locked in. Anthropic, OpenAI or anything else arrives through environment variables and inline or external config. A singleton resource named `OpenClawClusterDefaults` with the name `cluster` fills unset fields on every instance, and per-instance fields always win.
Non-root UID 1000, read-only root filesystem, default-deny NetworkPolicy
Pod hardening arrives as defaults rather than switches. Non-root at UID 1000, a read-only root filesystem, all capabilities dropped, seccomp RuntimeDefault, a default-deny NetworkPolicy and a validating webhook.
Gateway authentication is automatic too, with a persistent token Secret per instance holding a shared secret, so no one has to invent a token and then remember to rotate it. On managed clusters, ServiceAccount annotations carry AWS IRSA or GCP Workload Identity, and a CA bundle can be injected for a corporate proxy that terminates TLS inside the cluster.
Reachability has a second path. Tailscale Serve or Funnel exposes an instance with SSO auth and no Ingress, and the same sidecar pattern covers Chromium for browser automation or Ollama for local models, next to your own init containers and sidecars.
What those defaults do not cover is the agent's own behaviour. A default-deny NetworkPolicy limits what the pod can reach and a read-only root filesystem blocks writes outside mounted volumes, but nothing there stops the agent from calling a messaging API with whatever token you handed it through `envFrom`. The least privilege work happens in the secret you mount, not in the operator.
config.forcePaths decides who owns a key under mergeMode: merge
Two config modes, and the difference is who wins on restart. `overwrite` replaces config on restart. `merge` deep-merges the custom resource with the config already on the PVC, preserving what the agent changed at runtime, and an init container restores config on every container restart.
Merge on its own leaves a hole for platform operators, because a tenant edit inside the pod then survives every restart. `config.forcePaths` closes it. It lists dot-paths that the init container rebuilds from the custom resource on every restart even under `mergeMode: merge`, and the named examples are auth, allowed providers and the sandbox image, so a managed deployer keeps those operator-owned while user-owned config persists.
Rollouts are content-hashed, so a change to the rendered config produces a new pod rather than an in-place mutation, and drift detection runs every 5 minutes against live state. Setting `spec.suspended: true` scales the runtime to zero while every non-runtime resource stays managed, and `false` resumes it.
Workspace seeding fills the volume before the agent's first run, with the option to reference an external ConfigMap for GitOps workflows, so the first thing the agent reads is the filesystem you meant to give it.
Without selfConfigure, an agent editing its own config triggers no restart
OpenClaw agents can install skills, patch config, add environment variables and seed workspace files through the Kubernetes API. Enabling that takes two steps: switch the feature on with an allowlist, then let the agent create a second resource.
spec:
selfConfigure:
enabled: true
allowedActions: [skills, config, envVars, workspaceFiles]apiVersion: openclaw.rocks/v1alpha1
kind: OpenClawSelfConfig
metadata:
name: add-fetch-skill
spec:
instanceRef: my-agent
addSkills:
- "@anthropic/mcp-server-fetch"Every request is validated against the instance's allowlist policy. Protected config keys cannot be overwritten, and denied requests are logged with a reason.
The failure mode is stated plainly: without `selfConfigure` enabled, config or skill changes the agent makes inside the container do not trigger a pod restart, and the pod has to be deleted by hand, for example `kubectl delete pod <pod-name>`. That is the state of a fresh install for anyone who expects the agent to adapt itself, and it produces a confusing symptom. The file on disk says the skill is installed while the running process never loaded it.
Registry polling, a snapshot first, and a rollback if health checks fail
Auto-update is opt-in and works against the OCI registry rather than a git tag. It checks for new semver releases, backs up first, rolls out, and rolls back automatically if the new version fails its health checks.
Backups are S3-compatible snapshots and they fire in three situations: on deletion, before an update, and on a cron schedule. A restore writes into a new instance from any snapshot, so recovering does not mean reusing the failed instance's PVC. The pre-update snapshot is what makes the automatic rollback survivable in the first place.
The lifecycle around it is ordinary Kubernetes machinery rather than something the project invented: PodDisruptionBudgets, health probes, HPA integration on CPU and memory with min and max replica bounds, and automatic StatefulSet replica management. Skills and plugins install declaratively as well, through `spec.skills` with `npm:` and `pack:` prefixes plus `additionalWorkspaces[].skills` for workspace-scoped sets, and through `spec.plugins` resolved by the OpenClaw CLI ClawHub installer inside a secure init container. Runtime dependencies get their own init containers: pnpm via corepack, or Python 3.12 with uv for MCP servers and skills.
diskReadiness ANDs the gateway check with a free-space test
A full or read-only workspace PVC is the quiet failure this operator has an answer for. `spec.probes.diskReadiness` renders the readiness probe as an exec check that ANDs the gateway `/readyz` signal with a workspace writability and free-space check. A pod that cannot persist what it accepts gets drained from Service endpoints instead of quietly taking writes. Liveness and startup stay HTTP on purpose, so a full disk never turns into a CrashLoopBackOff. The guard is defaulted off.
Cluster-wide defaults are the other lever. `OpenClawClusterDefaults`, named `cluster`, fills unset fields on every instance, which suits air-gapped or China-region clusters where each instance would otherwise repeat the same registry and mirror environment boilerplate, while per-instance fields always win. The final row of the feature table describes zombie reaping through a shared PID namespace under `spec.shareProc`, and the table is cut off in the middle of that key, so its full name and behaviour are not visible.
Observability needs no separate wiring: Prometheus metrics with ServiceMonitor integration, structured JSON logging and Kubernetes events.
Go 1.25, controller-runtime 0.19 and a distroless manager image
The build side is a kubebuilder-shaped Go project. `go.mod` declares `go 1.25.0` with `sigs.k8s.io/controller-runtime v0.19.0` and `k8s.io/api`, `k8s.io/apimachinery` and `k8s.io/client-go` all at `v0.31.0`, alongside the Prometheus client and the OpenTelemetry metric SDK that back the observability features. The `Dockerfile` compiles `cmd/main.go` with `CGO_ENABLED=0` in a `golang:1.25-alpine` stage and ships the binary on `gcr.io/distroless/static:nonroot` as `USER 65532:65532` with `ENTRYPOINT ["/manager"]`.
CRDs are generated rather than hand-edited. The visible `Makefile` targets are `manifests`, which runs controller-gen for rbac, crd and webhook objects into `config/crd/bases`, then `sync-chart-crds`, `sync-bundle-crds`, `generate` and `fmt`, with `ENVTEST_K8S_VERSION = 1.31.0` for the test assets.
make manifests
make sync-chart-crdsThree distribution paths sit in the tree: `charts/` for Helm, `bundle/` with `artifacthub-repo.yml` for the OLM bundle, and `docs-site/`. Apache-2.0 covers the code. Releases move in small steps, v0.38.3 on 2026-07-28, v0.39.0 on 2026-08-13, v0.40.0 on 2026-09-07, with the last push on 2026-09-17.
One gap deserves flagging. The README stops mid-word in the last table row, and the install section that would follow the feature table is not part of the page. Until you read `docs/`, pick an install path deliberately rather than assuming the Helm chart is the intended one.
Editorial conclusion
Adopt this operator if you want OpenClaw running inside a cluster you already secure and monitor, and stay on the managed hosting from Paperclip Inc. if you would rather not own a StatefulSet, a NetworkPolicy and a default-deny posture for a workload that holds messaging tokens. Verify three things before rolling it out: whether you install from charts/, bundle/ or raw CRDs, because the README stops mid-sentence before that section; whether spec.probes.diskReadiness and selfConfigure are the defaults you want, because both are off unless you ask; and what forcePaths pins, since a merge mode that keeps tenant edits will also keep an agent's own config edits across restarts.
Frequently asked questions
What is OpenClaw and what does it do?
OpenClaw is an AI agent platform that acts on your behalf across Telegram, Discord, WhatsApp and Signal, handling an inbox, a calendar, a smart home and more through 50+ integrations. This project is the Kubernetes operator that runs such an agent on infrastructure you control.
What does the OpenClaw Kubernetes operator reconcile?
One OpenClawInstance custom resource in the openclaw.rocks/v1alpha1 group expands into 9 or more Kubernetes resources, among them a StatefulSet, a Service, RBAC, a NetworkPolicy, a PVC, a PodDisruptionBudget and an Ingress.
How do merge and overwrite config modes differ in the OpenClaw operator?
`overwrite` replaces config on restart, while `merge` deep-merges the custom resource with the config on the PVC and preserves runtime changes. `config.forcePaths` lists dot-paths the init container rebuilds from the custom resource on every restart even under merge, keeping auth, allowed providers and the sandbox image immune to edits made inside the pod.
Does the OpenClaw operator restart pods when the agent edits its own config?
Not unless selfConfigure is enabled. Config or skill changes the agent makes inside the container do not trigger a pod restart by default, and the pod has to be deleted manually, for example with `kubectl delete pod <pod-name>`.
How does auto-update roll out a new OpenClaw version?
It is opt-in and polls the OCI registry for new semver releases. The operator snapshots to S3-compatible storage first, rolls out, and rolls back automatically if the new version fails its health checks.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/paperclipinc-openclaw-operator)