shell-operator: an operator-sdk for scripts, running hooks inside the cluster
Shell-operator is a tool for running event-driven scripts in a Kubernetes cluster
At a glance
- What is it?
- A Go program that watches Kubernetes objects and runs your bash or Python when they change, treating scripts as hooks. It is not an operator for any particular product, which is both its point and the reason it needs a wrapper before it is useful.
- Who is it for?
- shell-operator is the right foundation when the behaviour you need in a cluster is imperative and you already know how to write it in bash, python or kubectl, and the wrong choice when you want a typed, tested controller for a CRD with a well-specified reconcile loop. It is Apache-2.0, written in Go against Kubernetes client libraries at 0.34.8, and release v1.20.6 shipped on the same day as the last push, 2026-09-24, so it is actively worked on.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 8, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Not an operator for a product, but an integration layer for events
The README draws the distinction early and it is the right one to start from. Shell-operator is not an operator for a particular software product such as prometheus-operator or kafka-operator. It provides an integration layer between Kubernetes cluster events and shell scripts, by treating scripts as hooks triggered by events. The framing it uses is an operator-sdk, but for scripts.
That framing has a consequence worth sitting with. A conventional operator wraps one API with typed structs, a reconcile loop and generated clients. This one does not wrap any API at all. It gives you a watcher and a trigger mechanism, and the behaviour inside the trigger is whatever you wrote. All the Kubernetes-specific machinery lives in the Go binary; all the domain logic lives in your script.
The feature list in the README covers the machinery rather than any domain. Hooks can be triggered by add, update or delete events. Object selectors and property filters let a hook monitor a particular set of objects and detect changes in specific properties. Configuration is simple in a specific sense: the hook binding definition is a JSON or YAML document printed on the script's stdout. There is also validating webhook machinery, so a hook can handle validation for Kubernetes resources, and conversion webhook machinery, so a hook can handle version conversion.
Those last two are worth pausing on, because admission webhooks are normally a Go coding exercise with TLS plumbing and API machinery. Being able to implement one as a script is the least obvious capability in the list.
The stdout binding is the entire contract
The most important design decision is the one about configuration: a hook declares what it wants by writing a JSON or YAML document to its standard output. There is no manifest file to keep in sync with the script, and no schema to declare up front.
The practical effect is that monitoring and reacting are separated. A script runs, inspects what changed, and returns a binding that says what it wants to watch next time. The event stream is therefore shaped by the script's own output, which means the watch set can be dynamic rather than fixed in a CRD.
It also means debugging is a matter of reading a script's output rather than reading operator logs, and that a misbehaving hook manifests as missing events rather than as a crash. For an imperative tool that is a reasonable trade. For anything where you want the platform to validate your intent before acting, it is a gap, which is presumably part of why the addon-operator layer exists.
The README points to docs/src/HOOKS.md for the detail on triggers, and that file is the one to read first. The main README is a description of intent; HOOKS.md is the specification.
Building and running it as a container image
There is no install command in the README, because this is a Go program that gets built into a container image and deployed into a cluster rather than installed on a machine. The Dockerfile shows what that involves, and it is a two-stage build with a prebuilt jq stage in front of it.
FROM --platform=${TARGETPLATFORM:-linux/amd64} registry.deckhouse.io/container-factory@sha256:4b36dcf53c35b50e0afbc445232713aff15f788a61b832cd720bf9e88fc9fba8 AS libjq
+FROM --platform=${TARGETPLATFORM:-linux/amd64} registry.deckhouse.io/container-factory@sha256:193e8ed6cd7fc19015ab615ccf92d0fe02471e66e3e5abf560b3a87fb05bdb62 AS builder
++ +The build stage runs go mod download, then builds a single binary from ./cmd/shell-operator with a version stamped in through ldflags:
RUN go build \
+ -ldflags="-s -w -X 'github.com/flant/shell-operator/pkg/app.Version=${appVersion}'" \
+ -o shell-operator \
+ ./cmd/shell-operator
++ +The runtime stage is deliberately minimal. It installs ca-certificates, bash, sed, tini and curl, then downloads a kubectl binary matching a kubectlVersion build argument, defaulting to v1.34.8, into /bin/kubectl. That kubectl binary is the clue to how this is meant to be used: your scripts are expected to shell out to kubectl against the cluster the operator runs in, so the tool has to be on the PATH inside the container. + +The image also creates a /hook directory, which is where the hook scripts are mounted. There is a shell_lib.sh in the repository root and a frameworks/ directory, which is where the bash and python conveniences live. + +For local work the Makefile gives you the standard targets. go-module-version regenerates the module version command against the current HEAD, lint runs golangci-lint with --fix, and test runs the suite under the race detector with coverage:
+test: go-check
+ @$(GO) test --race --cover ./...
+What the code actually depends on
go.mod is worth reading because it shows the shape of the integration. The module path is github.com/flant/shell-operator and it declares Go 1.26.4.
The Kubernetes surface is the expected one: k8s.io/api, apimachinery, client-go and apiextensions-apiserver all at v0.34.8. The presence of apiextensions-apiserver is what makes custom resource definitions and webhooks possible, rather than the tool only being able to watch built-in kinds.
Around that sit several libraries that explain specific features. gopkg.in/robfig/cron.v2 points to scheduled triggers alongside event-driven ones. gofrs/uuid/v5 is there for identities. prometheus/client_golang provides metrics, and the OpenTelemetry packages for otel, the otlptrace exporter and the trace SDK show that tracing is a first-class concern rather than an afterthought. hashicorp/go-multierror and pkg/errors suggest error aggregation across the watch and hook machinery. itchyny/gojq is notable: it means jq semantics are available in-process, which pairs with the prebuilt jq in the Dockerfile.
There are two Deckhouse dependencies, pkg/log and pkg/metrics-storage, plus a module-sdk. That coupling to Flant's own libraries is the clearest signal about where this code came from and what it assumes, and it is worth weighing before adopting it outside a Deckhouse-adjacent setup.
Who runs this in anger, and what addon-operator adds
The README names three prominent users, which is more informative than any adoption metric. Deckhouse, a Kubernetes platform, uses both shell-operator and addon-operator as core technology for configuring and extending Kubernetes features. KubeSphere uses it in its installer. Confluent uses it in a Kafka DevOps solution.
All three are platform or operations tooling rather than application code, which matches the design. This is infrastructure glue.
The addon-operator relationship is the key thing to understand, and the README states it directly: shell-operator is used as a base for the more advanced addon-operator, which supports Helm charts and value storages.
Read that as a layering statement about ambition. Shell-operator gives you events and scripts. Addon-operator adds chart rendering and persisted values, which is what you need if your hook is supposed to configure a chart rather than run a command. If your use case involves Helm and stored values, shell-operator alone is the lower layer and you probably want the one above it.
The example directory in the repository holds more use cases than the README describes, and there is a Show and tell discussion category where users post theirs, which is the fastest way to find something close to your own problem.
Where this loses to writing a real operator
The honest comparison is not against other shell tools, it is against operator-sdk and against just writing the controller in Go, and in several dimensions Go wins.
You lose compile-time checking on your logic. A bash hook that misreads an object field fails at three in the morning in a cluster rather than at build time, and there is no unit test story for a shell script that is as cheap to write as the Go equivalent would be. You lose a reconcile model: a hook fires on events, and nothing enforces that the cluster converges to a desired state if an event is missed or a hook fails midway.
You also lose the declarative contract. With a CRD and typed client, the shape of what you manage is in the API server, discoverable with kubectl get. Here the behaviour is in scripts mounted into a pod, which is harder to review, harder to version alongside cluster state, and easier to break with a bad edit.
What shell-operator buys for those losses is speed of development and the fact that your existing operational knowledge transfers directly. A runbook you already trust as a bash script becomes a hook by adding a watch and a trigger, instead of being rewritten in a language your team may not want to maintain.
The webhook support is where the calculus shifts most, because implementing a validating webhook in Go normally means certificate management, API registration and retry semantics that are tedious and easy to get subtly wrong. Being able to write that as a script is a real saving for a platform team that already runs this in production at Deckhouse and KubeSphere scale.
Editorial conclusion
shell-operator is the right foundation when the behaviour you need in a cluster is imperative and you already know how to write it in bash, python or kubectl, and the wrong choice when you want a typed, tested controller for a CRD with a well-specified reconcile loop. It is Apache-2.0, written in Go against Kubernetes client libraries at 0.34.8, and release v1.20.6 shipped on the same day as the last push, 2026-09-24, so it is actively worked on. Read HOOKS.md in the docs before writing anything, because the binding contract on script stdout is the whole interface, then decide honestly whether you need shell-operator alone or the addon-operator layer on top of it, which is how Deckhouse actually uses this code.
Frequently asked questions
What is shell-operator and what problem does it solve?
It is a Go program that runs inside a Kubernetes cluster and executes your scripts when cluster events occur, treating scripts as hooks. It provides the integration layer between Kubernetes events and shell scripts, in the way operator-sdk does for controllers, rather than managing any particular software product.
How does a shell-operator hook declare what it wants to watch?
By printing a hook binding definition as a JSON or YAML document on the script's standard output. That binding tells the operator which objects and properties to monitor, and the watch set is therefore determined by the script's own output rather than by a fixed manifest.
Should I use shell-operator or addon-operator?
The README describes addon-operator as a more advanced project built on shell-operator as its base, adding support for Helm charts and value storages. If your hook needs to render a chart or keep values between runs, you want the layer above; if it just reacts to events with a script, shell-operator is the right size.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/flant-shell-operator)