llmaz builds with Go 1.23 and asks for Go 1.24 in the same repository
☸️ Easy, advanced inference platform for large language models on Kubernetes. 🌟 Star to support our work!
At a glance
- What is it?
- A Kubernetes operator for serving large language models, with two API groups, five inference backends, a Python sidecar for model loading, and a Helm chart. The declared versions in the container build and the test harness do not match the ones the Go module requires.
- Who is it for?
- llmaz is worth a look if you want to serve a model on Kubernetes without hand-writing inference deployments, and you are willing to treat the API as alpha. That caveat is not boilerplate: both resource kinds live in v1alpha1 groups, and the documentation says outright that the API may change before Beta.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 37 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The example applies an OpenModel, not a Model
The quick start says you apply a Model and a Playground, and then shows two manifests that do not quite say that. The first declares `apiVersion: llmaz.io/v1alpha1` with `kind: OpenModel`. The second declares a different API group entirely, `apiVersion: inference.llmaz.io/v1alpha1`, with `kind: Playground`.
Two API groups in one toy example is worth pausing on. It means the model resource and the serving resource are versioned separately, so a cluster can be running one version of the model group and another of the inference group without either being obviously wrong.
The manifests themselves are short. The model names a `familyName`, a `source` pointing at a model hub entry with the model identifier, and an `inferenceConfig` with a list of flavors, each carrying a name and a limits block. In the example the flavor is called default and asks for one `nvidia.com/gpu`.
The Playground then claims the model by name and sets a replica count. That indirection is the design: the model describes what to load and on what hardware shape, the playground describes how many copies to run.
One preparatory step is easy to miss. If the model needs a HuggingFace token for the weight download, the secret has to exist before you apply anything.
Alpha means the version suffix on both groups is not a promise
The project states its own status in a callout above the overview: llmaz is alpha now, so the API may change before graduating to Beta.
Every resource in the example carries that status in its group name. `llmaz.io/v1alpha1` and `inference.llmaz.io/v1alpha1` are both alpha, which means both the shapes of the objects and the endpoints serving them are expected to move.
That matters more for this project than for a library, because an operator writes objects into a cluster and then has automations reading them back. A field rename is a normal library inconvenience and a migration in a cluster. Any tooling you build on top of the model or playground schemas has to assume that.
The release history is consistent with that posture. The listed tags are v0.1.2, v0.1.3, and v0.1.4, spaced across April and June of 2025, and the default branch was last pushed on 2026-08-27. There is no v1 tag and no stability badge beyond the alpha one at the top of the README.
The closest thing to a stability contract in the repository is the roadmap, and even that is a list of things not yet built: serverless support, prefill and decode disaggregated serving, KV cache offload, and model training and fine tuning in the long term.
The builder image pins Go 1.23 while the module declares Go 1.24
The manager image is built in two stages. The builder stage defaults to `golang:1.23.0`, and the final stage is a distroless static image running as UID 65532 with `/manager` as the entrypoint and CGO disabled. The final image is about as small as a Go operator gets.
The problem is one file over. `go.mod` declares `go 1.24.0` and pins `toolchain go1.24.4`.
A 1.23 toolchain cannot compile a module that declares 1.24 without either switching toolchains, which means downloading go1.24.4 at build time, or failing outright if the build is offline or the proxy is pinned. The Dockerfile does expose a `GOPROXY` build argument, so the download path exists, but the default builder image and the declared module version are still a mismatch.
The rest of the Dockerfile is careful in a way that suggests this was noticed elsewhere. The `TARGETARCH` argument is left without a default on purpose, with a comment explaining that leaving it empty lets Docker's build platform decide the architecture so the container and the binary inside it always match, including the Apple Silicon case. The manifest copy list is also narrow: only `cmd/main.go`, `api/`, `pkg/`, and `client-go/` are copied in, with a dependency download layer in front of them so source edits do not invalidate the module cache.
Envtest assets are pinned to Kubernetes 1.32 while the libraries are on 0.33
The Makefile pins the control plane binaries the test harness downloads:
ENVTEST_K8S_VERSION = 1.32.0The Go module, meanwhile, depends on `k8s.io/api`, `k8s.io/apimachinery`, `k8s.io/client-go`, `k8s.io/apiextensions-apiserver`, and `k8s.io/code-generator` all at `v0.33.5`, plus `sigs.k8s.io/controller-runtime v0.21.0`.
So the tests run against a 1.32 control plane while the client libraries are built for the 1.33 API surface. That is a defensible choice for compatibility, since envtest assets lag the libraries by design, but it is also a gap where a field the operator writes may not exist on the API server the tests boot. Worth knowing before you chase a failure that only reproduces under `make test`.
The rest of the harness is conventional and has one honest limitation written into it. `CONTAINER_TOOL` defaults to docker with a comment saying the target commands are only tested with Docker and suggesting podman as a substitute. Recipes run under bash with pipefail and errexit already set.
There is a Makefile-deps.mk pulled in at the top of the Makefile, alongside an index.yaml at the repository root, so the build manages its own pinned dependency set rather than resolving tool versions ad hoc.
A Go operator with a Python loader, and two different version numbers
The repository root contains `pyproject.toml` and `poetry.lock` inside a Go project, plus a second Dockerfile called `Dockerfile.loader`. That is the model-loading path, and it explains the model hub support in the feature list.
The Python dependencies are four lines: Python itself at `^3.10`, `huggingface-hub`, `modelscope`, and `omnistore`. Those map directly onto the three provider types the README names, HuggingFace, ModelScope, and object stores, with the loading handled automatically rather than by the user.
The version in that file is `0.0.1`. The newest GitHub release is v0.1.4. So the Python package, the Go module, and the release tags are three separate version lines, and only one of them is a release number anyone would recognise.
Development dependencies are minimal, just `black` and `pytest`, with black configured to a line length of 88. There is no linting or type checking configured for the Python side, which is consistent with the Python here being a sidecar rather than the main surface.
The rest of the layout is standard operator scaffolding: an API directory, a Helm chart, kustomize configuration under `config/`, a Kubebuilder `PROJECT` file, an `OWNERS` file, `client-go/`, `hack/`, a documentation site under `site/`, and an `index.yaml` at the root alongside `Makefile-deps.mk` for dependency management.
The verify path is completions, not chat
After applying the two manifests, llmaz creates a ClusterIP service named with an `-lb` suffix for load balancing. The port-forward targets that service rather than a pod:
kubectl port-forward svc/opt-125m-lb 8080:8080Then the endpoint list is short. `/v1/models` returns the registered models, and the query example posts to `/v1/completions` rather than `/v1/chat/completions`:
curl http://localhost:8080/v1/completions \
-H "Content-Type: application/json" \
-d '{
"model": "opt-125m",
"prompt": "San Francisco is a",
"max_tokens": 10,
"temperature": 0
}'That is worth noting because the chat interface is a separate feature. The built-in chat UI is an Open WebUI integration offering function calling, retrieval, and web search, and it is configured separately. The quick start deliberately exercises the raw completion path, which is the one thing every backend is expected to support.
The example model is `facebook/opt-125m`, small enough to schedule anywhere, and the README points CPU-only users at a llama.cpp example instead. The GPU request is a single `nvidia.com/gpu` in the flavor limits, which is also where heterogeneous scheduling comes in.
The service name convention is worth internalising if you script anything, because the generated service is not named after the object you applied. You apply a Playground named opt-125m and you forward a service named opt-125m-lb, and the two names are related by a suffix rather than by being identical.
Distributed inference is homogeneous today and heterogeneous is a promise
The feature list is precise about this. Multi-host and homogeneous xPyD support is available from day zero through the LeaderWorkerSet project, while heterogeneous xPyD is something that will be implemented in the future.
Heterogeneous serving appears elsewhere in the list, attached to a separate scheduler plugin project, and the stated reason is cost and performance: serving one model across mixed devices rather than requiring an entire homogeneous pool.
So the two mentions are not the same claim. One is about running many replicas of a prefill-decode setup that all look alike. The other is about a single request spanning different hardware. Only the first is in the current release.
The rest of the operational surface is delegated rather than built. Horizontal pod scaling uses HPA with LLM-based metrics, node autoscaling for spot instances uses Karpenter, and the gateway features such as token-based rate limiting and model routing come from integrating Envoy AI Gateway.
Backend support is the broadest part of the claim: vLLM, Text Generation Inference, SGLang, llama.cpp, and TensorRT-LLM are named, with the full list in the documentation site rather than in the README. Model providers are HuggingFace, ModelScope, and object stores.
Editorial conclusion
llmaz is worth a look if you want to serve a model on Kubernetes without hand-writing inference deployments, and you are willing to treat the API as alpha. That caveat is not boilerplate: both resource kinds live in v1alpha1 groups, and the documentation says outright that the API may change before Beta. Before you build anything on it, read the two version mismatches that will otherwise cost you an afternoon of confusing errors. The container build pins a Go 1.23 builder against a module that declares Go 1.24 with a 1.24.4 toolchain, and the test harness downloads envtest assets for Kubernetes 1.32 while the client libraries are on the 0.33 line. Also note the example applies an OpenModel, not a Model as the prose says, and that weight downloads for gated models need a HuggingFace token secret created first.
Frequently asked questions
What is llmaz and how do I deploy a model with it?
It is a Kubernetes inference platform for large language models. You apply two objects: an OpenModel in the llmaz.io API group describing the model and its hardware flavors, and a Playground in the inference.llmaz.io group claiming that model with a replica count. Then port-forward the generated load balancer service.
Which inference backends does llmaz support?
The README names vLLM, Text Generation Inference, SGLang, llama.cpp, and TensorRT-LLM, with the full list held in the documentation site. Model loading is handled for HuggingFace, ModelScope, and object stores through a separate Python loader.
Is the llmaz API stable?
No. The project states it is alpha and that the API may change before Beta. Both resource kinds in the quick start live in v1alpha1 groups, and the newest release tag is v0.1.4.
How do I download a gated model with llmaz?
Create the HuggingFace token secret before applying the manifests: `kubectl create secret generic modelhub-secret --from-literal=HF_TOKEN=<your token>`. The weight fetch happens when the model is loaded, not when the manifest is applied.
Does llmaz support heterogeneous GPU serving?
Not in the current release. Multi-host and homogeneous xPyD is supported from day zero through LeaderWorkerSet, while heterogeneous xPyD is listed as something to be implemented in the future. Cost and performance across mixed devices comes from a separate scheduler plugin project.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/inftyai-llmaz)