Kubeflow Katib: Kubernetes-native AutoML for Hyperparameter Tuning and Architecture Search
Automated Machine Learning on Kubernetes
At a glance
- What is it?
- Katib runs hyperparameter tuning, early stopping and neural architecture search as Kubernetes custom resources, so the search loop lives beside the training jobs it controls. The trade-off is that you need a cluster and a trial template before you get a single number back.
- Who is it for?
- Adopt Katib if your training already runs as Kubernetes workloads and you want the search loop to share that scheduler, quota and namespace model. Do not adopt it for a single laptop experiment: the control plane needs a cluster, and the README's install path assumes kubectl and kustomize are working.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Katib solves, and who ends up using it
Tuning a model means running the same training code many times with different values for learning rate, batch size, layer count or whatever else you expose, then keeping the run that scored best. Doing that by hand means a spreadsheet, a shell loop and a lot of copy-pasted job YAML. Katib replaces the loop with a controller that reads a search space, proposes configurations, launches each one as a Kubernetes workload and records the metric each run reports back.
The audience is narrower than "machine learning practitioners". The README describes Katib as a Kubernetes-native project for automated machine learning, and that word native carries the real constraint: the unit of work is a Kubernetes custom resource. If your training already runs on a cluster, Katib reuses the scheduler, the namespace boundaries and the resource quotas you already have. If your training is a notebook on a laptop, Katib asks you to build the cluster first and gain little from it.
Katib is framework agnostic in the sense the README states: it can tune hyperparameters of applications written in any language, with native support listed for TensorFlow, PyTorch and XGBoost among others. The tuning logic never imports your model. It launches something and reads a number.
Experiments, suggestions and trials: the control loop
The architecture splits into three pieces. An Experiment is the user-facing object: it holds the search space, the objective metric, the algorithm name, the maximum number of trials and a trial template. A Suggestion service is the optimiser, one of the algorithms the README lists, running behind a gRPC interface (the repository's go.mod pins google.golang.org/grpc and the goptuna, hyperopt, optuna and scikit-optimize integrations are named in the README). A Trial is one concrete run of the trial template with one set of parameter values.
The flow is a loop. The controller asks the Suggestion service for a new set of parameters, renders those values into the trial template, creates the Trial, and waits. When the workload finishes, a metrics collector reads the objective value from wherever the training code wrote it and reports it back. The Suggestion service takes that observation and proposes the next configuration. When the trial budget is exhausted or the algorithm converges, the Experiment reports the best trial.
That collector is the part people underestimate. Katib does not read your loss function in memory; it reads a file, a log line or a TensorFlow event stream that your training container produced. The repository layout reflects this, with a metrics collector under the Go packages and a TensorFlow event test fixture referenced in the Makefile. If your training script prints a number in a format the collector does not parse, the trial completes and the experiment learns nothing from it.
Two design consequences follow. First, the trial template is the integration surface: Katib can drive any Kubernetes custom resource, and the README names Kubeflow Training Operator, Argo Workflows and Tekton Pipelines as supported out of the box. Second, because each trial is a real pod, the cost of a bad search space is measured in cluster time, not CPU seconds on your desk.
Installing the control plane and running a first tuning job
The README points prerequisites and detailed instructions at the Kubeflow documentation, then gives the standalone control plane install as a kustomize reference. The command below is the one printed in the README, with its release ref left as written there. Note that the ref in the README is older than the most recent releases listed for the repository, so check which tag you actually want before applying it.
kubectl apply -k "github.com/kubeflow/katib.git/manifests/v1beta1/installs/katib-standalone?ref=v0.17.0"After this, the Katib controller and its supporting components appear in the cluster; the README does not spell out the namespace in the text above, so confirm placement with kubectl once the apply finishes. The README also offers the same command with ref=master for the latest changes, which is a moving target and not what you want on a shared cluster.
The Python SDK is the second half. It is published as kubeflow-katib and is described as a way to simplify creation of hyperparameter tuning jobs for data scientists.
pip install -U kubeflow-katibWith the SDK installed and the control plane running, the workflow is to define an Experiment in Python, submit it, and poll for the best trial. The README does not reproduce the SDK object model in the text above; it sends you to the getting started guide for that, and to examples/v1beta1 in the repository for the complete examples list. That examples directory is the honest starting point: copy a manifest that resembles your workload, replace the trial template with your own training job, and set the objective metric to whatever your container already writes out.
Where Katib is the wrong tool
The clearest failure mode is a metric that never reaches the collector. Katib's optimisers are only as good as the observations they receive, and the observation path runs through your container's output. A trial that exits before writing the metric, or writes it in a format the collector does not recognise, produces a completed trial with no usable objective. The experiment keeps going and the suggestion service keeps guessing, but the search is effectively random. Debugging this means reading trial pod logs, not the Experiment status.
Second, Katib is a poor fit when the training job cannot be expressed as a Kubernetes custom resource. The README frames trial templates around custom resources and names the operators it supports. If your training is a long-running interactive process, a proprietary managed service, or something that needs a GPU type your cluster does not schedule, the trial template becomes the blocker and no amount of algorithm choice helps.
Third, the algorithm catalogue is not a reason to choose Katib on its own. Random search, grid search, Bayesian optimization, TPE, multivariate TPE, CMA-ES, Sobol sequences, HyperBand and Population Based Training are all listed in the README, but several of them exist in standalone libraries that run in a single process with no cluster at all. Choosing Katib is a decision about where the loop runs, not about which optimiser is best. If your experiments fit on one machine, the cluster is overhead you are paying for scheduling and isolation you do not need.
Katib versus running the optimiser in-process
The closest alternative for many teams is Optuna used directly as a Python library, and the comparison is instructive because Katib's own suggestion layer integrates Optuna. With Optuna in-process, your script defines an objective function, calls study.optimize, and the trials run in the same Python process or in worker processes you manage. There is no controller, no custom resource, no metrics collector, and no cluster. The search space is a Python dictionary and the result is a Python object.
The difference is where state and scheduling live. Optuna in-process keeps the study in a storage backend you choose and relies on your own process management to parallelise trials. Katib moves both into Kubernetes: the study state lives in the Katib database (the go.mod pins MySQL and lib/pq drivers), and parallelism is expressed as the number of trials the controller may run at once, bounded by your cluster's capacity. You get retries, pod-level isolation and a record of every trial as a Kubernetes object. You also inherit the cluster's failure modes.
There is a middle path worth naming. Optuna and Hyperopt both appear in Katib's supported framework list, so a team can start in-process on a workstation and move the same search space into an Experiment when the workload outgrows one machine. That migration is not automatic, but the optimiser behaviour is at least familiar.
Maintenance, releases and what the Apache-2.0 licence means here
The repository is not archived, and the last push was on 2026-09-10, which is recent. Releases are not frequent: v0.18.0 landed on 2025-03-31 and v0.19.0 on 2025-10-30, with a release candidate in between. Between those tags the master branch moves, and the README's own install command offers ref=master as an option, which tells you the project expects people to track the branch. For a shared cluster, pin a tag instead.
Upgrade cost is dominated by the Kubernetes API surface rather than the optimiser code. The go.mod pins k8s.io/api, apimachinery and client-go at v0.34.0 and controller-runtime at v0.22.0, so a Katib release is tied to a Kubernetes version range. The Makefile sets ENVTEST_K8S_VERSION to 1.34.0 and the controller-gen version to v0.18.0, which is the project's own tested environment. If your cluster lags those versions, check compatibility before applying new manifests, because the custom resource definitions move with them.
Katib is Apache-2.0, the same licence as the wider Kubeflow project. That permits commercial use and modification with the usual attribution and notice obligations, and it includes a patent grant. It does not remove the operational cost of running a controller, a database and a suggestion service inside your cluster. For legal questions about redistribution or modification, the licence text is the source, not this summary.
Editorial conclusion
Adopt Katib if your training already runs as Kubernetes workloads and you want the search loop to share that scheduler, quota and namespace model. Do not adopt it for a single laptop experiment: the control plane needs a cluster, and the README's install path assumes kubectl and kustomize are working. Verify first that your trial template can be expressed as a Kubernetes custom resource, and check the manifests ref pinned in the install command against the release you intend to run.
Frequently asked questions
What is Kubeflow Katib used for?
Katib is a Kubernetes-native project for automated machine learning. The README states that it supports hyperparameter tuning, early stopping and neural architecture search, and that it is agnostic to machine learning frameworks.
What does the name Katib mean?
The README states that Katib stands for secretary in Arabic.
What is the origin of the name Katib?
The README only says that Katib stands for secretary in Arabic. It does not describe the word's further history or origin.
How does Katib compare with MLflow?
The README does not describe MLflow, so no comparison can be drawn from it. What it does say is that Katib is Kubernetes-native and performs training jobs using Kubernetes custom resources, with out of the box support for Kubeflow Training Operator, Argo Workflows and Tekton Pipelines.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/kubeflow-katib)