Self-hosted service
polyaxon/polyaxon avatar
polyaxon/polyaxon

Polyaxon: Kubernetes MLOps Platform for Experiment Orchestration

AI Infra / AI Orchestration / AI Control Plane

3,737 stars330 forksMDXApache-2.0

At a glance

What is it?
Polyaxon is a self-hosted MLOps platform that runs on Kubernetes and manages the full lifecycle of deep learning experiments. It handles reproducibility, hyperparameter search, distributed training, and pipeline automation from a single CLI and dashboard.
Who is it for?
Polyaxon suits ML engineering teams that already run Kubernetes and need one control plane for reproducibility, hyperparameter search, and distributed training across frameworks. Teams without Kubernetes have no supported path in.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly MDX, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What Polyaxon Solves and Who It Is For

ML teams at startups and large companies share a common problem: experiments are hard to reproduce, GPU servers are used inefficiently, and scaling a single training job to multiple nodes requires framework-specific plumbing that is separate from the experiment tracking tool. Polyaxon addresses all three at once. The README describes it as a system for reproducibility, automation, and scalability for machine learning applications. It deploys into any data center or cloud provider and supports all major deep learning frameworks: TensorFlow, MXNet, Caffe, Torch, and others. GPU servers become shared, self-service resources for a whole team, managed through smart container and node scheduling. The primary audience is a team of ML engineers or data scientists who own a Kubernetes cluster and want to stop coordinating jobs by hand.

How Polyaxon Orchestrates Workloads on Kubernetes

Polyaxon sits on top of Kubernetes as a control plane. Each experiment or service is defined in a polyaxonfile, a YAML configuration that describes the component to run, the resource requirements, and any hyperparameters. Polyaxon reads this file and submits the corresponding Kubernetes workloads, managing container scheduling and node allocation for you. Distributed jobs delegate execution to framework operators: TFJob for TensorFlow, PyTorchJob for PyTorch, MPIJob for MPI-based training, Dask, and Spark. The DAG engine, called the flow engine, lets you define pipelines where each operation can itself be a job, a service, a distributed run, or a nested DAG. This means you can chain data preprocessing, training, and evaluation into a single versioned pipeline without leaving the platform. A background Kubernetes operator (mloperator in the repository tree) watches cluster state and translates Polyaxon resource definitions into native Kubernetes primitives.

Installing the CLI and Deploying Polyaxon to Kubernetes

The CLI installs from PyPI with a single command:

bash
pip install -U polyaxon

The server deployment uses Helm. The README gives the following sequence:

bash
kubectl create namespace polyaxon
helm repo add polyaxon https://charts.polyaxon.com
polyaxon admin deploy -f config.yaml
polyaxon port-forward

Once the control plane is running, you create a project and run your first experiment:

bash
polyaxon project create --name=quick-start --description='Polyaxon quick start.'
polyaxon run -f experiment.yaml -u -l

The `-u` flag uploads local code and `-l` tails the logs. The dashboard starts with:

bash
polyaxon dashboard -y

From there, Jupyter notebooks and TensorBoard instances start as hub components:

bash
polyaxon run --hub notebook
polyaxon run --hub tensorboard -P uuid=UUID

The Helm chart path means teams can use standard GitOps tools to version and deploy the Polyaxon control plane alongside their other cluster services. The exact shape of config.yaml is documented in the polyaxon installation guide; the README points to polyaxon.com/docs/setup/ for the full reference.

Hyperparameter Search and Experiment Groups

Polyaxon implements hyperparameter search through a concept it calls experiment groups, which it compares in the README to Google Vizier. An experiment group combines a search algorithm, a search space, and a model definition. The supported algorithms are grid search, random search, Hyperband, Bayesian Optimization, Hyperopt, and a custom iterative approach for cases where none of the built-in algorithms fit. Parallel execution is handled through a mapping abstraction: you define a set of parameter combinations and Polyaxon runs them concurrently, managing resource allocation and collecting results automatically. This is the practical mechanism that lets a team use a shared GPU cluster without manually queuing jobs or writing scripts to coordinate parallel runs. The hypertune component in the repository handles this optimization engine.

Distributed Training Across Frameworks

Polyaxon does not implement distributed training itself. It provides the integration layer between your polyaxonfile and the Kubernetes operators that each framework uses. For TensorFlow, it delegates to TFJob. For PyTorch, to PyTorchJob. For MPI-based jobs and Horovod, to MPI Operator. Dask and Spark are also listed as supported distributed backends. The consequence of this design is that Polyaxon inherits the operational complexity of each operator: you need the relevant operator installed in the cluster before you can run that class of distributed job. The upside is that Polyaxon does not have to maintain its own distributed communication layer. Teams already using TFJob or PyTorchJob for other purposes can layer Polyaxon on top without replacing their existing infrastructure. The repository tree shows separate top-level directories for each major component: cli, haupt, hypertune, traceml, and mloperator, indicating a modular architecture.

Where Polyaxon Falls Short

The most direct limitation is the Kubernetes requirement. There is no supported path for running Polyaxon on a single developer machine without Kubernetes. Teams that do not already operate a cluster face significant setup overhead before they can run a single experiment. The README itself says installation requires helm and kubectl, and the quick start assumes a running Kubernetes cluster. A second limitation is that the repository's primary language is MDX, which means the bulk of the content is documentation rather than code. The actual inference engine and experiment tracking logic live in the component subdirectories (haupt, hypertune, traceml), but no releases are published on GitHub, which makes it harder to understand the version lifecycle from the repository alone. The README points to polyaxon.com for documentation and release notes. Teams that prefer everything self-contained in the open repository may find this inconvenient.

Polyaxon vs Kubeflow

Kubeflow is the other prominent Kubernetes-native MLOps platform. Both deploy onto Kubernetes, both support distributed training through framework operators, and both offer hyperparameter tuning. The primary architectural difference is scope: Kubeflow is a collection of loosely coupled components (Kubeflow Pipelines, Katib for hyperparameter tuning, Training Operator, KServe for serving) that teams assemble and integrate themselves. Polyaxon presents a unified control plane: one CLI, one dashboard, one polyaxonfile format covering experiments, hyperparameter search, pipelines, and serving configuration. Teams that want to pick individual MLOps components and connect them manually may prefer Kubeflow's modular model. Teams that want a single opinionated platform with one configuration format and one CLI may find Polyaxon's integrated approach easier to operate. Neither is a self-contained solution without Kubernetes.

Editorial conclusion

Polyaxon suits ML engineering teams that already run Kubernetes and need one control plane for reproducibility, hyperparameter search, and distributed training across frameworks. Teams without Kubernetes have no supported path in. Before adoption, check that the distributed training operators your frameworks need (TFJob, PyTorchJob, MPIJob) are already installed in the cluster, since Polyaxon delegates execution to those operators and does not bundle them.

Frequently asked questions

What deployment targets does Polyaxon support?

Polyaxon deploys to any data center, cloud provider, or can be managed by Polyaxon directly, according to the README. It requires a Kubernetes cluster and installs via Helm.

Can Polyaxon run on a single machine without Kubernetes?

The README does not document a path for running Polyaxon without Kubernetes. The install sequence requires kubectl and Helm, so a local Kubernetes installation such as minikube or kind would be the minimum.

How does Polyaxon track hyperparameter experiments?

Polyaxon uses a concept called experiment groups, comparable to Google Vizier per the README. Each group defines a search algorithm, a search space, and a model; supported algorithms include grid search, random search, Hyperband, Bayesian Optimization, and Hyperopt.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. polyaxon/polyaxon on GitHub
  4. Project website
  5. README
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/polyaxon-polyaxon.svg)](https://hysenlabs.com/projects/polyaxon-polyaxon)
Community notes

Community notes