# OpenPAI: Microsoft's Kubernetes Cluster Manager for Shared GPU Training

> OpenPAI is a full stack platform for sharing GPU, FPGA and InfiniBand capacity across teams, built on Kubernetes and now in read-only stable mode. It fits organizations that run their own hardware and need job queuing, quota and a web portal more than they need the newest scheduler.

**microsoft/pai** — Resource scheduling and cluster management for AI

- Repository: https://github.com/microsoft/pai
- Website: https://openpai.readthedocs.io
- Stars: 2,694 · Forks: 554
- Language: JavaScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/microsoft-pai

## The problem OpenPAI solves: idle GPUs and competing teams

Most AI groups do not start with a cluster. They start with a few machines under a desk, then a rack, then a room, and at some point the question changes from "can we train this model" to "who gets the GPUs on Thursday". That is the problem OpenPAI targets. The README lists four situations where it is worth considering: sharing powerful AI computing resources such as a GPU or FPGA farm among teams, sharing and reusing common assets like models, data and environments, giving IT a single operations platform for AI work, and running a complete training pipeline in one place.

The audience is therefore not a single researcher. It is the platform team inside an organization that owns the hardware and has to answer for its utilization. The README frames the value as on-premises, hybrid or public cloud deployment, including a single-box option, which matters because the same tool is expected to run on a one-machine proof of concept and on a production farm. If your team rents GPU time per hour from a cloud provider and never sees a physical card, OpenPAI is solving a problem you do not have.

## How OpenPAI is put together: Kubernetes plus four standalone pieces

The architecture diagram in the README is the clearest statement of how the system works. At the bottom sits Kubernetes cluster management, and below that the hardware layer the README labels CPU/GPU/FPGA/InfiniBand. Everything OpenPAI adds is layered above Kubernetes rather than replacing it, which is why the project describes itself as a full stack solution.

Four components carry the actual scheduling and runtime work, and each is a separate repository linked from the README: frameworkcontroller for job orchestration, hivedscheduler for job scheduling, openpai-runtime for job runtime, and a job error analysis component. Around those sit the user-facing pieces: a web portal, an API, a VS Code extension, an SDK, and services for user authentication, user and group management, storage management, and cluster and job monitoring. There is also a marketplace repository for sharing assets.

The important design consequence is modularity. The README states that OpenPAI became more modular so the platform can be customized and expanded, and the standalone components are published separately, which means a team can in principle adopt hivedscheduler or frameworkcontroller without the rest of OpenPAI. That is a real escape hatch, and it is also the reason upgrades require attention: the platform version and the component versions are not the same thing.

## Installing OpenPAI and submitting a first job

The README splits getting started into two tracks, one for cluster administrators and one for cluster users. The administrator track is the one that matters for an install, and the repository ships the tooling under deployment/ with a top-level paictl.py entry point. The README does not reproduce the full command sequence, so treat the deployment directory and the documentation site at openpai.readthedocs.io as the source of truth for the exact flags.

The examples directory contains runnable starting points rather than abstractions. The repository layout lists examples/pytorch_cifar10, examples/tensorflow_cifar10, examples/MXNet_cifar10, examples/cntk_mnist, examples/mnist_500_tasks, examples/Distributed-example, examples/Dockerfiles and examples/cluster-configuration as sibling directories under examples/. For a first real use, the PyTorch CIFAR-10 example is the most direct path, since it exercises a real training loop rather than a synthetic sleep job. The pattern OpenPAI uses is that a job references a container image plus a command, and the cluster configuration under examples/cluster-configuration describes the machines a job can land on. Administrators edit that configuration to describe their own hardware before deploying. Users do not touch it.

For the user side, the README points at the web portal and the SDK, and the SDK lives in its own repository, microsoft/openpaisdk, linked from the README. A user submits a job through the portal or the SDK, the API accepts it, frameworkcontroller turns it into Kubernetes objects, and hivedscheduler decides which GPUs it gets. The job error analysis component then reports failures back through the portal.

What a reader should expect to see after a successful install is a reachable web portal with authentication, a cluster view showing the nodes and their accelerators, and the ability to submit one of the examples and watch it move through queued and running states. If the portal loads but no nodes appear, the problem is almost always in the cluster configuration rather than in the scheduling layer.

## OpenPAI is in stable mode, and that changes the calculus

The first paragraph of the README is a status notice, not marketing. It states that after the release of v1.8.1, OpenPAI entered stable mode with no major feature release planned, and that the repository was changed to read-only mode to save maintenance effort, with collaboration directed to the repository admin. The most recent release listed for the project is v1.8.0 from 2021-07-16, and the repository's last push was on 2026-09-08.

The practical reading is that OpenPAI is finished in the sense that its maintainers intended, not abandoned in the sense that nobody ever touched it again. The distinction matters when you evaluate it against a scheduler that ships new features every quarter. You are buying a design that was proven in Microsoft's own production environment, as the README puts it, and you are accepting that the feature set is frozen. For an on-premises platform that a company will run for five years, a frozen feature set is often acceptable and occasionally preferable. For a team that wants the newest gang-scheduling or topology-aware placement work, it is a dead end.

A second limitation is subtler. Because the heavy lifting lives in separate repositories, the read-only status of this repo does not automatically tell you the status of frameworkcontroller or hivedscheduler. Anyone evaluating OpenPAI should check those repositories independently rather than assuming the whole stack shares one maintenance posture.

## Where OpenPAI is the wrong tool

OpenPAI assumes you own or control the machines. The README's framing is on-premises first, with hybrid and public cloud as options, and the deployment tooling is oriented around configuring a cluster you administer. If your training runs are ephemeral and you want a provider to manage the nodes, adding OpenPAI in front of a managed Kubernetes service means you inherit a portal, an authentication layer and a scheduler to operate without gaining the hardware control that justifies them.

It is also a poor fit for small teams. A single researcher with two GPUs does not need user and group management, storage management, quota and a marketplace. The single-box deployment mentioned in the README exists, but deploying a full stack to schedule work on one machine is a lot of moving parts for a problem that a shell script solves.

Finally, consider the failure mode that the documentation is least explicit about: the boundary between OpenPAI and Kubernetes. Because OpenPAI layers on top of Kubernetes, debugging a stuck job can require understanding both the OpenPAI job state and the underlying Kubernetes pod state. The README does not document a rollback procedure for a failed cluster upgrade, and the captured text does not describe how to downgrade a deployment. That is a gap worth confirming with the documentation site before you upgrade a production cluster.

## OpenPAI compared with running Kubeflow or plain Kubernetes

The nearest comparison is Kubeflow, and the difference is where each puts its weight. Kubeflow is a set of machine learning tooling on Kubernetes, oriented around pipelines and notebook-centric workflows, and it expects you to bring your own scheduling and multi-tenancy story or assemble one. OpenPAI is the opposite emphasis: the README's four reasons to consider it are all about sharing hardware and assets across teams, and the architecture puts scheduling and runtime at the center, with hivedscheduler and frameworkcontroller as first-class components rather than add-ons.

A second alternative is simply Kubernetes with a GPU-aware scheduler and your own conventions. That gives you full control and no extra portal, at the cost of building the quota, job history and error analysis layers yourself. OpenPAI's job error analysis component is the piece teams most often underestimate when they plan to build their own.

The honest summary is that OpenPAI is a platform, and platforms trade flexibility for a working default. If your organization's bottleneck is that researchers cannot get a GPU without filing a ticket, OpenPAI addresses that directly. If your bottleneck is that your pipelines need a specific orchestration model, Kubeflow or a custom stack is the better starting point.

## Licence, upgrade cost and what to check before adopting

OpenPAI is released under the MIT licence, and the repository carries both a LICENSE and a NOTICE.txt file. MIT is permissive: it allows use, modification and redistribution with the copyright notice preserved. The NOTICE file exists because the project bundles or depends on other work, and anyone redistributing OpenPAI as part of a product should read it rather than assume the MIT terms cover every file. This is not legal advice; if you plan to ship OpenPAI inside a commercial offering, have counsel review the NOTICE alongside the dependency tree.

The upgrade story is the weakest part of the picture. Releases listed for the project stop at v1.8.0 from 2021-07-16, and the README states that v1.8.1 was followed by stable mode and read-only status. That means there is no stream of patch releases to pick up security fixes in the platform layer, and any fixes you need land on your own fork. The repository did receive a push on 2026-09-08, so the code is not frozen in a literal sense, but the README's statement about maintenance intent is the thing to plan around.

Before adopting, verify that the Kubernetes version your infrastructure runs is supported by the deployment assets under deployment/, that your shared storage can be presented the way the cluster configuration expects, and that the standalone components you rely on are still available from their own repositories. Those three checks cover most of the ways an OpenPAI rollout stalls.

## Conclusion

Adopt OpenPAI if you run your own GPU or FPGA hardware and need a single control plane for quotas, job queuing and a web portal that data scientists can use without touching kubectl. Do not adopt it if you need features merged after the v1.8.1 release, since the README states the repository was switched to read-only mode with no major feature release planned, or if your workloads are already well served by plain Kubernetes with a scheduler you trust. Before committing, verify three things: that the deployment path in deployment/ matches your Kubernetes version, that your storage layer can be mounted the way the cluster configuration expects, and that the standalone components you depend on (frameworkcontroller, hivedscheduler, openpai-runtime) are still reachable from their own repositories, because OpenPAI's read-only status applies to this repo, not necessarily to those.

## FAQ

### Is OpenPAI still maintained?

The README states that after the v1.8.1 release OpenPAI entered stable mode with no major feature release planned, and that the repository was changed to read-only mode to save maintenance effort. The last push to the repository was on 2026-09-08, and the most recent release listed is v1.8.0 from 2021-07-16.

### What is OpenPAI used for?

The README describes it as a platform for sharing AI computing resources such as a GPU or FPGA farm among teams, sharing and reusing assets like models, data and environments, giving IT a single operations platform for AI, and running a complete training pipeline in one place.

### What licence does OpenPAI use?

OpenPAI is released under the MIT licence, and the repository includes both a LICENSE file and a NOTICE.txt file that covers bundled or dependent work.

## Sources

- [License: MIT](https://github.com/microsoft/pai/blob/master/LICENSE)
- [microsoft/pai on GitHub](https://github.com/microsoft/pai)
- [Project website](https://openpai.readthedocs.io)
- [README](https://github.com/microsoft/pai/blob/master/README.md)
- [Releases](https://github.com/microsoft/pai/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/microsoft-pai
