Crater is the layer between a shared GPU cluster and the people who share it, and the billing model is the feature
Crater is a cloud-native AI training & inference platform.
At a glance
- What is it?
- A Kubernetes-native control plane that adds accounts, quotas, approvals and cost visibility on top of an existing scheduler, so that a training job submitted by a student can be attributed, bounded and billed. Written in Go, with a web console, a CLI and agent skills.
- Who is it for?
- This is an infrastructure governance project, not a training framework, and reading it that way is the fastest way to be disappointed by it or to use it well. If you have one team and one cluster, none of this matters and Kubernetes plus the underlying scheduler is a better answer.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
It is a control plane, not a scheduler
The subtitle is the most useful line in the readme: a Kubernetes-native control plane for shared AI computing clusters. That word choice rules out the interpretation most people arrive with. Crater does not schedule your training jobs better than the batch scheduler underneath it does. It decides who is allowed to run them, how many, at what cost, and what happens when two teams want the same accelerator.
The platform description says this directly. It builds on Kubernetes and on a separate batch scheduling project, connecting users, accounts, queues, quotas, images, datasets, models, jobs, services and observability into one workflow. Every one of those nouns except the first two is an organisational concept rather than a technical one, and that is the shape of the entire product.
The justification section makes the gap concrete with a five-row comparison. Raw command-line access is easy to misuse. GPU usage is hard to attribute and hard to bound. Everyone rebuilds training and serving manifests by hand. Assets scatter across nodes. Operators and users debug from different tools. Each row is a governance problem that the two underlying projects genuinely do not solve, because they have no concept of a user.
Five audiences, and the same platform serving each differently
The designed-for table is the most useful thing in the repository for deciding whether this fits, because it maps five concrete scenarios to the workloads and the features that answer them.
Research and engineering gets long-running jobs, reusable environments, mounted data and models, logs, monitoring and lifecycle controls. Teaching and training gets account and quota management, job templates, burst handling, fair access and a web-based submission path, which is a genuinely different set of requirements because a classroom has hundreds of users who are not going to read documentation.
Model training and serving gets deployment templates, hardware-aware placement, managed assets, service access, and governance that separates training resources from serving resources. Enterprise services get managed runtime environments, service access, operational visibility and resource governance. Data processing gets storage integration, asset management, schedulable batch jobs and observability.
Read as a list, that is five products with one backend. Read as a design constraint, it explains the multi-tenant governance feature existing at all: the teaching scenario and the enterprise scenario have opposite requirements, and a platform that claims both has to make its accounting model flexible enough for both.
Queues, quotas and approvals are the load-bearing features
Of the eight feature blocks, the two that would be hardest to build later are the multi-tenant governance one and the scheduling one.
Governance covers users, accounts, queues, quotas, approvals and billing-oriented resource visibility, and the stated goal is to turn a raw cluster into an accountable shared service. Accountable is the operative word. A quota you can exceed is a suggestion, and a cost figure you can only reconstruct from a billing system after the fact is not visibility. Both of those are the difference between a cluster you can share and a cluster you have to ration.
Scheduling sits on the existing scheduler and adds queue-based admission, priority-aware execution, pre-queue policies, and placement across heterogeneous resources, explicitly including mixed training and serving workloads. Pre-queue admission control is the part worth understanding: it means a job that cannot run right now is refused with a reason rather than queued indefinitely, which is what stops one team's interactive notebook work from sitting behind a week-long training job.
The mixed training and serving claim is the load-bearing one. A training job wants the whole card for hours; a serving job wants a slice of it for milliseconds. Putting both on the same cluster without policy is how a research cluster becomes unusable for its own researchers.
Interactive environments are the feature researchers actually adopt
Alongside the governance features there is a block for interactive development that is worth reading on its own, because in most research clusters this is the feature that decides whether people use the platform or route around it.
It provides containerised notebooks, a web development environment, web terminals, direct shell access and custom environments, none of which require the user to know how to write a manifest or how the cluster is configured. The stated benefit is a reproducible workspace close to the data and the accelerators, and that is the correct framing: the value is not the notebook, it is that the notebook is already attached to the right storage and the right hardware.
The workload lifecycle block covers the other half: submit, clone, monitor, stop and inspect, with reusable templates spanning interactive sessions, fine-tuning and long batch jobs. Cloning is a small feature that matters disproportionately in research, where the third run of an experiment is almost always the first run with one thing changed.
Taken together, these two blocks are the adoption strategy. Governance is what a platform team needs and users do not care about. A notebook launcher is what a researcher needs on day one and will route around if it is missing. Build the second first, and the first gets paid for later.
Heterogeneous accelerators, and an admission that it is hard
The accelerator support block is the one to read with the most care, because it is where the claims are broadest. It represents graphics processors and other accelerator models as schedulable resources, and it names four specific things it supports: the dominant vendor's cards, domestically produced accelerator cards, virtualised processor resources, and device integration based on the newer dynamic resource allocation mechanism.
The first is unremarkable. The fourth is the one to understand: the dynamic resource allocation path is the modern replacement for advertising devices as opaque strings in a resource request, and it is where the next few years of device plugin work is going. Supporting it means the platform is written against the direction rather than against the current state.
The second is the reason this project exists in the shape it does. Supporting a second accelerator vendor at the scheduler level means abstracting over the device rather than assuming one vendor's model, and that abstraction is expensive and mostly invisible until you need it. Every research platform in this space either supports one vendor properly or several vendors approximately.
The honest caveat is the issue count: roughly seventy-eight open issues on a project with this much surface area. Not all of them will be in the device layer, but a heterogeneous-accelerator feature is exactly the kind of thing that accumulates issues per vendor, per driver version and per cluster configuration. Read the issue list before you commit to running a second vendor's hardware on it.
Four interfaces, including one aimed at agents
The last feature block is about how you talk to the platform, and it names four ways: a web console, a command-line interface, HTTP interfaces, and agent-oriented command skills.
Four interfaces is a considered choice rather than feature creep, because each audience arrives through a different one. An administrator configures through the console. A researcher submits through the command line. An existing platform integrates through the interfaces. And the skills block is for automation, scripted workflows and agent-driven operation, which is a category that barely existed two years ago and which every infrastructure tool is now adding whether or not it has done the work.
What makes the agent angle interesting here is that the readme lists AI-assisted operations twice. Once as a headline capability, and once in the comparison table as the answer to operators and users debugging from different tools. That second framing is the better one: the point is not that an assistant can summarise your dashboards, it is that the person who ran the job and the person who owns the cluster can ask the same system the same question.
The repository structure matches the interface split. There is a backend, a frontend, a command-line tool, a Helm chart, a documentation directory, a directory of dashboards, a hack directory, a website, and a skills directory that is separate from the rest, which is how you would organise skills if they were meant to be consumed by something other than this repository.
A monorepo that installs its own git hooks
The build file is Go-adjacent in the way a polyglot monorepo is: four components, a chart, a website and a skills directory, coordinated by one makefile.
The first thing it does is parse itself. The help target extracts category headers and the commands beneath them with an embedded script, so the command list is generated from the file rather than maintained separately. Every target carries a comment that becomes its description. It is a small thing and it is the difference between documentation that is right and documentation that was right once.
Then it sets up git hooks, which is more interesting than it sounds. Rather than asking you to copy a hook file by hand, the makefile locates the hooks directory through the git executable itself, falling back to a conventional path, creates it if it does not exist, copies the pre-commit hook in, and marks it executable. There is also a target that runs the hook without installing it.
That is a solve for the standard failure of commit hooks: they work on the machine of whoever set them up and silently do nothing on everyone else's. Asking the build tool to install its own hooks, using git's own answer for where they belong, means the check runs for a contributor who cloned the repository yesterday. Everything else in the file follows the same pattern, which is what you expect from a project where the engineering time went into the operator experience.
Editorial conclusion
This is an infrastructure governance project, not a training framework, and reading it that way is the fastest way to be disappointed by it or to use it well. If you have one team and one cluster, none of this matters and Kubernetes plus the underlying scheduler is a better answer. If you have six teams and a grant that requires you to show where the money went, then attribution and quotas are the whole product, and they are not things either of the underlying tools will do for you. The practical caution is the open issue count: high for a project with this much surface area, and the heterogeneous accelerator support and the device-integration path are where you should look before committing.
Frequently asked questions
What is Crater?
A Kubernetes-native control plane for operating a shared AI computing cluster. It sits on top of Kubernetes and a batch scheduler and adds the organisational layer neither provides: users, accounts, queues, quotas, approvals, images, datasets, models, jobs, services and observability in one workflow, exposed through a web console, a command-line tool, HTTP interfaces and agent skills.
Is Crater a replacement for Kubernetes?
No, and the readme says so: it is a control plane built on Kubernetes and on a separate batch scheduler. It does not claim to schedule better than those do. What it adds is everything that requires a concept of a user, which is what makes a cluster shareable between teams rather than merely operable by one.
Which GPU types does Crater support?
It represents accelerators as schedulable resources and names four categories: the dominant vendor's cards, domestically produced accelerator cards, virtualised processor resources, and device integration through the newer dynamic resource allocation mechanism. The last is the forward-looking one, and the issue list is worth reading before relying on the second.
How does Crater handle a team asking for more GPUs than there are?
With queue-based admission, quotas, approvals, priority-aware execution and pre-queue policies. Quotas are per account, and pre-queue admission control means a job that cannot run right now is refused with a reason instead of waiting indefinitely behind work that will take a week.
Can researchers get a notebook without writing manifests?
Yes. The interactive development feature provides containerised notebooks, a web development environment, web terminals, direct shell access and custom environments, none of which require the user to know how the cluster is configured. Job lifecycle covers submitting, cloning, monitoring and stopping, with reusable templates.
Who is Crater built for?
Shared clusters in universities, research institutes, enterprise AI teams and internal platform teams. The readme maps five scenarios to workloads and features: research and engineering, teaching and training, model training and serving, enterprise services, and data processing. Teaching has different requirements from enterprise services, which is a large part of why the accounting model is built the way it is.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/raids-lab-crater)