Self-hosted service
run-house/kubetorch avatar
run-house/kubetorch

kubetorch states its iteration speed twice with two different numbers, and installs as a client extra plus a Helm chart you keep in step by hand

Distribute and run AI workloads on Kubernetes magically in Python, like PyTorch for ML infra.

1,228 stars61 forksPythonApache-2.0

At a glance

What is it?
kubetorch is a Python interface for building and running machine learning workloads on Kubernetes, presented as a serverless-style abstraction over a cluster with no local runtime and no code serialization. The abstraction is real and the SDK is small. The documentation around it is where the value is thin: three headline performance numbers with no method, one of them given twice with different figures, and a two-part install whose two halves are version-coupled by hand.
Who is it for?
kubetorch is worth trying if you already run Kubernetes and want to dispatch ordinary Python functions to it without building an image or managing a job manifest, because the example shows the whole surface and the SDK installs as a single extra.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 133 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The iteration figure appears twice with two different values

Three headline numbers carry the whole value proposition, and one of them is stated twice with different figures. The opening section says the interface enables extremely fast iteration of one to two seconds. The capability list, a few lines later, says one hundred times faster iteration, from ten or more minutes down to one to three seconds. Both cannot be the same measurement. The other two claims are 50 percent or more in compute cost savings and 95 percent fewer production faults, attributed to bin-packing, dynamic scaling, and built-in fault handling with programmatic error recovery. None of the three names a baseline system, a workload beyond a general reference to reinforcement learning and distributed training, or a method for measuring it. There is no benchmark table, no reproduction instructions, and no indication whether the figures come from the authors' own deployments or from customers. A reader deciding whether to install this will find that the documentation's most persuasive sentences are also its least verifiable ones, and the two conflicting iteration figures make it harder to know which one to hold you to.

No code serialization is claimed at the same time as shipping a local function

The stated reason for the speed is architectural, and it is worth reading twice because it is in tension with the example. The documentation says there is no local runtime and no code serialization, and that this is why cluster compute can be reached from any Python environment, an editor, a notebook, a pipeline, or production code, as if it were a local process pool. Serialization is what normally makes remote Python execution expensive: you pickle a function and its closure graph, ship the bytes, and unpickle them in a fresh interpreter, which is where the seconds go. So the claim is that this step is gone. The hello world example then defines a function in the file, wraps it, points it at a compute object, and calls it, and the comment above that call says it sends a local function to freshly launched remote compute. A locally defined function has to reach the cluster somehow. Either the bytes are shipped by a mechanism the documentation does not name, or the definition is recovered from source the cluster can read, or the function runs somewhere closer than the cluster suggests. The three possibilities have very different latency characteristics and very different failure modes, and the documentation resolves none of them. That single unstated mechanism is the load-bearing question behind all three headline numbers.

The client and the chart install separately, and only the chart version is written down

Installation has two halves that a user must keep aligned by hand. The Python side is one command with an extra:

bash
pip install "kubetorch[client]"

The cluster side is a Helm chart, and the documentation offers two ways to install it, one pointing Helm straight at an OCI registry reference and one pulling the chart first and installing from the local directory:

bash
# Option 1: Install directly from OCI registry
helm upgrade --install kubetorch oci://ghcr.io/run-house/charts/kubetorch \
  --version 0.5.0 -n kubetorch --create-namespace

# Option 2: Download chart locally first
helm pull oci://ghcr.io/run-house/charts/kubetorch --version 0.5.0 --untar
helm upgrade --install kubetorch ./kubetorch -n kubetorch --create-namespace

Note what those commands do and do not pin. The chart is pinned twice, to the same version. The Python client is not pinned at all, so it installs whatever the index currently serves. That asymmetry is the practical risk: a fresh install gives you a current client against a chart frozen at the version the documentation was written for. It also means an air-gapped or reproducible setup has to pin the client by hand, because nothing in the documented commands does it. The chart version in the documentation matches the newest published tag, while the last push to the branch is dated 2026-05-29, roughly three months after that tag, so the documentation describes a release rather than the current state of the code.

The chart, the controller, and the base images moved into this repository from elsewhere

One line in the documentation explains the shape of the tree. This repository now includes the customer-facing deployment components that were previously split across internal and open source repositories. The source layout names five areas: a Python client directory for the SDK, a charts directory for the Helm chart, a services directory for the controller and data store sources, a release directory for the workload base images, and a release directory for release scripts and version sync. That consolidation is the right thing to do and it introduces a specific new hazard. Components that were versioned in separate repositories are now released together from one repository, which means their version numbers have to be reconciled with each other as well as with the tags. The presence of both a base images directory and a version sync script suggests that is happening deliberately. A separate version file at the repository root, alongside a release directory for release scripts and version sync, is the third place a version can live, and the documentation adds a fourth by hardcoding the chart version in both install commands. None of that is a design flaw so much as a description of what happens when a multi-repository deployment is folded into one, and it is the thing to check after an upgrade.

The repository points at a managed platform that you have to ask for

After the self-hosted instructions, the documentation has a section inviting readers to contact the company by email or through a chat workspace to try the project on a fully managed serverless platform. That tells you the commercial shape of the thing: the open source repository is the self-serve path, and the hosted version is gated behind a conversation. Nothing in the repository contradicts that, and Apache-2.0 licensing means the self-hosted path is genuinely open. One small detail in that section is worth noting because it is verifiable and slightly odd: the invite link used for the managed platform contact and the invite link given later as the community channel are different invite identifiers for what appears to be the same workspace. Two tokens means two invites, and one of them will have been rotated or revoked at some point without the documentation being updated. The rest of the closing material is three links, a documentation site with an API reference, an examples site, and the chat invite, plus a licence line and a credit to the company. The examples directory in the repository itself holds only a readme, so the working code lives off the repository.

The word serverless is in quotation marks in the project's own subtitle

The subtitle puts the word in quotes, which is the author's own hedge and a fair one. Nothing here is serverless in the sense the word normally carries: there is a Helm chart, a controller, a data store, base images for the workloads, and a cluster that must already exist. What the design removes is not the cluster but the per-job packaging step. Instead of building an image, writing a manifest, submitting a job, and tailing logs, you write a function and call it, and the framework handles placement, resource requests, log propagation, and fault reporting. The documentation makes three specific promises about that experience which are more testable than the headline numbers. Logs, exceptions, and hardware faults are propagated back to you in real time. A freshly launched compute is available to a call almost immediately. And the whole thing behaves like a local process pool. Those three are the ones to verify first, because a framework that hides packaging but swallows a hardware fault silently is worse than the manual path it replaces. The third bullet's claim that faults are reduced by 95 percent is the least verifiable of the set, since fault reduction depends entirely on what the retry logic is allowed to do.

The hello world asks for a tenth of a CPU, and the examples directory is empty

The worked example is the entire API surface, and it is worth reading closely because there is not much to it:

python
import kubetorch as kt

def hello_world():
    return "Hello from Kubetorch!"

if __name__ == "__main__":
    # Define your compute
    compute = kt.Compute(cpus=".1")

    # Send local function to freshly launched remote compute
    remote_hello = kt.fn(hello_world).to(compute)

    # Runs remotely on your Kubernetes cluster
    result = remote_hello()
    print(result)  # "Hello from Kubetorch!"

A compute object requests a tenth of a CPU and nothing else, no memory and no accelerator is mentioned, which tells you resources are expressed as fractions of a node's capacity rather than as absolute amounts. Whether fractional CPU is right depends on your scheduler and your cluster, and it is the kind of detail that needs to be configurable for anything beyond a demo. Two further observations. The compute is described as freshly launched, so the first call pays for a pod start, which is the opposite of the steady-state latency the headline numbers describe. And the examples directory in the repository contains a single readme, so the examples the documentation points to live on the company website rather than in version control. The repository itself is otherwise conventionally organised, with a code of conduct, contributing guidelines, a maintainers file, and a pre-commit configuration.

Editorial conclusion

kubetorch is worth trying if you already run Kubernetes and want to dispatch ordinary Python functions to it without building an image or managing a job manifest, because the example shows the whole surface and the SDK installs as a single extra. Treat the performance claims as marketing until you reproduce them, and note specifically that the iteration figure is stated twice with two different values while none of the three numbers names a baseline, a workload, or a measurement method. Before you adopt it, check four things. Whether the two install halves are actually in step, since the client package and the chart version are pinned independently and only the chart version appears in the documentation. What your Kubernetes cluster actually needs, since the chart pulls base images and a controller and a data store that the client package does not contain. Whether self-hosting is the path you want, since the project points people at a managed platform and a sales conversation for the hosted version. And what the version file in the repository root says, since a checked-in version plus release scripts plus documentation that hardcodes a number is three places where the version can disagree.

Frequently asked questions

What is kubetorch?

It is a Python interface for building and running machine learning workloads on Kubernetes, described as a fast, Pythonic serverless-style layer over a cluster. You define a compute object with fractional resources, wrap a local function, point it at that compute, and call it, with logs, exceptions, and hardware faults propagated back.

How do I install kubetorch?

In two halves. The Python client goes on with a pip install of the client extra, and the cluster side is a Helm chart installed from an OCI registry reference or pulled first and installed from a local directory. The chart version is pinned in both documented commands while the client version is not pinned at all.

How fast is iteration with kubetorch?

The documentation gives two different figures in two places, one to two seconds in the opening section and one to three seconds in the capability list, both as part of a one hundred times improvement claim. No baseline, workload, or measurement method is given for that or for the cost and fault-reduction claims that sit beside it.

Does kubetorch serialize my Python code?

The documentation says there is no local runtime and no code serialization, and presents that as the reason iteration is fast. The hello world example simultaneously sends a locally defined function to remote compute, so how the definition reaches the cluster is not explained.

Is there a hosted version of kubetorch?

Yes, and it is gated behind contact rather than a signup. A section invites readers to email or join a chat workspace to try a fully managed serverless platform, while the self-hosted Helm path in the repository remains open under the Apache-2.0 licence.

Where are the kubetorch examples?

Off the repository. The examples directory in this repository contains only a readme, and the documentation points to an examples site hosted by the company. The worked example in the readme is the hello world, which requests a tenth of a CPU and launches its compute fresh.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. run-house/kubetorch on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/run-house-kubetorch.svg)](https://hysenlabs.com/projects/run-house-kubetorch)