Self-hosted service
meta-pytorch/torchx avatar
meta-pytorch/torchx

TorchX: a universal job launcher for PyTorch applications

TorchX is a universal job launcher for PyTorch applications. TorchX is designed to have fast iteration time for training/research and support for E2E production ML pipelines when you're ready.

427 stars156 forksPythonNOASSERTION

At a glance

What is it?
TorchX wraps a training script in a named spec and hands it to Kubernetes, Slurm, Docker or the local machine. It is for teams that move the same job between a laptop and a cluster, and it is a launcher, not a training framework.
Who is it for?
Adopt TorchX if you already run PyTorch jobs on Kubernetes, Slurm or Docker and want one spec format that survives the move between them. Do not adopt it expecting a training loop, a checkpoint manager or a distributed launcher: the README states that PyTorch itself is a requirement and that Docker is optional, needed only for docker based schedulers.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap TorchX fills between a training script and a scheduler

A PyTorch training script usually runs one way on a laptop and another way on a cluster. The laptop invocation is python train.py with a few flags. The cluster invocation is a YAML manifest, or an sbatch script, or a docker run line, each with its own way of passing arguments, images and resource requests. The script does not change; the wrapper around it does, and the wrapper is where teams lose time.

TorchX puts a named spec in that gap. The README calls it a universal job launcher for PyTorch applications and lists the targets it currently supports: Kubernetes (EKS, GKE, AKS and similar), Slurm, Docker and Local. The intended audience is stated in the same paragraph: people who want fast iteration during training and research, and who want the same jobs to feed end to end production ML pipelines later.

The audience is therefore not everyone writing PyTorch. If you train on a single machine and never leave it, the Local scheduler adds a layer you do not need. If your jobs already run through a hand written sbatch template that your team understands, TorchX is a second abstraction to maintain. It pays off when the same job has to exist in more than one of those places.

How a spec becomes a running job

The unit TorchX works with is a component, and a component is turned into a spec. The spec describes the roles in the job, and a role carries the image, the entrypoint, the arguments and the resource request for one group of processes. The CLI command torchx run takes a component name, applies the component's arguments, and produces that spec.

The scheduler is the backend that consumes the spec. TorchX ships schedulers for the four targets named in the README, and the repository's own topic list mentions Airflow, AWS Batch, Kubernetes, Ray and Slurm, which reflects the breadth of the integrations rather than a single execution model. A scheduler is selected per run, so the same component can be pointed at Local during development and at a cluster later.

Two details of the layout are worth noting. The .torchxconfig and .torchxignore files at the top level suggest that a working directory can carry project level configuration and an ignore list, similar in spirit to how other tools treat a project file. The pyproject.toml also declares a torchx.tracker entry point group, which means tracker implementations are pluggable rather than hard coded, though the README does not describe what a tracker does.

The honest summary is that TorchX is a translation layer. It does not schedule anything itself, it does not manage checkpoints, and it does not replace torch.distributed. It builds a description and hands it to something that already knows how to run work.

Installing TorchX and running a first job

The README gives four installation lines. The minimum install is the SDK and CLI with no scheduler extras, and the extras add the dependencies for specific backends. If you intend to target Kubernetes, the kubernetes extra is the one that matters, because pyproject.toml lists kubernetes>=11 only under that extra.

bash
# minimum: SDK and CLI
pip install torchx

# Kubernetes / Volcano support
pip install "torchx[kubernetes]"

The README also documents a nightly channel and a source install. For development from a clone, the repository recommends uv:

bash
# from a clone of the repository
uv sync --extra dev

# or with pip
pip install -e ".[dev]"

Note that the README's requirements section says python3 (3.8+), while pyproject.toml sets requires-python to >=3.10. Where the two disagree, the packaging metadata is what a pip install will enforce, so plan for 3.10 or newer.

The README points to a quickstart guide at meta-pytorch.org/torchx/latest/quickstart.html for the first real run, and that is the correct next step: the README itself does not show a complete torchx run invocation with its arguments, so any command I wrote here would be a guess. The CLI entry point is declared in pyproject.toml as torchx = "torchx.cli.main:main", so after installation the torchx command is on the path. TorchX also publishes a container image at ghcr.io/pytorch/torchx, which the README describes as being for use as part of a TorchX role rather than as a general purpose development image.

Where TorchX is the wrong tool

The clearest limitation is scope. TorchX launches jobs; it does not train them. There is no optimizer, no data loader, no checkpointing policy and no fault tolerance for a job that dies halfway. If your problem is that a training run crashes at epoch 40 and you lose the work, TorchX is not the answer, and adding it will not change that.

Second, the project describes itself as Beta. The classifier in pyproject.toml is Development Status :: 4 - Beta, and the release history is sparse: v0.5.0 in April 2023, v0.6.0 in October 2023, v0.7.0 in July 2024, with no tagged release since. The repository's last push was on 2026-09-10, so work continues on main, but a team that pins to tagged releases is working with code from 2024. That is a real operational fact, not a criticism of the code.

Third, the documentation is uneven. The README covers installation thoroughly and then defers almost everything else to the docs site. It does not document rollback, it does not explain the tracker entry point, and it does not show a full run command. Anyone evaluating TorchX should expect to read the hosted documentation, not the repository front page.

Finally, if your organisation has standardised on exactly one scheduler and has no intention of moving, TorchX is an extra dependency between you and that scheduler. The value is portability across targets, and a single target does not need portability.

TorchX compared with torchrun and with writing your own launcher

The most common comparison is torchrun, and the difference is the layer each one occupies. torchrun starts distributed processes within a job: it sets up the rendezvous, assigns ranks and launches the workers on the machines you already have. TorchX decides where the job goes and describes it to a scheduler. They are not substitutes. A TorchX role can invoke a distributed entrypoint, and torchrun remains the thing that makes the processes talk to each other.

The second alternative is a hand written launcher: a shell script that calls kubectl apply, or an sbatch template with substitutions. That approach has no dependency and no abstraction to learn, and for one team on one cluster it is often sufficient. Its cost appears when you add a second target, because now you maintain two templates that drift apart. TorchX's claim is that one component definition replaces both.

A third alternative is a full pipeline orchestrator. TorchX integrates with Airflow and the README's own examples reference Hydra, Ax and PyTorch Lightning in the dev extras, which shows the project expects to sit inside a larger toolchain rather than replace it. If you need scheduling across many dependent steps, retries and a UI, the orchestrator is the right layer and TorchX is a component inside it.

Maintenance, licensing and what to check before adopting

The repository is not archived, and the last push was on 2026-09-10, so it is being worked on. The release cadence is the weaker signal: the newest tagged release listed is v0.7.0 from 2024-07-16. A team that needs tagged releases with a predictable upgrade path should check the changelog before planning, because the gap between the last tag and current main is where the recent work lives.

Upgrade cost is mostly dependency cost. The base install pulls docstring-parser, pyyaml, docker, filelock, fsspec, tabulate and typing-extensions. The fsspec floor is 2023.10.0, and the docker package is a base dependency even though the README calls Docker optional and needed only for docker based schedulers, which means a Kubernetes only user still installs it. The kubernetes extra adds kubernetes>=11. The dev extra is heavy: it pulls torch, torchvision, torchtext, pytorch-lightning, ax-platform[mysql], captum, hydra-core, s3fs and the lint tooling. Do not install [dev] in a production image.

On licensing, the README states that TorchX is BSD licensed as found in the LICENSE file, and pyproject.toml declares license = "BSD-3-Clause". The repository metadata reports the licence as NOASSERTION, which is a metadata classification rather than a statement about the terms. Read the LICENSE file itself before you rely on it; I am not giving legal advice, and the file is the authority, not the badge or the classifier.

Editorial conclusion

Adopt TorchX if you already run PyTorch jobs on Kubernetes, Slurm or Docker and want one spec format that survives the move between them. Do not adopt it expecting a training loop, a checkpoint manager or a distributed launcher: the README states that PyTorch itself is a requirement and that Docker is optional, needed only for docker based schedulers. Before committing, verify that the scheduler you actually use is in the supported list, that your Python version satisfies the requires-python >=3.10 floor in pyproject.toml, and that your own launcher code calls the torchx API rather than the CLI, because the repository still classifies the project as Development Status 4 - Beta.

Frequently asked questions

Which is better, PyTorch or Torch?

This question is about PyTorch and Torch, not about TorchX, so the README does not address it. TorchX is a job launcher for PyTorch applications and lists PyTorch as a requirement.

Is PyTorch free to use?

This question is about PyTorch rather than TorchX. For TorchX itself, the README states that it is BSD licensed as found in the LICENSE file.

How do I install PyTorch?

This question is about PyTorch, not TorchX. TorchX's requirements section lists PyTorch as a dependency and links to the PyTorch installation page for it.

Is PyTorch built on Python?

This question is about PyTorch rather than TorchX. TorchX itself is a Python project: pyproject.toml sets requires-python to >=3.10 and declares the torchx CLI entry point.

Official sources

  1. Issues
  2. meta-pytorch/torchx on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/meta-pytorch-torchx.svg)](https://hysenlabs.com/projects/meta-pytorch-torchx)