# TorchIO: medical image preprocessing and augmentation for PyTorch pipelines

> TorchIO is a Python library that turns medical volumes into patch-based PyTorch training data, with spatial metadata preserved across every transform. It fits research code that already uses PyTorch, and it is the wrong tool for DICOM archive management or 2D natural-image augmentation.

**TorchIO-project/torchio** — Medical imaging processing for AI applications.

- Repository: https://github.com/TorchIO-project/torchio
- Website: https://docs.torchio.org/
- Stars: 2,446 · Forks: 280
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/torchio-project-torchio

## The problem TorchIO solves: 3D medical volumes that survive augmentation

Standard image augmentation libraries assume 2D arrays of pixels with no physical meaning. A horizontal flip of a chest radiograph is a flip of a grid of numbers. A horizontal flip of a brain MRI is a flip of a grid of numbers plus a change in the patient's left-right anatomy, and if the affine matrix is not flipped with it, the image and its segmentation mask no longer agree with the scanner coordinate system. TorchIO exists to keep that bookkeeping correct while you apply random flips, affine transforms, elastic deformation, blur and noise.

The library targets researchers and engineers who train deep learning models on 3D medical images: MRI, CT, PET and similar modalities stored as NIfTI. The pyproject.toml describes the package as "Medical image preprocessing, augmentation, and patch-based training" and lists science and research as the intended audience. If your data is JPEGs of skin lesions, this is not the tool you want. If your data is a directory of NIfTI volumes with voxel spacing that varies between scanners, the metadata problem is the whole problem, and TorchIO is aimed squarely at it.

## How TorchIO represents a volume and its spatial metadata

The central object is a Subject, which groups one or more images that share geometry: a T1 scan, its segmentation label map, and possibly a PET scan from the same session. Each image carries its affine matrix, so a transform applied to the subject can update the affine alongside the voxel data. That is the mechanism that separates TorchIO from generic tensor augmentation.

Transforms are composed into pipelines. The README links to documentation pages for RandomFlip, RandomAffine, RandomElasticDeformation, RandomBlur and RandomNoise, and shows a Compose pipeline in its animated examples. Because the transforms operate on the subject rather than on a bare tensor, a single random affine can be applied consistently to image and label, and the label can be resampled with nearest-neighbour interpolation while the image uses linear interpolation.

Patch-based training is the other half of the architecture. Rather than loading whole volumes, which for a high-resolution CT can exceed available memory, the library supports extracting patches and sampling them during training. The dependencies list confirms the stack: torch, numpy, nibabel and simpleitk for I/O, einops and torch-interpol for tensor operations, jaxtyping for typed array annotations, and tyro for command-line configuration. The package is typed, according to its classifier list, which matters if you run a type checker over your training code.

## Installing TorchIO and running a first augmentation

TorchIO is published on PyPI and on conda-forge, and the README carries badges for both. The pyproject.toml requires Python 3.10 or newer and lists support through 3.14. Install it with pip:

```bash
pip install torchio
```

The conda-forge channel is the alternative if your environment is conda-managed. Note that torch is a declared dependency, so pip will pull a PyTorch build; if you need a specific CUDA variant, install torch first and then TorchIO.

For a first real use, the repository ships tutorials in the tutorials directory, linked from the README as Google Colab notebooks. Those notebooks are the reference for constructing a subject and composing transforms, and they run against real data without a local environment. The README itself documents the transform classes by name: RandomFlip, RandomAffine, RandomElasticDeformation, RandomBlur and RandomNoise, composed with Compose.

What you should watch for when you build that pipeline is how label maps are handled. If a segmentation comes out with interpolated grey values instead of integer class indices, the label was loaded as a scalar image rather than a label map, which is the most common first mistake.

## Where TorchIO stops: DICOM, viewers, and the v2 alpha

TorchIO reads and writes medical image formats through nibabel and SimpleITK. It is not a DICOM archive, it does not speak DICOMweb, and it has no query/retrieve layer. If your data still lives in a PACS, you need something else to get it out first, and the README does not claim otherwise.

There is no viewer. The README shows animated GIFs of transforms, but those are documentation artifacts, not an interactive tool. For visual inspection you will export to a NIfTI file and open it in a separate application.

The release situation deserves attention. The most recent releases are v2.0.0a2 and v2.0.0a1, both dated 2026-07-20, with v1.2.1 on 2026-06-02 as the previous stable tag. The alpha suffix means the 2.0 line is pre-release. The pyproject.toml in the repository declares version 2.0.0a2, so the main branch is tracking the alpha. If you need a stable dependency for a production training pipeline, pin to the 1.2.x line and read the release notes before moving. The README does not document a rollback or migration procedure between the 1.x and 2.0 lines, so that is something to establish from the changelog yourself.

Memory is the other practical boundary. Patch-based training exists because whole-volume loading does not scale, and the library gives you the sampling machinery rather than solving the memory problem for you. Choosing patch size and sampling strategy remains your responsibility.

## TorchIO versus MONAI for medical image pipelines

MONAI is the obvious comparison, and the difference is scope rather than quality. MONAI is a broader framework: it includes network architectures, losses, metrics, workflows, bundle packaging, and its own transform system for both 2D and 3D data. TorchIO is narrower. The pyproject.toml describes it as preprocessing, augmentation and patch-based training, and that is essentially the whole surface area.

That narrowness is the argument for it. If you already have a PyTorch training loop you like, and the only thing missing is geometry-aware augmentation and patch sampling, TorchIO drops in without asking you to adopt a framework. You keep your model, your optimizer, your logging. The Subject abstraction is the only new concept you have to learn.

If you want an end-to-end medical imaging framework with reference architectures and deployment bundles, MONAI covers ground TorchIO does not attempt. The two are not mutually exclusive, but mixing two transform systems in one pipeline means two sets of conventions for affine handling, and that is where subtle bugs live. Pick one for the geometry-sensitive stage.

## Licence, maintenance and what upgrading costs

TorchIO is licensed under Apache-2.0, declared both in the repository LICENSE file and in the pyproject.toml license field. Apache-2.0 is a permissive licence that includes an explicit patent grant, which is why it is common in research code that companies later want to use. It does not require you to open your own code. This is not legal advice; if your organisation has a licence review process, run it.

The repository is not archived, and the last push was on 2026-08-01. There is an active CI setup with test and documentation workflows, and the project uses pre-commit and ruff for code style. The maintainer is listed as Fernando Pérez-García in the pyproject.toml.

The upgrade cost is the 1.x to 2.0 transition. The dependencies changed: the 2.0 pyproject.toml lists einops, fsspec, jaxtyping, loguru, rich, torch-interpol and tyro, and requires Python 3.10 or newer. If you are on an older Python, upgrading TorchIO means upgrading your interpreter first. The presence of tyro suggests a command-line configuration layer that did not exist in the 1.x line, and the typed-array annotations via jaxtyping suggest the API surface is annotated more strictly. Expect to re-read the transform documentation rather than assume the 1.x calls still work, and pin your version explicitly in requirements.

## Conclusion

Adopt TorchIO if your training code is already PyTorch and your data is NIfTI or SimpleITK-readable, and you want augmentation that respects voxel spacing and affine geometry. Do not adopt it as a DICOM store, a viewer, or a 2D image augmentation library for photographs; it is built around 3D medical volumes. Before committing, verify that your installed Python is 3.10 or newer, that your preprocessing steps map onto the transforms documented at docs.torchio.org, and that the v2.0.0a2 release is acceptable for your project, since the stable line is v1.2.1.

## FAQ

### What is TorchIO?

TorchIO is a Python library for medical image preprocessing, augmentation and patch-based training, built on PyTorch. It represents volumes as Subjects that carry their affine geometry, so transforms applied for data augmentation keep image and label maps aligned.

### How does TorchIO compare with MONAI?

TorchIO is narrower: the pyproject.toml describes it as preprocessing, augmentation and patch-based training. MONAI is a broader framework covering architectures, losses, metrics and workflows. TorchIO is the lighter choice when you already have a PyTorch training loop and only need geometry-aware augmentation.

### How do you install TorchIO?

It is published on PyPI and conda-forge, and the README carries badges for both. The pyproject.toml requires Python 3.10 or newer and lists torch among its dependencies, so pip will install a PyTorch build alongside it.

### Does TorchIO read DICOM files?

TorchIO handles image I/O through nibabel and SimpleITK, and its documentation and README describe work with volumes such as NIfTI files. The README does not present it as a DICOM archive or a DICOMweb client, so retrieving studies from a PACS is outside what the material documents.

### Which TorchIO version should I use?

The most recent releases are v2.0.0a2 and v2.0.0a1 from 2026-07-20, both alpha pre-releases, with v1.2.1 from 2026-06-02 as the earlier stable tag. The repository pyproject.toml declares 2.0.0a2, so the main branch tracks the alpha; pin to 1.2.x if you need a stable dependency.

## Sources

- [License: Apache-2.0](https://github.com/TorchIO-project/torchio/blob/main/LICENSE)
- [Project website](https://docs.torchio.org/)
- [README](https://github.com/TorchIO-project/torchio/blob/main/README.md)
- [Releases](https://github.com/TorchIO-project/torchio/releases)
- [TorchIO-project/torchio on GitHub](https://github.com/TorchIO-project/torchio)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/torchio-project-torchio
