PathML: a Python toolkit for whole-slide and spatial omics pathology pipelines
Tools for computational pathology
At a glance
- What is it?
- PathML is a GPL-2.0 research toolkit from Dana-Farber AIOS for loading, tiling and modelling pathology images. It covers slide formats, multiplex immunofluorescence and graph analysis, and its pinned dependency set is the first thing to check before adopting it.
- Who is it for?
- Adopt PathML if you work with whole-slide or multiplex images in Python and want a single library that covers slide reading, tiling, stain normalisation, graph construction and HoverNet-style training, and if the pinned dependency list in setup.py fits an environment you control.
- Can I use it commercially?
- Yes, with conditions. GPL-2.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository last received commits 47 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What PathML is for, and who actually needs it
PathML targets one specific bottleneck in cancer research: whole-slide and multiplex images are large, come in incompatible vendor formats, and need to be cut into tiles before any model can see them. The README frames the project around three themes, scalability, standardization and ease of use, and describes the goal as lowering the barrier to entry to digital pathology. That is a research framing, not a clinical one. The repository is maintained by Dana-Farber AIOS and the package author field lists Jacob Rosenthal, Ryan Carelli et al. with [email protected] as contact.
The audience that benefits is a computational researcher who already writes Python and needs to go from a .svs or .tiff slide to a tensor, or from a CODEX or multiplex immunofluorescence experiment to a graph, without assembling five libraries by hand. If you only need to look at slides and draw regions, PathML is the wrong layer. If you need a regulated diagnostic pipeline, the GPL-2.0 licence and the research framing both point elsewhere.
How the pipeline is put together: slides in, tiles and graphs out
The repository layout shows the mechanism more clearly than the README prose does. The pathml/ package sits alongside a requirements/ directory with separate environment files, including environment_mac.yml and environment_docker.yml, and an examples/ directory that reads as a map of the supported workflows: loading_images_vignette.ipynb, workflow_HE_vignette.ipynb, stain_normalization.ipynb, multiplex_if.ipynb, codex.ipynb, tile_stitching.ipynb, construct_graphs.ipynb, Graph_Analysis_NSCLC.ipynb, train_hovernet.ipynb, train_hactnet.ipynb and InferenceOnnx_tutorial.ipynb.
Read together, those notebooks describe a data flow. Slide images are loaded through OpenSlide and Bio-Formats bindings, which is why openslide-python and python-bioformats appear in install_requires and why openjdk<=18.0 is a prerequisite. Tiles are produced from the loaded slide, normalised for stain variation, and either fed to a segmentation or classification model trained with PyTorch, or converted into a graph with networkx and torch-geometric for the NSCLC graph analysis example. Spatial omics data moves through anndata and scanpy, which is how the CODEX and multiplex examples connect image features to transcriptomic measurements. ONNX appears twice, as onnx and onnxruntime, with a dedicated inference tutorial, so trained models can be exported and run without the training stack.
The dependency pins are tight and deliberate: numpy>=1.26.4,<2, pandas<=2.1.4, scikit-image<=0.22.0, scanpy==1.9.6, torch==2.12.0, onnxruntime>=1.17.0,<1.18, networkx<=3.2.1. That is a snapshot of a working environment rather than a loose set of ranges, and it is the single most important fact for anyone planning to install PathML into an existing project.
Installing PathML and running a first slide through it
The README gives two paths. The fastest is the published Docker image, which starts a Jupyter server on port 8888. The maintainers recommend Micromamba for a local environment, with Miniconda as an alternative if you have a licence; the Dockerfile itself builds from continuumio/miniconda3 and creates a conda environment from requirements/environment_docker.yml.
Start with the container if you only want to evaluate the library.
docker pull pathml/pathml && docker run -it -p 8888:8888 pathml/pathmlAfter the container starts, the entrypoint script launches Jupyter and you reach it on port 8888. The examples/ directory is copied into the image at /opt/pathml/examples, so the vignettes are available inside the container without any extra download.
For a local install, the README uses Micromamba and pins the JDK, because Bio-Formats needs a JVM.
micromamba create -n pathml 'openjdk<=18.0' -c conda-forge python=3.9
micromamba activate pathml
pip install pathmlOn Linux the README also asks for system packages before either install method, since OpenSlide and the BLAS/LAPACK headers are not Python wheels.
sudo apt-get install openslide-tools g++ gcc libblas-dev liblapack-devOn macOS the equivalent is a single Brew install of openslide. On Windows the README offers vcpkg install openslide or pre-built OpenSlide binaries extracted to a directory such as C:\OpenSlide\.
Developers cloning the repository create the environment from the checked-in file instead, and the README notes that the default CUDA version in environment.yml is 11.6.
conda env create -f environment.yml
conda activate pathmlOnce the environment is active, the first real use is loading a slide. The loading_images_vignette.ipynb example in examples/ is the intended entry point, and the H&E workflow vignette, workflow_HE_vignette.ipynb, shows the path from a loaded slide to tiles. The README does not print the exact API call for slide loading in the text quoted here, so follow the notebook rather than guessing at a constructor signature.
Where PathML will fight you: pinned dependencies and a JVM
The install_requires list is the main practical constraint. PathML pins torch==2.12.0, scanpy==1.9.6, pydicom==3.0.2, h5py==3.10.0, opencv-contrib-python==4.8.1.78, python-bioformats==4.1.0, python-javabridge==4.0.4 and onnxruntime below 1.18, and it holds numpy below 2.0 and pandas at or below 2.1.4. Dropping PathML into an existing environment that already has a different torch build or a numpy 2.x stack will force resolution conflicts, and the README's suggested mitigation is to update Conda and switch to the libmamba solver rather than to relax the pins. There is no documented compatibility matrix beyond the environment files themselves.
The second constraint is the JVM. Bio-Formats support comes through python-javabridge and jpype1, and the README pins openjdk<=18.0 in both the Micromamba and Conda instructions. That is a hard ceiling on the JDK version, and it means PathML carries a Java runtime dependency that a pure-Python image library would not.
The third is scope. PathML is a research toolkit published under GPL-2.0. The README does not describe a validation or regulatory pathway, and nothing in the repository suggests it is intended for clinical decision making. If your requirement is a diagnostic-grade viewer with audit trails, look at tools built for that purpose instead of treating PathML as one.
PathML versus QuPath, and why the comparison is not a ranking
QuPath is the alternative most people reach for, and the difference is architectural rather than a matter of features. QuPath is an interactive desktop application with a scripting layer: you open a slide, draw annotations, and write Groovy scripts against the objects you have drawn. PathML is a Python library with no viewer. There is no slide canvas in PathML, no annotation tool, and no project file format; the image handling happens inside a script or a notebook.
That split determines which one fits. If your work is exploratory, if a pathologist needs to inspect and correct regions, or if the deliverable is an annotated slide collection, QuPath is the natural home. If your work is a training pipeline, if tiles need to flow into PyTorch, ONNX export, graph construction with networkx and torch-geometric, or into anndata and scanpy for spatial omics, PathML sits where QuPath does not. The two are often used together in practice, with QuPath for inspection and PathML for the modelling stage, but the repository does not document an integration between them, so treat that as a workflow you would build yourself.
Maintenance, licensing and what an upgrade costs
The repository is not archived, and the last push was on 2026-08-14. Releases have been frequent through 2026: v3.0.5 on 2026-03-24, v3.0.6 on 2026-04-14 and v3.0.7 on 2026-07-09. The project also ships a CITATION.cff, which tells you the maintainers expect academic use and want it cited.
The licence is GPL-2.0, and setup.py carries the matching classifier. For research code that is unremarkable. For a company that wants to embed PathML inside a proprietary product, copyleft is a real consideration, and the repository offers no separate commercial licence or exception. That is a question for your own legal review, not something the README answers.
Upgrade cost is dominated by the pinned set. Moving between PathML releases means moving the whole environment, because torch, scanpy, onnxruntime and numpy are pinned together. The README documents installation and environment creation but does not document a rollback procedure or a migration guide between major versions, so plan on rebuilding the environment from requirements/ rather than upgrading in place.
Editorial conclusion
Adopt PathML if you work with whole-slide or multiplex images in Python and want a single library that covers slide reading, tiling, stain normalisation, graph construction and HoverNet-style training, and if the pinned dependency list in setup.py fits an environment you control. Do not adopt it as a viewer, an annotation tool or a clinical reporting system; it is a research toolkit, not a diagnostic product, and the README does not document rollback or migration for the pinned packages. Verify first that openjdk<=18.0 and the pinned torch, onnxruntime and scanpy versions resolve in your environment, then run the published Docker image and open one slide from your own cohort before rewriting any existing pipeline around it.
Frequently asked questions
What is the fastest way to try PathML without installing it locally?
The README gives a two-command Docker route: pull pathml/pathml and run it with port 8888 mapped, which starts a Jupyter server with the examples directory already copied into the image at /opt/pathml/examples.
Does PathML require Java or a specific JDK version?
Yes. Bio-Formats support comes through python-bioformats, python-javabridge and jpype1, and the README pins openjdk<=18.0 in both the Micromamba and Conda environment creation commands.
Can PathML be used for spatial transcriptomics and multiplex imaging?
The examples directory includes multiplex_if.ipynb and codex.ipynb, and the dependency list includes anndata and scanpy, which is how image features connect to spatial omics data in the documented workflows.
What licence does PathML use?
PathML is published under GPL-2.0, and setup.py carries the matching OSI classifier. The repository does not mention a separate commercial licence.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/dana-farber-aios-pathml)