torchxrayvision: 18 outputs per model, and which ones are trained
TorchXRayVision: A library of chest X-ray datasets and models. Classifiers, segmentation, and autoencoders.
At a glance
- What is it?
- A library that puts many chest X-ray datasets behind one interface and ships classifiers, autoencoders and anatomical segmentation models on top. The single most important line in its documentation is a warning: every pretrained model returns 18 outputs, and for weights other than the all-trained one, the targets that were not in the training set predict randomly. The only way to know which is to read the corresponding dataset's own pathology list.
- Who is it for?
- torchxrayVision is built for two situations, and it does both well: a clinical group that needs working feature extractors rather than training runs, and an algorithm group that needs to evaluate across datasets whose metadata conventions differ. Two things to check before you build on it.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Eighteen outputs, and the trained ones are listed elsewhere
The models section contains a note that deserves to be read twice by anyone loading a checkpoint.
Every pretrained model has eighteen outputs. One weights identifier, the all-trained one, has every output trained. For the other weights, some targets are not trained, and the page says plainly that they will predict randomly, because they do not exist in the training dataset.
So the failure mode is silent and looks like a result. Load a model trained on one source and it returns a full dictionary of eighteen plausible-looking probabilities, and the ones for diseases that source never labelled are noise with a decimal point in front of it.
The fix is documented and it costs one lookup: the only valid outputs are the ones listed in the pathologies field of the dataset that corresponds to those weights. In other words the checkpoint does not know what it knows, and the dataset object is the authority on which of the eighteen to trust.
The sample output makes the shape concrete. A processed image returns a dictionary keyed by pathology names, from atelectasis and cardiomegaly through consolidation, edema, effusion, emphysema, an enlarged cardiomediastinum entry, fibrosis, fracture, hernia, infiltration, lung lesion, lung opacity, mass and a nodule entry, with the highest value in that particular sample being cardiomegaly.
Labels are one, zero or NaN, and the third case carries the meaning
The dataset interface is three fields, and the third value in one of them is the design decision that makes the library usable across sources.
The first field is the list of pathologies a dataset contains, which is also the set of keys the labels field will use. The second holds, for each of those, either a one, a zero, or not-a-number.
That third option is the point. A missing label and a negative label are different facts, and a library that collapsed them would silently train and evaluate on assumptions its sources do not support. Keeping them distinct is what lets a dataset that marks findings as uncertain be merged with one that does not, and why the page recommends merging and filtering datasets to construct specific distributional shifts.
The third field is the metadata as a dataframe. It is the raw metadata file that comes with the data, with rows aligned to dataset elements so positional indexing works. Where possible the shared columns are aligned across datasets, with a patient identifier named as the first common field.
The page also states when these fields stay consistent: they are maintained when the subset and merge dataset wrappers are used, which is what makes filtering and merging safe operations rather than manual surgery.
The preprocessing chain is a fixed window, not a resize
The getting started example is five lines, and each one is a decision you have to keep.
An image is read, then normalized against the value 255, which converts an 8-bit image into a range running from minus 1024 to plus 1024. That is not a conventional normalization and it is the single most important line in the example, because it maps ordinary pixel data onto the scale a radiograph is usually displayed at.
Next the colour channels are averaged and a single channel is added, so the model receives one plane where the file had three. Then a transform is composed from two pieces: a center crop that knows where the chest is in a radiograph, and a resizer that produces the 224-pixel input the weights expect.
Finally the result is wrapped as a tensor and passed to the model, with the page noting that the feature-only path is available as well, so the same network can be used as an extractor without its classification head.
The consequence for reuse is worth stating plainly: the pipeline is a contract, not a convenience. An image normalized into a different range still runs, still returns eighteen numbers, and no part of the library will tell you the result is meaningless.
Segmentation returns fourteen regions at 512, classifiers return eighteen labels
The library exposes three kinds of model, and their output shapes are worth comparing because they encode different assumptions.
The classifiers all share one architecture family, described as pre-trained models currently all on one backbone, with the weights distinguished by training source. The sample identifiers name that source directly: the pneumonia challenge set, a national chest X-ray collection, a Spanish chest dataset from a university, a research benchmark from a US university, and two variants of a critical care dataset from an academic medical centre.
Segmentation is a different code path. It loads through a baseline models namespace and returns a tensor whose shape is one batch, fourteen channels, and five hundred and twelve pixels square. The fourteen channels are named anatomical regions rather than diseases: left and right clavicle, left and right scapula, left and right lung, left and right hilum, heart, aorta, a diaphragmatic surface, mediastinum, a structure the list spells as weasand, and spine.
Autoencoders are the third kind, and they are the only ones that reconstruct. They load through their own namespace with a weights identifier naming an elastic variant, and the usage is a two-call round trip that encodes an image to a latent vector and decodes it back.
Two benchmark documents are referenced rather than reproduced: one in the repository for the models, and a paper for the performance of some of them.
Autoencoders are trained on four datasets and expose one weights name
The autoencoder section is short, and the details are concentrated in two sentences.
A pre-trained autoencoder can be loaded, and it is trained on four datasets, named as the Spanish chest set, the national collection, the research benchmark and the critical care set. The API is two calls on the same object, one to encode an image into a vector and one to decode a vector back into an image.
The weights are given as a single identifier with a layer count and an elastic variant name. That is worth contrasting with the classifiers, where the identifier's whole job is to say which dataset trained the network, and where seven or more identifiers are listed.
The reason is structural rather than accidental. An autoencoder does not predict pathologies, so it does not inherit the eighteen-output contract, and it does not need to be re-trained per source to be useful. It needs to be trained on enough data to learn the appearance of radiographs, and four large sources will do.
Which means the risk profile of the autoencoder is different too. A classifier loaded from the wrong weights gives you noise in specific fields, and you can check that by reading the dataset. An autoencoder loaded from the wrong weights still produces a plausible image, and the page offers no field list to check it against.
Packaging is a setup script with dependency floors from the last decade
The packaging metadata is worth reading next to the dependency file, because they say different things about how old this library expects its stack to be.
There is a setup script and no project file, so the version is not static. It is read out of a private version file inside the package with a regular expression, and the script raises if the pattern is not found. The long description is read from the readme at build time, and the dependency list is read line by line from a requirements file rather than being declared inline.
The install itself is a single line:
$ pip install torchxrayvisionSo there is no extra step to fetch weights or datasets at install time, and no lockfile to reason about.
That dependency file uses floors with no upper bounds: a first-generation torch, a torchvision from before the modern releases, a scikit-image from 2019, numpy at one, pandas at one, a tqdm at four, a pillow from 5.3, requests at one, and imageio with no version at all.
Meanwhile the declared interpreter floor is Python 3.6, which predates several of those libraries' own support windows.
Read together, the intent is compatibility across a wide range rather than a validated environment. There is no lock on the upper side, so an install takes whatever is current, and the classifiers do not claim anything narrower than the Apache license line and an independent operating system.
The metadata does not assert a license and the packaging declares one
There is a licensing inconsistency in this repository, and it is the kind that is easy to miss and awkward to resolve later.
The license field recorded for the repository reports no assertion, meaning the detector could not classify it. The packaging script declares a classifier for the OSI-approved Apache license. And a license file sits in the repository root.
Those three facts point in different directions, and nothing in the visible documentation resolves them. The honest statement is that the repository ships a license file and its packaging claims Apache, while the recorded license metadata is unresolved, and anyone building on this should read the license file and the citation file rather than trusting the package listing.
The rest of the repository layout is worth noting because it explains how a research library stays usable. There is a benchmarks document, a citation file, a contributing guide, a documentation directory, a scripts directory holding the demo notebooks and the sample processing script, a tests directory, and a development requirements file.
Two deployment shell scripts and a style script sit at the root, along with a top-level package init file that is not part of the installed package, which is a small sign that the repository root and the distributed package are not the same surface.
Editorial conclusion
torchxrayVision is built for two situations, and it does both well: a clinical group that needs working feature extractors rather than training runs, and an algorithm group that needs to evaluate across datasets whose metadata conventions differ. Two things to check before you build on it. Read the pathology list for the weights you plan to load, because the eighteen numbers a model returns are not eighteen trained outputs, and the untrained ones are random rather than zero. And decide whether you want the pinned floor or the current stack, since the packaging declares support from a Python version old enough that its dependency floors are equally old and unbounded. On the legal side, the repository's own metadata does not assert a license while its packaging declares one, so check the license file rather than the package listing before you build a product on it.
Frequently asked questions
What does torchxrayvision provide?
A library for chest X-ray datasets and pre-trained models. It gives a common interface and a common pre-processing chain across a wide set of publicly available chest X-ray datasets, plus classification and representation learning models trained on different data combinations that can be used as baselines or feature extractors, along with pre-trained autoencoders and anatomical segmentation models.
Do all eighteen outputs of a torchxrayvision model mean something?
No. The page states that every pretrained model has 18 outputs and that the all-trained weights have all of them trained, while for other weights some targets are not trained and will predict randomly because they do not exist in the training dataset. The valid outputs for a given weights file are the pathologies field of the dataset that corresponds to those weights.
How do I preprocess an image for torchxrayvision?
Normalize it against 255, which converts an 8-bit image into a range of minus 1024 to plus 1024, average the colour channels down to one, then compose a center crop with a resizer to 224 before wrapping the result as a tensor and passing it to the model. The feature-only path is available on the same model object.
Which datasets does torchxrayvision cover?
A wide set of publicly available chest X-ray datasets behind a uniform interface, with classifier weights trained on sources that include the RSNA Pneumonia Challenge, a national chest X-ray collection, PadChest, CheXpert and MIMIC-CXR. Its pre-trained autoencoder is trained on the PadChest, NIH, CheXpert and MIMIC datasets.
What license is torchxrayvision released under?
The repository's recorded license metadata reports no assertion, while the packaging script declares a classifier for the OSI-approved Apache license and a license file is present at the repository root. The visible documentation does not resolve which of those governs, so read the license file directly.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mlmed-torchxrayvision)