PyHealth 2.0: A Deep Learning Toolkit for Clinical Healthcare AI
A Deep Learning Python Toolkit for Healthcare Applications.
At a glance
- What is it?
- PyHealth is a Python toolkit for clinical predictive modeling that provides a modular five-stage pipeline, 33 or more pre-built models, and native support for healthcare datasets including MIMIC-III, MIMIC-IV, and eICU. Version 2.0 requires Python 3.12 or 3.13 and installs from PyPI with a single pip command.
- Who is it for?
- PyHealth 2.0 is well-suited for ML researchers who need a reproducible baseline for clinical predictive modeling tasks, and for healthcare data scientists who want to run deep learning experiments on MIMIC or eICU without writing all the preprocessing and model scaffolding themselves. It is not the right choice for teams that need Python 3.9 to 3.11 support, since version 2.0 requires 3.12 or 3.13.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Who PyHealth Is For and What Problem It Addresses
Clinical deep learning research is slowed by the absence of shared infrastructure. Each research group typically writes its own data loaders for MIMIC-III or MIMIC-IV, its own medical code converters, and its own training loop, making results difficult to reproduce and compare. PyHealth addresses this by providing a common toolkit that both ML researchers and medical practitioners can use to develop, test, and deploy healthcare AI applications without rebuilding that infrastructure from scratch.
The README describes the target audience explicitly as both ML researchers and medical practitioners, positioning the library at the intersection of technical flexibility and domain specificity. Version 2.0 was accompanied by a paper submitted to arXiv in January 2026, titled PyHealth 2.0: A Comprehensive Open-Source Toolkit for Accessible and Reproducible Clinical Deep Learning. The earlier 1.0 version was presented at KDD 2023.
The Five-Stage Modular Pipeline
PyHealth organizes the healthcare ML workflow into five stages: dataset loading, task definition, model selection, training, and evaluation. Each stage can be used independently, so a team that already has preprocessed tensors can plug in at the model stage without running the dataset loaders.
The healthcare-first design is most visible in the dataset stage. PyHealth ships native loaders for MIMIC-III, MIMIC-IV, eICU, and OMOP-formatted datasets. These loaders understand medical coding systems, handle the hierarchical structure of clinical visits, and convert ICD codes, medication codes, and lab event sequences into formats suitable for the included models.
The task layer maps raw clinical data to specific prediction problems. The README lists 10 or more supported healthcare tasks. The model layer provides 33 or more pre-built implementations, covering architectures for sequential EHR data, graph-based patient representations, and transformer-based models.
The README states that data processing runs roughly three times faster than a comparable pandas-based pipeline. This is attributed to the version 2.0 additions: parallelized processing, dynamic memory support, and multimodal dataloaders.
Installation and First Steps
PyHealth 2.0 requires Python 3.12 or 3.13. The README states this explicitly: version 2.0 dropped support for earlier Python releases to take advantage of parallel processing and memory management improvements available in those versions.
The recommended installation uses pip:
pip install pyhealthThis installs the latest PyHealth 2.0 release and automatically pulls PyTorch and other deep learning dependencies. If backward compatibility with older code is needed, the previous stable version is available:
pip install pyhealth==1.16Version 1.16 supports Python 3.9 and later but does not include the 2.0 performance improvements. For contributors or developers who need the latest unreleased features, the source installation path is:
git clone https://github.com/sunlabuiuc/PyHealth.git
cd PyHealth
pip install -e .The project uses pixi for environment management (a pixi.lock file is present in the repository root), though pip remains the recommended route for most users. After installation, tutorials are available as Jupyter notebooks in the examples/ directory and on Google Colab via links in the documentation.
Supported Datasets, Models, and Tasks
The dataset support centers on the four major open clinical data sources: MIMIC-III (ICU records from Beth Israel Deaconess Medical Center), MIMIC-IV (the updated version of the same database), eICU (a multicenter ICU database), and OMOP CDM (a standardized clinical data model used by a range of hospital systems). The examples/ directory contains notebooks demonstrating dataset loading for MIMIC-III demo data (examples/event_timestamps_mimic3demo.py) and MIMIC-IV (examples/cnn_mimic4.ipynb, examples/gat_mimic4.ipynb, and others).
The model catalogue spans CNN, GAT, GCN, ConCare, DeepR, EHRMamba, HALO, and GraphCare, among others. These are visible from the examples directory names. The pyproject.toml lists graph model support as an optional dependency (torch-geometric) and NLP support (rapidfuzz, rouge_score, nltk) as another optional group.
Supported healthcare prediction tasks include drug recommendation and cardiology detection, as visible from examples/drug_recommendation/ and examples/cardiology_detection_isAR_SparcNet.py. The full task list is documented in the pyhealth.readthedocs.io documentation rather than the README.
Limitations and Cases Where PyHealth Is Not the Right Choice
The Python 3.12 or 3.13 requirement is a hard constraint for version 2.0. Any environment pinned to Python 3.9, 3.10, or 3.11 cannot use the current release without a Python upgrade. The older 1.16 version supports Python 3.9 but lacks the 2.0 performance and feature improvements.
PyHealth is built around clinical EHR datasets in the MIMIC, eICU, and OMOP formats. Teams working with non-EHR healthcare data, such as medical imaging without attached clinical records, radiology report text without structured codes, or genomics data, would find the pipeline less applicable. The library's medical code handling assumes ICD, ATC, or similar coding systems that do not apply to all healthcare AI use cases.
The auto-installation of PyTorch and deep learning dependencies is convenient for new setups but is a potential conflict in existing research environments that pin specific PyTorch versions for other projects. The pyproject.toml pins torch~=2.7.1, which may clash with environments requiring a different PyTorch version.
A comparable alternative for clinical NLP specifically is the medspaCy ecosystem, which focuses on text extraction from clinical notes using spaCy-based pipelines rather than the structured EHR prediction models that PyHealth centers on. For structured EHR tasks, PyHealth's pipeline approach is more directly applicable.
Maintenance, Versioning, and License
PyHealth is licensed under MIT for its own code; a separate BSD-3-Clause is listed in pyproject.toml. The last push was on 2026-09-26 and the current release is v2.0.2, released on 2026-09-02. Earlier 2.0 releases were v2.0.1 (2026-04-01) and v2.0.0 (2026-03-23).
The project maintains a CI pipeline visible in the repository's GitHub Actions configuration and a CHANGELOG.md documenting version changes. Contributors are invited via a Discord server and a mailing list linked from the README. The README note at the top warns that the README itself may be out of date and points to pyhealth.readthedocs.io for the most current documentation.
Editorial conclusion
PyHealth 2.0 is well-suited for ML researchers who need a reproducible baseline for clinical predictive modeling tasks, and for healthcare data scientists who want to run deep learning experiments on MIMIC or eICU without writing all the preprocessing and model scaffolding themselves. It is not the right choice for teams that need Python 3.9 to 3.11 support, since version 2.0 requires 3.12 or 3.13. Before starting, check that your EHR dataset fits the MIMIC-III, MIMIC-IV, eICU, or OMOP format, because the pipeline is designed around those schemas. The last push was on 2026-09-26 and the latest release is v2.0.2.
Frequently asked questions
What is PyHealth?
PyHealth is a Python deep learning toolkit for clinical predictive modeling. It provides a five-stage modular pipeline, 33 or more pre-built models, and dataset loaders for MIMIC-III, MIMIC-IV, eICU, and OMOP, designed for both ML researchers and medical practitioners building healthcare AI applications.
Which Python version does PyHealth 2.0 require?
PyHealth 2.0 requires Python 3.12 or 3.13. The README states this explicitly and explains that the requirement enables the parallel processing, memory management, and dependency compatibility improvements introduced in version 2.0. The legacy 1.16 release still supports Python 3.9 and later.
What clinical datasets does PyHealth support out of the box?
The README lists MIMIC-III, MIMIC-IV, eICU, and OMOP CDM as the natively supported datasets. The examples/ directory contains Jupyter notebooks demonstrating loading and modeling workflows for MIMIC-III and MIMIC-IV specifically.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/sunlabuiuc-pyhealth)