Open-source project
MIT-LCP/mimic-code avatar
MIT-LCP/mimic-code

MIT-LCP/mimic-code: SQL Build Scripts for the MIMIC Critical Care Databases

MIMIC Code Repository: Code shared by the research community for the MIMIC family of databases

3,388 stars1,739 forksJupyter NotebookMIT

At a glance

What is it?
The mimic-code repository is the shared SQL and Python tooling layer for the MIMIC family of PhysioNet datasets. It is useful if you already have credentialed access to MIMIC-IV or MIMIC-III and need reproducible derived tables rather than a quick CSV.
Who is it for?
Adopt mimic-code if you hold PhysioNet credentials for a MIMIC dataset and want the community's derived concept definitions rather than writing your own cohort SQL. Do not adopt it if you have no data access, since the repository ships no data, or if you expect a packaged Python library that reads MIMIC tables for you.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 30 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What mimic-code solves, and who it is actually for

MIMIC is not a tidy table. It is a set of relational databases covering ICU and hospital stays, with item identifiers, chartevents, labevents and note tables that require a long chain of joins and filters before anything resembles a cohort. Every research group that used it wrote that chain again, slightly differently, which made published results hard to compare. The MIMIC Code Repository exists to hold one shared version of those chains. The README describes it as "a central hub for sharing, refining, and reusing code used for analysis of the MIMIC critical care database".

The audience is narrow and specific. You need credentialed access to one of the PhysioNet datasets before any of this runs: MIMIC-III, MIMIC-IV, MIMIC-IV-Note, MIMIC-IV-ED or MIMIC-CXR. The repository contains build scripts, derived concept definitions and tutorials, not data. A data scientist with a MIMIC-IV extract in hand is the intended user. Someone looking for a downloadable sample dataset will find nothing here, and someone who wants a scikit-learn style loader for MIMIC tables is looking at the wrong project.

How the repository is organised across MIMIC versions

The layout is one top-level folder per dataset: mimic-iii, mimic-iv, mimic-iv-note, mimic-iv-cxr, mimic-iv-ed and a mimic-iv-waveforms placeholder that the README marks as TBD. Each holds build scripts, derived concepts and tutorials, and each has its own README with the detail. That per-folder split matters because the schemas differ between MIMIC-III and MIMIC-IV, so SQL written against one will not run against the other without editing.

The derived concepts are also published, not just generated. The README states that MIMIC-III derived concepts are available on the physionet-data.mimiciii_derived dataset on BigQuery, and MIMIC-IV concepts on release-specific datasets such as physionet-data.mimiciv_3_1_derived. That is the single most consequential fact in the repository for a new user: if the concept you need is already in the published derived dataset and your access is on BigQuery, you may not need to run the build at all. The scripts matter when you want to change a definition, run on your own PostgreSQL instance, or work with a MIMIC version that has no published derived dataset.

Installing mimic_utils and running the first build

The Python side of the repository is packaged as mimic_utils, described in the README as utilities for working with the MIMIC datasets, primarily transpiling SQL code. The pyproject.toml declares requires-python >=3.10 and depends on sqlglot, pandas, numpy and tqdm. The README's installation step is an editable install with the test extra:

bash
pip install -e ".[test]"

For reproducible builds the repository ships requirements-lock.txt, generated by pip-compile for Python 3.9 according to the README. Note the mismatch: the lock file targets 3.9 while pyproject.toml requires 3.10 or newer. The README gives two commands for the locked path, installing dependencies first and then the package without resolving them again:

bash
pip install -r requirements-lock.txt
pip install -e . --no-deps

Tests run through pytest, with testpaths set to tests in pyproject.toml:

bash
pytest tests/

What you should see is pytest collecting from the tests directory. The README does not state an expected pass count, so treat any failure as something to inspect rather than a known quantity. The installed console entry point is mimic_utils, mapped to mimic_utils.__main__:main, which is the CLI the README points at for the SQL transpiling work. If you only need derived tables on BigQuery, skip this section entirely and query the published derived dataset.

The transpiling design and where it breaks down

The core design decision is that the concept definitions are written in SQL and then transpiled between dialects rather than maintained as separate per-dialect copies. sqlglot is the dependency that does this, and the test extra pulls in duckdb and sqlfluff, which indicates the intended verification path: parse and lint the SQL, then compare results across engines. The optional equivalence extra adds psycopg2-binary alongside duckdb, which points at cross-engine comparison between DuckDB and PostgreSQL as a supported workflow rather than an afterthought.

The limitation is inherent to transpiling. Dialect translation handles syntax, not semantics: functions that exist in one engine and not another, type coercion rules, and date arithmetic differences are exactly where generated SQL drifts from the original. The repository's own test suite is the only signal you get, and the README does not document a rollback path if a build produces different numbers than the published derived dataset. There is also no documented compatibility matrix between mimic_utils versions and MIMIC-IV releases, so a v3.0.1 checkout against an older MIMIC-IV extract is a combination you have to validate yourself. The repository is not archived and the last push was on 2026-09-01, so the code is being touched, but the README does not commit to a release cadence.

Licence, citation and the cost of upgrading

The repository is MIT licensed and ships the LICENSE file at the top level, which is permissive for the code. The data is a separate matter entirely: MIMIC datasets are governed by PhysioNet credentialing and a data use agreement, and the MIT licence on this repository grants you nothing with respect to the data. The README also asks that you cite the datasets you use, cite the Zenodo archive at record 6818823 by selecting the release you used, and cite the MIMIC code repository paper in JAMIA. If you publish, that is three citations, and the Zenodo one is version-specific, which means recording which release you built against.

Upgrade cost is driven by the release-specific derived datasets. The recent releases are v3.0.1 on 2025-04-24, v3.0.0 on 2025-04-23 and v2.5.0 on 2024-08-26. Because MIMIC-IV derived concepts are published per release, moving from one MIMIC-IV version to the next can change table names in BigQuery, and any query pinned to physionet-data.mimiciv_3_1_derived will need editing. The README does not document a migration guide between derived dataset versions, so budget time for diffing your own queries. Treat the licence question as one for your institution's research office, not as something the repository answers.

Alternatives and when mimic-code is the wrong tool

The README itself lists adjacent projects, and the differences are real. MIMIC Extract is described as a Python package for transforming MIMIC-III data into a machine learning friendly format, so it targets feature matrices rather than SQL concept definitions; if your endpoint is a model, that is a shorter path than building concepts and reshaping them yourself. FIDDLE is described as a flexible data-driven pipeline transforming structured EHR data into a machine learning friendly format, which is a pipeline abstraction rather than a set of SQL scripts. Bloatectomy solves a different problem entirely, removing duplicate text in clinical notes, and Medication categories extracts medications from free-text notes.

The honest case against mimic-code is when you want a stable Python API. mimic_utils is described as utilities for transpiling SQL, not as an ORM or a dataframe loader; you still supply the database connection and run the SQL. If your team has no SQL fluency, or if you need MIMIC-IV Waveforms, note that the README marks that dataset as TBD and the folder as a placeholder, so there is nothing to build against yet. And if you only need a handful of variables, writing your own queries against the raw tables may be less work than learning the concept layer and its versioning.

Editorial conclusion

Adopt mimic-code if you hold PhysioNet credentials for a MIMIC dataset and want the community's derived concept definitions rather than writing your own cohort SQL. Do not adopt it if you have no data access, since the repository ships no data, or if you expect a packaged Python library that reads MIMIC tables for you. Before committing, check which top-level folder matches your dataset version, confirm whether the derived tables you need already exist in the release-specific BigQuery dataset, and read the subfolder README because the top-level README does not document rollback or version pinning for the SQL builds.

Frequently asked questions

Does mimic-code include the MIMIC patient data?

No. The repository holds build scripts, derived concept definitions and tutorials. The README states that access to the datasets is requested through PhysioNet, and that cloud access requires adding the relevant cloud identifier to your PhysioNet profile and requesting access via the PhysioNet project page.

Which Python version does mimic_utils need?

pyproject.toml sets requires-python to >=3.10. The README separately notes that requirements-lock.txt was generated by pip-compile for reproducible builds on Python 3.9, so the locked dependency path and the package metadata do not agree on the interpreter.

Do I have to run the SQL builds myself?

Not necessarily. The README states that derived concepts are available on BigQuery, with MIMIC-III concepts in physionet-data.mimiciii_derived and MIMIC-IV concepts in release-specific datasets such as physionet-data.mimiciv_3_1_derived. Running the scripts yourself is for when you need to change a definition or work on another engine.

How do I run the tests for mimic-code?

The README gives pytest tests/ as the command, and pyproject.toml sets testpaths to tests. The test extra installs pytest, duckdb and sqlfluff, so the environment for running them comes from the editable install with the test extra.

How should mimic-code be cited?

The README asks for three things: cite the datasets you use as described on their PhysioNet project pages, cite the Zenodo repository at record 6818823 selecting the specific release you used, and cite the MIMIC code repository paper published in JAMIA.

Official sources

  1. License: MIT
  2. MIT-LCP/mimic-code on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mit-lcp-mimic-code.svg)](https://hysenlabs.com/projects/mit-lcp-mimic-code)