# pydicom: reading DICOM metadata in Python without installing a C library

> pydicom is a pure Python library for reading, editing and writing DICOM datasets, with optional NumPy and plugin support for pixel data. It deliberately stops at the dataset layer and leaves networking and anonymisation to sibling projects.

**pydicom/pydicom** — Read, modify and write DICOM files with python code

- Repository: https://github.com/pydicom/pydicom
- Website: https://pydicom.github.io/pydicom/dev
- Stars: 2,220 · Forks: 559
- Language: Python
- License: NOASSERTION
- Published: 2026-10-07 · Updated: 2026-10-07 · Language: en
- Canonical page: https://hysenlabs.com/projects/pydicom-pydicom

## Installing the package and the extras you actually need

The base install is deliberately tiny, with two package managers shown:

```bash
pip install pydicom
```

```bash
conda install -c conda-forge pydicom
```

What you get by default is more limited than most first-time users expect. The package manifest lists `dependencies = []`, which is unusually honest for a medical imaging library: nothing is required to read a dataset, because a DICOM dataset is a structured set of data elements and reading the tags does not need an image decoder.

The extras exist for when you cross that line. `basic` adds `numpy` and `types-pydicom`, so you get ndarray access to pixel data and editor types. `pixeldata` is the one to read carefully: `numpy`, `pillow`, `pyjpegls`, `pylibjpeg[openjpeg]`, `pylibjpeg[rle]`, `pylibjpeg-libjpeg` and `python-gdcm`. That is a list of native dependencies, and the point of the split is that a script that only reads StudyInstanceUID and Modality should not have to install gdcm.

There is also a `gpl-license` extra, which resolves to `pylibjpeg[libjpeg]`. The comment left in the manifest next to it explains why: they would rather use libjpeg-turbo, which has a BSD-like licence, but 12-bit support is not available in it yet. So the default path stays permissively licensed and the GPL path is opt-in. The manifest also pins `requires-python = ">=3.10"` and carries classifiers through Python 3.14.

## Reading a file and looking at raw PixelData bytes

The entry point is `dcmread`, and the first example in the README is deliberately about raw bytes rather than an image array:

```python
from pydicom import dcmread
from pydicom.data import get_testdata_file
path = get_testdata_file("CT_small.dcm")
ds = dcmread(path)
type(ds.PixelData)
len(ds.PixelData)
ds.PixelData[:2]
```

`get_testdata_file` fetches a bundled test file rather than one on disk, so the snippet runs without a path. The documented output shows a `bytes` object of length 32768 whose first two bytes are `b'\xaf\x00'`, which is the point: compressed and uncompressed Pixel Data is always available as bytes, readable, changeable and writable, with no optional dependency involved.

The dataset is also the editing surface. Elements are addressed by keyword, which the second example demonstrates on the Patient ID:

```python
from pydicom import dcmread
ds = dcmread("/path/to/file.dcm")
ds.PatientID = "12345678"
ds.save_as("/path/to/file_updated.dcm")
```

Three lines, and the file is rewritten with the new value. There is no separate serialization step, no round trip through a DICOM dictionary and no separate toolkit. For de-identification work this is the operation you perform thousands of times, and it is why the library sits at the centre of the pydicom organisation's other projects.

## Getting a numpy array, and which codecs need plugins

With NumPy present, the `pixel_array` property converts Pixel Data into an array, and the README shows the shape and dtype the conversion produces:

```python
arr = ds.pixel_array
arr.shape
```

which gives `(128, 128)` for the bundled CT file with a dtype of `int16`. Displaying it is then matplotlib's job:

```python
import matplotlib.pyplot as plt
from pydicom import dcmread, examples
path = examples.get_path("ct")
ds = dcmread(path)
arr = ds.pixel_array
plt.imshow(arr, cmap="gray")
plt.show()
```

The codec story is where this gets genuinely conditional, and the README is upfront about it. For JPEG, JPEG-LS and JPEG 2000 decompression you must install additional libraries, and the docs page on pixel data handlers says which. For RLE, only NumPy is needed, but the README warns it can be quite slow and suggests installing one of the additional libraries to speed it up.

The same split applies in reverse when compressing. JPEG-LS compression requires `pyjpegls`. JPEG 2000 requires `pylibjpeg` plus the `pylibjpeg-openjpeg` plugin. RLE compression requires nothing extra, with the same slowness caveat, and `pylibjpeg` with the `pylibjpeg-rle` plugin or `gdcm` as the speedups.

So the honest summary is that pydicom is pure Python for everything except image codecs, and the codecs are delegated rather than reimplemented. The trade is real: you gain a package that installs anywhere, and you lose the property that every format works out of the box.

## Where pydicom deliberately stops

This is the most useful paragraph in the README and it is easy to skim past. The package is described as a general-purpose DICOM framework concerned with reading and writing DICOM datasets, and it states that to keep the project manageable, it does not handle the specifics of individual SOP classes or other aspects of DICOM.

Two named sibling projects show what that means in practice. `pynetdicom` is a Python library for DICOM networking, and `deid` supports anonymisation of DICOM files. Both are in the pydicom GitHub organisation and both build on pydicom rather than replacing it. If your task involves a modality-specific workflow, a conformance statement, or a network association, you are outside pydicom's scope by design, and the README tells you so instead of leaving you to discover it.

The repository layout backs this up. Beyond `src/` and `tests/`, the tree contains `benchmarks/`, `build_tools/`, `doc/`, `examples/` and `util/`, plus a `Makefile` and a `_typos.toml`. The examples folder is subdivided rather than flat, into `input_output/`, `image_processing/`, `metadata_processing/` and `memory_dataset.py`, which maps onto the same division of labour: datasets in and out, pixel data, metadata, and the in-memory representation.

There is also an `asv.conf.json` at the top of the tree, which is an asv configuration for performance benchmarking, and a `benchmarks/` directory to match. That tells you someone is tracking the speed of this library deliberately, which matters more than usual when a single `pixel_array` call can decompress a large multi-frame study.

## Version lines and what the release history shows

The release list on the repository shows two live lines. pydicom 3.0.2 was published on 2026-03-19 and 3.0.1 on 2024-09-22, and then 2.4.5 was published a day after 3.0.2, on 2026-03-20. That ordering is informative: the 2.x line was still getting a patch release after 3.0.2 shipped, which is what a project does for users who have not migrated. A reader who installs today and finds only 3.x behaviour in the documentation should not assume 2.4.5 is dead.

The repository manifest gives a clearer picture of where development currently sits. The version field reads `3.1.0.dev0`, which is a development version rather than a release, and the description reads "A pure Python package for reading and writing DICOM data". The build backend is flit_core with a constraint of `>=3.12,<5`, and the licence field says MIT with `LICENSE` as the licence file.

That last point is worth a note because the GitHub repository metadata does not classify the licence at all, reporting no recognised licence type. The `pyproject.toml` in the repository says MIT and names the LICENSE file. The README also links to a Zenodo DOI badge, which is how a research library gets a citable archive snapshot, and the project describes itself in the README as volunteer-run with limited resources, which is a reasonable expectation to set about response times on issues.

## The comparison worth making, and when to reach elsewhere

The usual alternative is not another Python DICOM library but a different kind of tool: a command line utility like `dcmdump` from the DCMTK, or a viewer application, or a commercial PACS archive. The difference in approach is the interesting part. Those tools give you a formatted dump or an image on screen; pydicom gives you a Python object you can branch on, loop over, annotate and save back. For a pipeline that has to make decisions, that is the whole point.

The cost is the one already described. If your data is uniformly uncompressed, or if you are willing to install `pydicom[pixeldata]`, pydicom is not meaningfully harder. If your data is a mixed archive of JPEG 2000 studies from several vendors and you want zero extra native packages, you will spend your time on codec installation rather than on your task. Note also that pydicom is a dataset library, not a viewer, and it will not give you a window to look at a series.

One more practical detail: if you need to compare two datasets field by field, the examples tree contains `plot_dicom_difference.py`, and there is a separate `show_charset_name.py` for the character set problems that plague older DICOM files. Those small scripts are often more useful as a starting point than the official tutorials, because they solve the two problems newcomers actually hit.

## Conclusion

pydicom is the right choice when the DICOM metadata is what you need and you do not want a compiled toolchain or a heavyweight imaging stack in your environment. The pure Python guarantee means it installs anywhere Python runs, the dataset model gives you keyword access to every element, and the write path lets you change a tag and save without understanding the file layout. It is the wrong choice if you need network DICOM, because that is `pynetdicom`'s job, if you need de-identification, because that is `deid`'s, and if you need to decode JPEG 2000 pixel data without also installing `pylibjpeg` and a codec plugin. Check the extras before you start: `pip install pydicom` gives you bytes, `pydicom[basic]` adds NumPy for arrays, and `pydicom[pixeldata]` pulls in Pillow, pyjpegls, the pylibjpeg family and gdcm. The last push was 2026-09-19 and the package version in the manifest reads 3.1.0.dev0, so the 3.x line is where any fixes land.

## FAQ

### How do I install pydicom?

The README gives two options: `pip install pydicom` or `conda install -c conda-forge pydicom`. The base package has no required dependencies, and the manifest pins Python 3.10 or newer. For pixel data you want the extras, with `basic` adding NumPy and `pixeldata` adding Pillow, pyjpegls, the pylibjpeg family and gdcm.

### How do I open a DICOM file?

Call `dcmread` with the path. `dcmread("/path/to/file.dcm")` returns a Dataset whose elements are addressed by keyword, so you can read `ds.PatientID` directly. The README also uses `get_testdata_file` to open a test file bundled with the package, which avoids needing a path on disk.

### How can I convert a DICOM file to a JPEG?

The README documents the pieces separately rather than as a one-step converter. Pixel Data is always available as bytes with no dependencies, and with NumPy the `pixel_array` property returns an ndarray, which the example then passes to `plt.imshow` with a grayscale colormap. Decoding JPEG, JPEG-LS or JPEG 2000 pixel data requires additional libraries, and the guides on pixel data handlers list which ones.

### Is DICOM still relevant?

The README describes pydicom as a general-purpose DICOM framework and points at the DICOM standard, and the ecosystem around it is still expanding rather than shrinking. The library itself has two maintained lines, with 3.0.2 published on 2026-03-19 and a 2.4.5 patch on 2026-03-20, and the last push was on 2026-09-19.

## Sources

- [Issues](https://github.com/pydicom/pydicom/issues)
- [Project website](https://pydicom.github.io/pydicom/dev)
- [pydicom/pydicom on GitHub](https://github.com/pydicom/pydicom)
- [README](https://github.com/pydicom/pydicom/blob/main/README.md)
- [Releases](https://github.com/pydicom/pydicom/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/pydicom-pydicom
