# Datumaro: a Python library and CLI for converting and auditing computer vision datasets

> Datumaro reads, filters and rewrites datasets across COCO, VOC, YOLO, CVAT and other computer vision formats from one Python API or one datum command. The trade-off is a heavyweight dependency set and a common representation that only Datumaro understands.

**open-edge-platform/datumaro** — Dataset Management Framework, a Python library and a CLI tool to build, analyze and manage Computer Vision datasets.

- Repository: https://github.com/open-edge-platform/datumaro
- Website: https://open-edge-platform.github.io/datumaro
- Stars: 692 · Forks: 162
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/open-edge-platform-datumaro

## What Datumaro solves, and for whom

Computer vision teams rarely keep one annotation format. A detection model trains on COCO JSON, an annotation vendor exports CVAT XML, a pretrained checkpoint expects YOLO text files, and a publication or audit needs statistics over all of it. Datumaro's README describes the project as a framework and CLI tool to build, transform, and analyze datasets, and its diagram puts a single Datumaro step between VOC, COCO and CVAT inputs on one side and annotation tools, model training and publication on the other. That diagram is the whole pitch: one in-memory representation, many readers and writers.

The audience is narrower than the format list suggests. If you are a Python user who already writes pandas or NumPy preprocessing scripts, Datumaro fits as a library. If you would rather not write code, the datum console script covers conversion, filtering, splitting, statistics and validation. People who only ever read a single COCO file and hand it to a training loop get little from it.

## The dataset model behind the format list

The README does not spell out the internal data model, but the feature list implies one: datasets are read into a common structure, then transformed, then written out. Everything Datumaro advertises happens on that intermediate structure rather than on the source files. Merging multiple datasets, filtering by custom criteria, converting polygons to instance masks, renaming or removing labels, and splitting into train, val and test are all operations on the in-memory dataset, not on COCO JSON or VOC XML directly.

The supported formats are listed with their annotation types, and that pairing matters. COCO is listed for image_info, instances, person_keypoints, captions, labels, panoptic and stuff. PASCAL VOC is listed for classification, detection, segmentation, action_classification and person_layout. Kitti is listed for segmentation, detection and 3D raw with velodyne points. YOLO appears only with bboxes. So a conversion between two formats is only as complete as the intersection of their annotation types. Moving VOC segmentation into YOLO is not a supported path, because YOLO's entry names bounding boxes only.

The build configuration shows a second layer under the Python API. pyproject.toml declares setuptools-rust and pybind11 in the build requirements, and the repository has a top-level rust/ directory. Something in Datumaro is compiled, not pure Python, which is worth knowing before you plan a cross-platform wheel build.

## Installing Datumaro and running a first conversion

The package is named datumaro and requires Python 3.10 through 3.14 according to pyproject.toml. The README points to a quick start guide on the documentation site rather than giving an install command, so the exact pip invocation follows the package name and the declared console script, which is datum.

```bash
pip install datumaro
```

After installation the datum command should be on your PATH. The README's own diagram is the smallest useful exercise: take one dataset in one format and write it out in another. The README does not reproduce the CLI reference, so the exact flags for a conversion are documented in the user manual rather than in the repository's front page. The format identifiers to pass are the ones listed in the README's feature list, such as coco and voc.

Expect a new directory tree at the output path in the target format's layout, and expect the command to fail loudly if the source directory does not match the format you declared. Confirm the flag spelling for your installed version against the user manual before scripting it.

## Filtering and re-splitting without touching the source files

Conversion is the visible feature; the filtering and splitting operations are the ones that save real time. The README lists concrete criteria rather than abstractions: remove polygons of a certain class, remove images without annotations of a specific class, remove occluded annotations, keep only vertically-oriented images, remove small area bounding boxes. Each of those is a rule you would otherwise implement by walking XML or JSON and rewriting it, with the usual risk of breaking a sibling field.

Splitting is more interesting than a random shuffle. Datumaro documents task-specific splits based on annotations that keep the initial label and attribute distributions: for classification, based on labels; for detection, based on bounding boxes; for re-identification, based on labels while avoiding the same IDs in training and test splits. The re-identification case is the one that catches people out, because a naive random split leaks identities across the boundary and inflates evaluation scores. Datumaro treats that as a first-class split mode.

The same layer supports quality checking: simple error checks, comparison with model inference, merging and comparison of multiple datasets, and annotation validation based on task type. Comparison against model inference is the least conventional item here and the one the README gives the least detail about, so treat it as something to inspect in the user manual before you build a pipeline around it.

## Where Datumaro is the wrong tool

The dependency list is the first honest constraint. A Datumaro install pulls in numpy, opencv-python-headless, pandas, pyarrow, pycocotools, shapely, lxml, orjson, ruamel.yaml and more. That is a reasonable footprint for a dataset tool and an unreasonable one for a training container that only needs to read one COCO file. If your only job is parsing COCO annotations, pycocotools is the smaller dependency, and Datumaro's value comes from the second format and the third operation, not the first.

Format coverage is uneven in a way the README is transparent about but easy to skim past. The parenthetical annotation types are the contract. A format listed without the type you need is not a partial success; it is a conversion that will not carry your annotations. Before planning a migration, check the type in the README's list, not just the format name.

There is also the question of what happens to anything Datumaro does not model. Dataset formats carry sidecar metadata, unusual attribute encodings and vendor extensions. A round trip through a common representation can normalize or drop what the common representation has no field for. The README does not document round-trip fidelity guarantees, so if your source format has fields outside the listed annotation types, test the round trip on a sample before trusting it on the full dataset.

## Alternatives and the actual difference in approach

The closest comparison is FiftyOne, which also ingests many computer vision formats and exposes a Python API and CLI. The difference is where the data lives. FiftyOne's model is a database of samples you query and browse, with a UI for inspection, so the tool is oriented around exploring and curating a persistent store. Datumaro's model, as the README describes it, is a transformation pipeline: read, filter, convert, split, write. You get files back, not a queryable database, and there is no browsing interface in the feature list. If your workflow is producing a clean dataset directory for training, Datumaro's shape matches. If it is interactively inspecting and tagging samples across months, it does not.

The other alternative is writing the conversion yourself against pycocotools or lxml. That is genuinely fine for one format pair and one operation, and it avoids the dependency weight entirely. It stops being fine at the second format, because that is when you start maintaining per-format readers and writers, which is exactly the code Datumaro already ships.

## Maintenance, licensing and upgrade cost

Datumaro is MIT licensed, with license-files pointing at LICENSE and a separate NOTICE file at the repository root. MIT is permissive, so embedding it in a commercial pipeline is not the obstacle. The practical licence question is the dependency set: opencv-python-headless, pycocotools, pandas and the rest carry their own licences, and 3rd-party.txt at the repository root is where the project tracks them. Read that file rather than assuming the MIT label covers everything you install.

The project is not archived, and the last push was on 2026-09-07. Releases v1.13.9, v1.13.10 and v1.13.11 landed on 2026-08-27, 2026-09-02 and 2026-09-07 respectively, so the cadence is frequent. Frequent patch releases cut both ways: fixes arrive quickly, and a pinned version can fall behind within weeks. The declared Python range is >=3.10,<3.15, and the dependency pins are tight (numpy >=2.2,<2.5, pillow ~=12.0, lxml ~=6.0), which means a Datumaro upgrade can force upgrades elsewhere in your environment. Pin Datumaro in your lockfile and upgrade it deliberately rather than letting it float.

## Conclusion

Adopt Datumaro when you have more than one annotation format in play and need a scriptable step between them, or when you need to filter and re-split a dataset without hand-editing XML. Do not adopt it if you only ever read COCO and never write anything else: pycocotools is smaller and does not pull in pandas, pyarrow and OpenCV. Before committing, verify that your target format's importer covers the annotation types you actually use, because the format list is broader than the per-type support, and check the release cadence against your own upgrade window.

## FAQ

### What is the Datumaro dataset format used for?

Datumaro does not define a single on-disk format you must adopt. It reads datasets in formats such as COCO, VOC, YOLO, CVAT and Kitti into a common in-memory representation, then writes them back out in whichever of those formats you choose, so the operations act on that intermediate representation.

### What is Datumaro used for?

The README describes it as a framework and CLI tool to build, transform, and analyze datasets. Concretely that covers format conversion, merging datasets, filtering by criteria such as removing occluded annotations, annotation conversion between polygons and masks, splitting, quality checking and dataset statistics.

### How do you describe a dataset?

Datumaro's README lists the annotation types each supported format carries, such as COCO with instances, captions and panoptic, or PASCAL VOC with detection, segmentation and person_layout. That per-format type list is how the project characterizes a dataset it can read.

## Sources

- [License: MIT](https://github.com/open-edge-platform/datumaro/blob/develop/LICENSE)
- [open-edge-platform/datumaro on GitHub](https://github.com/open-edge-platform/datumaro)
- [Project website](https://open-edge-platform.github.io/datumaro)
- [README](https://github.com/open-edge-platform/datumaro/blob/develop/README.md)
- [Releases](https://github.com/open-edge-platform/datumaro/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/open-edge-platform-datumaro
