# datatrove: HuggingFace's pipeline library for filtering and deduplicating web text at scale

> datatrove is a Python library of prebuilt blocks for reading, filtering, deduplicating and tokenizing text data, run through an executor on a laptop, a Slurm cluster or Ray. It is built for people preparing LLM training corpora, not for general data engineering.

**huggingface/datatrove** — Freeing data processing from scripting madness by providing a set of platform-agnostic customizable pipeline processing blocks.

- Repository: https://github.com/huggingface/datatrove
- Stars: 3,358 · Forks: 308
- Language: Python
- License: Apache-2.0
- Published: 2026-09-24 · Updated: 2026-09-24 · Language: en
- Canonical page: https://hysenlabs.com/projects/huggingface-datatrove

## The scripting mess datatrove is aimed at

Preparing a web-scale training corpus usually starts as a handful of scripts: one to read WARC files, one to extract text, one to apply filters, one to deduplicate, one to tokenize. Each has its own argument parsing, its own output layout and its own idea of how to resume after a crash. datatrove's stated goal is to free data processing from that scripting madness by shipping a set of platform-agnostic, customizable pipeline blocks. The intended reader is someone with a very large text workload, which the README names directly: processing an LLM's training data.

That framing matters when you decide whether to use it. datatrove is not a general ETL tool. Its atom is a Document with text, an id and a metadata dictionary, and its examples are all corpus preparation: a reproduction of FineWeb, a Common Crawl dump pipeline that runs on Slurm, a C4 tokenization script using the gpt2 tokenizer, MinHash and sentence-level deduplication. If your job is moving rows between a database and a warehouse, the vocabulary here will not fit.

## Pipelines, executors and the task-sharding model

The architecture separates three things. A pipeline is a list of processing steps. An executor runs that pipeline on a given environment. A job is one execution of a pipeline on an executor, and a job is split into tasks that parallelize the work, usually one shard of data per task.

The sharding rule is the part worth reading twice. A shard is a group of input files assigned to a specific task, and each task processes a different, non-overlapping shard. datatrove tracks which tasks have completed, so relaunching a job runs only the incomplete ones. That is the resumption mechanism, and it is the main reason the block design pays off on long jobs.

Two constraints follow from the README. Each file is processed by a single task, and datatrove does not automatically split a file into multiple parts, so fully parallelizing means having multiple medium-sized files rather than one large one. And if tasks outnumber files, some tasks process nothing, so there is usually no point setting tasks higher than files. The README's own worked example: 10000 files on a 100-core machine with 1000 tasks gives each task a shard of 10 files, and workers=100 means 100 tasks run at once. The trade-off is stated plainly too. Few tasks means each one is long, and a failure restarts a large unit; many small tasks means failures are cheap to rerun.

## Installing datatrove and running a first pipeline

The README requires Python 3.10 or newer and installs with uv sync. Dependencies are split into extras you combine by repeating --extra. The available extras are all, io, processing, s3, cli, ray, inference, decont and multilingual, so a text extraction and filtering setup needs io and processing.

```bash
uv sync --extra processing --extra s3
```

That command resolves the project's dependencies plus the processing and s3 extras. The README lists io as the extra for reading warc/arc/wet files and arrow or parquet, processing for text extraction, filtering and tokenization, s3 for s3 support, cli for command line tools, ray for the distributed compute engine, and inference for LLM inference pipelines.

For a first real use, the repository's examples directory is the entry point rather than a tutorial in the README. The README points to examples/process_common_crawl_dump.py as a full pipeline that reads Common Crawl WARC files, extracts their text, filters, and saves the result to s3, running on Slurm. To run something end to end without a cluster, examples/fineweb.py is described as a full reproduction of the FineWeb dataset.

```bash
python examples/fineweb.py
```

The README does not document the exact arguments these scripts accept, so read the file before running it. What you should expect from a datatrove job is not a single output file but a set of shards written through the executor, with logging that reports per-task progress.

## Choosing an executor: local, Slurm, Ray or jobs

The same pipeline runs on different backends because the executor is a separate object. The README names four: LocalPipelineExecutor, SlurmPipelineExecutor, RayPipelineExecutor and JobsPipelineExecutor. The platform-agnostic claim rests entirely on this split, and it is the strongest design decision in the library, because it means a pipeline you debug on one machine is the pipeline you submit to a cluster.

The mapping is not free, though. The extras tell you what each backend costs: ray is a separate extra, and the Slurm path is exercised by examples/process_common_crawl_dump.py, which the README describes as running on Slurm. Local execution uses workers as a proxy for CPU cores; the README's example treats workers=50 as 50 simultaneous tasks on 50 cores. If you are on a machine with a scheduler you do not control, or a managed platform that is none of these four, the abstraction stops helping and you are back to writing glue.

There is also a jobs variant in the examples: filter_hf_dataset_jobs.py, minhash_deduplication_jobs.py and tokenize_hf_dataset_jobs.py. The naming suggests a mode where a pipeline is submitted as multiple jobs rather than one, which is worth reading in the source if your cluster has wall-clock limits that a single long job would exceed.

## Where datatrove is the wrong tool

The file granularity limit is the sharpest one. Because each file goes to exactly one task and datatrove does not split files, a single enormous JSONL file caps your parallelism at one task. You cannot fix this with configuration; you fix it by repartitioning the input beforehand. Anyone whose data arrives as one giant export will spend their first day splitting files, not running pipelines.

The second limit is scope. Everything in the README is text: WARC and WET archives, parquet and arrow, tokenization, MinHash and sentence deduplication, URL and exact-substring deduplication. There is no evidence of a relational source, a streaming source, or a schema contract. If your records are structured events with types you need to validate, the Document model (text, id, metadata) is a poor fit and you will be carrying your real structure inside an untyped metadata dictionary.

The third is dependency weight. A full install pulls in trafilatura, fasttext-numpy2-wheel, nltk, inscriptis, tldextract, tokenizers, ftfy, pyahocorasick and more through the processing extra, plus lighteval for decontamination and spacy for multilingual. The extras exist precisely because you should not install all of it. Installing --extra all to get one filter is the common mistake, and it also pins trafilatura to >=1.8.0,<1.12.0, so an unrelated environment that needs a newer trafilatura will conflict.

## How datatrove compares with hand-written Spark or Dask jobs

The closest real alternative for this workload is writing the same steps as Spark or Dask jobs, or as a bespoke script set. The difference is where parallelism and resumption live. In Spark, you express transformations over distributed collections and the framework owns partitioning, retries and shuffle. In datatrove, you express a pipeline of blocks and choose an executor; parallelism comes from the number of tasks you declare, and resumption comes from datatrove tracking completed tasks so a relaunch only runs the incomplete ones.

That makes datatrove lighter to start and more predictable to reason about per unit of work: a task is a shard of files, a worker is a core, and the README's arithmetic (10000 files, 1000 tasks, 100 workers) is something you can do in your head. Spark gives you a query optimizer, a wider connector ecosystem and a much larger hiring pool. datatrove gives you the specific corpus-preparation blocks (WARC reading, MinHash deduplication, tokenization) already written, and no cluster of its own to operate beyond Slurm or Ray. If your team already runs Spark and your data is not web text, switching buys you little.

## Maintenance, licensing and upgrade cost

datatrove is not archived, and the last push to main was on 2026-09-17. Releases are infrequent and deliberate: v0.8.0 on 2026-01-19, v0.9.0 on 2026-03-04, and v0.10.0 on 2026-08-13, with pyproject.toml at version 0.10.0. That cadence means upgrading is a decision you make a few times a year, not a weekly chore, but it also means fixes you are waiting for may sit for months.

Upgrade cost is dominated by the extras, not the core. The core dependency list is short: dill, fsspec, huggingface-hub, humanize, loguru, multiprocess, numpy and tqdm. The constraints to watch are numpy>=2.0.0, huggingface-hub>=1.6.0 (pyproject.toml notes this is needed for HfFileSystemResolvedBucketPath, the bucket-write path) and trafilatura>=1.8.0,<1.12.0. A minor bump that moves any of those can collide with the rest of your environment.

The licence is Apache-2.0, declared both in the repository LICENSE file and in the pyproject.toml license field, with the OSI classifier set accordingly. That is a permissive licence with an explicit patent grant, which is generally the easy case for commercial use, but the terms also require you to preserve notices and state changes. Read the LICENSE text yourself; nothing here is legal advice. Note that the optional dependencies carry their own licences, and the processing extra pulls in several, so a redistribution of a built artifact is a different question from using the library internally.

## Conclusion

Adopt datatrove if you are building a large text corpus for language model training and you want the FineWeb-style steps (WARC reading, extraction, filtering, MinHash deduplication, tokenization) as reusable blocks instead of a folder of one-off scripts. Do not adopt it if your data is small, tabular, or already clean; the pipeline, executor and shard vocabulary is overhead you will not recover. Before committing, verify three things: that your input files can be split into many medium-sized files, because the README warns each file is processed by a single task and datatrove does not split a file automatically; that the extra you need (io, processing, s3, ray, inference) matches the formats and engines you actually use; and that your environment is Python 3.10 or newer, since pyproject.toml sets requires-python to >=3.10.0.

## FAQ

### What is datatrove and what does it do?

datatrove is a HuggingFace library for processing, filtering and deduplicating text data at very large scale. It provides prebuilt processing blocks plus a framework for adding custom ones, and runs the same pipeline locally or on a Slurm cluster.

### How do I install datatrove?

The README requires Python 3.10 or newer and installs with uv sync, combining extras as needed, for example uv sync --extra processing --extra s3. Available extras include all, io, processing, s3, cli, ray, inference, decont and multilingual.

### What is a datatrove pipeline?

A pipeline is a list of processing steps to execute, such as reading data, filtering and writing to disk. It is run by an executor on a given environment, and each execution is a job made up of tasks that process non-overlapping shards of the input files.

### Does datatrove run on Ray?

Yes. The README lists RayPipelineExecutor alongside LocalPipelineExecutor, SlurmPipelineExecutor and JobsPipelineExecutor, and there is a separate ray extra for the distributed compute engine.

### Which Python version does datatrove need?

Python 3.10 or newer. pyproject.toml sets requires-python to >=3.10.0 and lists classifiers for 3.10 through 3.13.

## Sources

- [huggingface/datatrove on GitHub](https://github.com/huggingface/datatrove)
- [Issues](https://github.com/huggingface/datatrove/issues)
- [License: Apache-2.0](https://github.com/huggingface/datatrove/blob/main/LICENSE)
- [README](https://github.com/huggingface/datatrove/blob/main/README.md)
- [Releases](https://github.com/huggingface/datatrove/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/huggingface-datatrove
