allenai/dolma: the toolkit behind AI2's three-trillion-token corpus
Data and tools for generating and inspecting OLMo pre-training data.
At a glance
- What is it?
- Dolma is two things at once: an open 3-trillion-token pre-training corpus and the Python and Rust toolkit that curates data for language models. Here is what it does, how it installs, and where it stops being the right tool.
- Who is it for?
- Adopt Dolma if you are curating a large text corpus for pre-training and you want the taggers, deduplication and parallelism that AI2 already used for OLMo. Skip it if you only need to clean a few gigabytes of text, or if the documents you are filtering are not plain text with JSONL metadata.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 36 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Dolma actually is, and the two projects sharing one name
The repository README opens by separating two things that share the name. Dolma Dataset is an open corpus of 3 trillion tokens drawn from web content, academic publications, code, books and encyclopedic materials, assembled as the training corpus for OLMo at the Allen Institute for AI. Dolma Toolkit is the code in this repository: a system for curating datasets for language modeling. If you arrived looking for the corpus, the README points you at huggingface.co/datasets/allenai/dolma, where it is distributed under ODC-BY. If you arrived to build your own corpus, you are in the right place.
The audience is narrow and specific. This is not a text-cleaning library for a weekend script. It is aimed at teams preparing pre-training data at the scale where a single pass over the documents is measured in hours and where the filtering decisions are documented in a paper. The README calls out four capabilities: built-in parallelism, portability across a single machine, a cluster or a cloud environment, ready-made taggers for Gopher, C4 and OpenWebText style filtering, and deduplication implemented as a Rust Bloom filter. That list is the product.
Taggers, a Rust core, and where the parallelism comes from
The architecture is a Python layer over a compiled Rust extension. The repository layout shows both halves: python/ and tests/ sit alongside src/ and Cargo.toml, and the Makefile builds the extension with maturin develop --extras="all" before running pytest. The Rust crate is declared in Cargo.toml as a cdylib named dolma, with pyo3 as the binding layer, rayon for parallelism, and a dependency set that includes aws-sdk-s3, zstd, tokenizers and an adblock crate for content blocking.
That split explains the performance claim. The README says the toolkit can process billions of documents concurrently thanks to built-in parallelism. In practice the heavy per-document work (deduplication via the Bloom filter, tokenization, some filtering) runs in Rust across threads via rayon, while the Python side wires up taggers, configuration and I/O. The Rust manifest also carries jaq-core, jaq-std and jaq-parse, a Rust implementation of jq, which is the mechanism behind JSON-path style attribute extraction and filtering on documents. The Python dependency list mirrors this with jsonpath-ng and jq, so both sides can express the same kind of selection.
Data flows as documents in, tagged documents out. Each document gets attributes attached by taggers; those attributes can then be used to filter. The taggers named in the README (Gopher, C4, OpenWebText) are the ones the project ships and tests. Writing a custom tagger is listed as a supported extension point, but the README does not specify the interface, so check the docs directory before assuming the plugin surface is stable.
Installing Dolma and running a first tagging pass
The README gives exactly one installation instruction: pip install dolma. The package requires Python >=3.10,<3.13 according to pyproject.toml, and it pulls in a large dependency set including fasttext-wheel, tokenizers, s3fs and numpy<2. Expect a build step, because the Rust extension is compiled.
pip install dolmaThat is the whole of the documented install path. The README does not print a full example command for tagging or filtering a corpus; it directs readers to the documentation under docs/ for how to use the toolkit. The Makefile shows how the maintainers build the extension for development, which is the closest thing to a second install recipe in the repository files.
maturin develop --extras="all"What you should see after the pip install is a dolma package with a compiled extension. What you will not find in the README is a worked command line, a flag list, or an output example. That gap matters: the tagger and filter subcommands, their argument names and their defaults are documented in docs/ rather than in the README, so read that directory before you design a pipeline around the toolkit. Treating the README as the interface reference will leave you guessing at flags.
Deduplication is a separate step. The README describes it as fast document deduplication using a Rust Bloom filter, which means it is approximate by construction. A Bloom filter can report a document as a duplicate when it is not. The README does not document a false-positive rate or a way to tune the filter size, so if exact deduplication matters to you, this is a design choice you are accepting rather than configuring.
Where Dolma is the wrong tool
The first limitation is scale. Everything about the toolkit assumes a corpus large enough that parallelism and a Bloom filter earn their keep. For a few gigabytes of text, the Rust build, the dependency tree and the tagger configuration cost more than they return. A single-process Python script with a handful of regexes will finish before you have finished reading the tagger documentation.
The second is language coverage. The built-in taggers named in the README are Gopher, C4 and OpenWebText. Those are heuristics developed for English web text. The README says nothing about non-English filtering, and nothing in the listed dependencies is a language identifier (the commented-out pycld2 and pycld3 lines in pyproject.toml are a hint that language detection was considered and left out). If your corpus is multilingual, you are writing your own tagger.
The third is the boundary between the dataset and the toolkit. Dolma Dataset is licensed ODC-BY and lives on HuggingFace. Dolma Toolkit is licensed Apache-2.0 and lives here. Downloading the corpus does not give you the toolkit's filtering pipeline, and installing the toolkit does not give you the corpus. Teams that conflate the two end up surprised.
Finally, deduplication being Bloom-filter based is a real failure mode, not a footnote. If your downstream evaluation depends on exact duplicate removal, an approximate filter will leave traces, and the README offers no guidance on sizing the filter to reduce that.
Dolma compared with a general-purpose data pipeline
The obvious alternative for a team already running a data stack is to build the same pipeline out of general-purpose tools: Spark or Dask for parallelism, a text-processing library for heuristics, and a hashing scheme for deduplication. The difference in approach is where the work lives. A Spark pipeline expresses filtering as distributed transformations over a dataframe and leaves deduplication to a join or a group-by on a hash. Dolma instead compiles the hot path into Rust and exposes a fixed set of taggers, with parallelism handled inside the extension rather than by an external scheduler.
That trade is legible. The Spark route gives you arbitrary expressiveness and an ecosystem you probably already operate. The Dolma route gives you taggers that already encode published filtering recipes, so you are not reimplementing Gopher heuristics from a paper, and you get a single binary extension that runs on one machine or on a cluster without a cluster manager. The cost is that you adopt the project's notion of what a document is, and you work within its tagger interface when you need something it does not ship.
A second alternative is to use the Dolma Dataset directly and skip the toolkit entirely. If your goal is to pre-train on a well-documented 3-trillion-token corpus rather than to curate your own, pulling the dataset from HuggingFace is the shorter path, and the data sheet in the docs directory is the document to read first.
Maintenance, versions and licence
The repository is not archived, and its last push was on 2026-08-24. The most recent tagged release is v1.2.1 from 2025-07-07, following v1.2.0 in June 2025 and v1.1.2 in February 2025. The version in pyproject.toml matches the v1.2.1 tag. Note that Cargo.toml still declares version 1.1.1 for the Rust crate, so the Python and Rust version strings are not kept in lockstep; if you are pinning by version, pin the Python package.
Upgrade cost is driven by the dependency pins rather than by API churn between the three recent releases. pyproject.toml pins blingfire==0.1.8, fasttext-wheel==0.9.2, s3fs==2023.6.0 and tokenizers>=0.15.0,<=0.19.1, and constrains numpy<2. Those upper bounds are where an upgrade will bite, because a newer tokenizers or a numpy 2 environment will conflict at install time rather than at runtime. The Python version range >=3.10,<3.13 is the other hard constraint.
On licensing: the toolkit is Apache-2.0, which is permissive and includes a patent grant. The dataset is ODC-BY, which requires attribution. The README asks that you cite the Dolma paper if you use either, and the citation block is in the README and in CITATION.cff. This is a description of what the files say, not legal advice; if attribution obligations matter to your distribution, read the ODC-BY text and the blog post the README links for the rationale behind the switch.
Editorial conclusion
Adopt Dolma if you are curating a large text corpus for pre-training and you want the taggers, deduplication and parallelism that AI2 already used for OLMo. Skip it if you only need to clean a few gigabytes of text, or if the documents you are filtering are not plain text with JSONL metadata. Before committing, verify that your Python version falls inside the >=3.10,<3.13 range declared in pyproject.toml, and check whether the built-in taggers match your language coverage, because the README names Gopher, C4 and OpenWebText and nothing about non-English text.
Frequently asked questions
What is Dolma?
Dolma is two things per its README: an open dataset of 3 trillion tokens for language model pre-training, and the Dolma Toolkit in this repository, which curates datasets for language modeling. The dataset is distributed on HuggingFace under ODC-BY; the toolkit is Apache-2.0.
How do I install the Dolma toolkit?
The README says to run pip install dolma. The package requires Python >=3.10,<3.13 according to pyproject.toml and compiles a Rust extension, so expect a build step during installation.
Which taggers does Dolma ship?
The README lists ready-to-use taggers for Gopher, C4 and OpenWebText. It also lists custom taggers as a supported extension point, but does not describe the interface in the README.
What licence does Dolma use?
The toolkit is Apache-2.0, per pyproject.toml and Cargo.toml. The dataset is licensed ODC-BY, and the README links a blog post explaining that choice.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/allenai-dolma)