RedPajama-Data-V2: A 30 Trillion Token Dataset Pipeline for LLM Training
The RedPajama-Data repository contains code for preparing large datasets for training large language models.
At a glance
- What is it?
- RedPajama-Data-V2 is an open dataset and processing pipeline that produced over 30 trillion tokens from 84 CommonCrawl snapshots across five languages. The repository contains the code to reproduce the three-stage pipeline: artifact preparation, quality signal computation, and deduplication.
- Who is it for?
- RedPajama-Data-V2 is a good starting point for teams that want to build or fine-tune a large language model on open web-crawl data without relying entirely on proprietary datasets. It is not a model, not a training recipe, and not a ready-to-use dataset in the sense of a neatly labeled benchmark: it is a processing pipeline and the output of running that pipeline.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 121 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What RedPajama-Data-V2 provides and who it targets
Training a large language model from scratch requires a large, diverse corpus of text. Building that corpus involves crawling the web, filtering low-quality text, removing duplicate documents, and computing signals that let downstream users decide how to weight each document. RedPajama-Data-V2 is both the output of this process and the code that produced it.
The dataset covers five languages: English, German, French, Italian, and Spanish. The annotated and deduplicated portion of the dataset, called `head_middle`, contains 20.8 billion documents and an estimated 30.4 trillion tokens. English makes up the largest share at 14.5 billion documents and 20.5 trillion tokens. The full dataset, before deduplication, includes over 100 billion documents from 84 CommonCrawl snapshots.
The intended users are machine learning research teams and companies who want a large, open corpus for LLM pre-training. The README directs users to HuggingFace to download the finished dataset if they do not want to run the pipeline. The repository itself is for teams who want to understand the processing steps, reproduce the dataset, or extend it.
Three-stage pipeline structure
The README divides the pipeline into three sequential steps. Step one creates artifacts used in later stages. Step two computes quality signals for each document. Step three runs deduplication. Each step can produce or consume data from S3-compatible object storage, and several steps use Apptainer (formerly Singularity) containers rather than Docker for execution on HPC clusters.
The pipeline assumes a Dockerized or Apptainerized environment. The README notes that steps can run without containers, but the provided scripts assume Docker and Apptainer are available. For step two, the `PYTHONHASHSEED` environment variable must be set to a consistent value:
export PYTHONHASHSEED=42This ensures that hash functions used in computing DSIR importance weights are consistent across runs. Forgetting this step would produce different hashes on different machines or Python versions, making the weights irreproducible.
The Docker image for the pipeline is built with:
. configs/default.conf
cd app
docker build -t "${DOCKER_REPO}:" .The configuration file at `configs/rp_v2.0.conf` sets the environment variables for the pipeline, including the Docker repository name, data root paths, and S3 configuration.
Creating artifacts: classifiers and quality models
The first pipeline step builds the tools used to score document quality. This includes training a bag-of-ngram generative model for DSIR importance weight computation, building quality classifiers, fetching a list of bad words from the LDNOOBW repository, and downloading a blacklist of URLs from the UT1 blacklist.
The English Wikipedia reference classifier is a pre-built fasttext model that must be downloaded separately. The README provides the download URL and the expected file path: `${DATA_ROOT}/wikiref-model/en/en-model.bin`. This is the same classifier used in RedPajama-V1.
Once the downloaded classifier is in place, the artifact creation script runs against a listing of CommonCrawl keys:
bash scripts/run_prep_artifacts.sh \
--config configs/rp_v2.0.conf \
--listings /path/to/listings/file.txt \
--max_workers 32The listings file contains keys like `2023-06/0000/en_head.json.gz`, which identify specific CommonCrawl partitions. The `--max_workers` flag controls the number of parallel processes. The script outputs an artifact ID that is stored in `ARTIFACTS_ID` for use in the next step.
Computing quality signals per document
The second step runs each document through the set of quality signals defined in the pipeline. These include heuristic filters (such as word count, sentence count, and proportion of repeated lines), classifier-based scores (such as the Wikipedia reference classifier), and minhash signatures used for fuzzy deduplication in step three.
The quality signals step runs via an Apptainer container:
bash scripts/apptainer_run_quality_signals.sh \
--config configs/rp_v2.0.conf \
--dump_id "2022-49" \
--input_base_uri "file:///path/to/data/root" \
--output_base_uri "file:///path/to/outout/data/root" \
--max_docs -1The `--dump_id` flag identifies which CommonCrawl dump to process. The input and output URIs accept both local file paths and S3 URIs. Setting `--max_docs -1` processes all documents in the dump.
The README lists a table of quality annotation tags. The blog post and HuggingFace dataset card contain the complete signal list. The signals are stored alongside each document in the output, so downstream users can filter or weight the dataset without reprocessing.
Deduplication with Bloom filters and locality-sensitive hashing
The third step removes duplicate documents using two complementary techniques. Exact deduplication uses a Bloom filter to check whether a document's content hash has been seen before. The Bloom filter implementation uses the `pybloomfiltermmap3` library and operates on S3-hosted data. The README warns that the capacity parameter must be set to a value greater than the number of documents, otherwise the configured error rate will not hold and more false positives will appear.
The exact deduplication command:
python3 app/src/bloomfilter.py \
--listings /path/to/listings/file.txt \
--input_base_uri "s3://path/to/ccnet/data" \
--output_dir "/path/to/output" \
--s3_profile "..." \
--endpoint_url "..." \
--parallel_readers 32 \
--batch_size 10 \
--capacity "..." \
--error_rate "..."Fuzzy deduplication uses locality-sensitive hashing (LSH) on the minhash signatures generated in step two. The implementation uses Polars and was tested on 200 million documents on a 64-core machine with 500 GB of RAM. That hardware requirement is a meaningful constraint: teams without large-memory servers will need to subset the data.
bash scripts/apptainer_run_lsh.sh \
--config configs/rp_v2.0.conf \
--dump_id "2022-49" \
--input_base_uri "file:///path/to/data/root" \
--output_dir "/path/to/output" \
--similarity "<similarity_threshold>" \
--listings "/minhash/listings/file.txt" \
--max_docs -1Infrastructure requirements and dataset limitations
Running the full pipeline against all 84 CommonCrawl snapshots requires significant infrastructure. The deduplication step was tested on a 64-core, 500 GB RAM machine for 200 million documents. Processing 20 billion documents would require proportionally larger resources or a distributed processing approach. The pipeline scripts assume S3-compatible storage for the raw CommonCrawl data, which means either AWS S3 or a compatible alternative.
The dataset covers only English, German, French, Italian, and Spanish. Teams working on languages outside this set would need to add language-specific classifiers and extend the listings. The README notes that the full list of quality annotation tags would grow as more signals are developed, but the baseline set is fixed in the current version.
The repository was last pushed on 2026-06-03. It has no GitHub releases. The code is open-source under Apache-2.0. The HuggingFace dataset itself is available under a separate license described on the dataset card.
RedPajama-Data-V2 versus RedPajama-V1 and other open datasets
RedPajama-V1 was the first release from the same team, targeting one trillion tokens to replicate the LLaMA training data. Its code lives on the `rp_v1` branch of the same repository. V2 replaces V1 in scale: 30 trillion tokens versus one trillion, and coverage of 84 CommonCrawl snapshots versus the handful used in V1.
The Pile is another large open dataset for LLM training, assembled by EleutherAI. It draws from a curated list of sources rather than CommonCrawl alone, which gives it more diversity per token but also a smaller total size. C4 (Colossal Clean Crawled Corpus) is a single CommonCrawl snapshot processed by Google and is widely used for T5-family models. RedPajama-Data-V2 is larger than either and includes quality signals that let users decide their own filtering strategy rather than accepting a fixed filter.
For a team that wants a large, open corpus with known provenance and documented quality signals, RedPajama-Data-V2 is a reasonable choice. The pipeline code is Apache-2.0, and the HuggingFace dataset is available for direct download without running the pipeline.
Editorial conclusion
RedPajama-Data-V2 is a good starting point for teams that want to build or fine-tune a large language model on open web-crawl data without relying entirely on proprietary datasets. It is not a model, not a training recipe, and not a ready-to-use dataset in the sense of a neatly labeled benchmark: it is a processing pipeline and the output of running that pipeline. Teams that only need a pre-processed, ready-to-train slice can download the annotated and deduplicated portion from HuggingFace without running the pipeline at all. Teams who want to add new quality signals, extend the language coverage, or process new CommonCrawl snapshots will need the pipeline code and a large-scale compute environment.
Frequently asked questions
How do I download the RedPajama-Data-V2 dataset?
The README directs users to the HuggingFace dataset at huggingface.co/datasets/togethercomputer/RedPajama-Data-V2 for direct download without running the pipeline. The annotated and deduplicated head_middle portion with 20.8 billion documents is the primary release.
What languages does RedPajama-Data-V2 cover?
The dataset covers five languages: English, German, French, Italian, and Spanish. English is the largest portion with 14.5 billion documents and an estimated 20.5 trillion tokens.
What is the difference between RedPajama-V1 and RedPajama-V2?
RedPajama-V1 targeted one trillion tokens to replicate LLaMA training data, and its code is on the rp_v1 branch of the repository. V2 covers 84 CommonCrawl snapshots and includes 30 trillion tokens with per-document quality signals for downstream filtering.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/togethercomputer-redpajama-data)