Data-Juicer: a YAML-configurable operator pipeline for LLM and multimodal data
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
At a glance
- What is it?
- Data-Juicer packages 200+ data operators into reproducible YAML recipes and runs them on Ray, from a laptop to a cluster. It is built for teams preparing pre-training, fine-tuning, agent and RAG data, and its main cost is the breadth of its own configuration surface.
- Who is it for?
- Adopt Data-Juicer if you already have a data corpus and need repeatable, versioned cleaning and filtering steps rather than a one-off script, and if you are willing to run Ray for anything large. Do not adopt it if your pipeline is a single regex pass over a few thousand rows, or if you cannot take a dependency on the pinned generic extras such as torch 2.8.0 and vllm 0.11.0.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 14 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Data-Juicer targets: raw corpora that need repeatable cleaning
Most teams preparing data for a foundation model start with a script. It reads JSONL, filters by length, strips HTML, deduplicates, and writes JSONL back. The script works once. Then a second person adds a language filter, someone else changes the length threshold, and nobody can reproduce the corpus that produced last month's checkpoint. Data-Juicer is aimed at that gap: it turns the individual cleaning steps into named operators and the sequence into a YAML recipe that can be versioned and shared like code.
The README frames this as treating data processing as composable infrastructure, and the repository backs the claim with a specific number: 200+ operators spanning text, image, audio, video, and multimodal data. The audience is not the person writing a one-off script. It is a team that runs the same curation repeatedly across pre-training corpora, fine-tuning sets, agent interaction traces, or RAG indices, and needs the pipeline itself to be an artifact.
There is a second audience implied by the repository layout. The presence of app.py, service.py and a label_studio_localhost_connection.json suggests a workflow where human annotation and a served interface sit alongside the batch pipeline. That is a heavier commitment than a filtering library, and it is worth knowing before you start.
How the operator pipeline actually runs
The core abstraction is the operator, or OP. Each OP does one thing: TextLengthFilter drops short records, WhitespaceNormalizationMapper rewrites text, image_ohem_selector picks high-loss image samples. OPs are grouped by role in the package layout under data_juicer/ops, with filter and mapper among the visible subpackages.
The Python path makes the data flow concrete. A NestedDataset is built from a dict of columns, then ds.process() is handed a list of OP instances and returns a processed dataset. Input records go in, each OP in the list is applied in order, and the surviving records come out. That is the whole model: a dataset and an ordered list of transformations.
The YAML path is the same model with the list written down. A recipe file names the dataset config and the pipeline of OPs, and dj-process executes it. This is the form that gets versioned. Release v1.6.0 added config validation, described as a preflight that catches invalid operator settings and executor or schema mismatches before processing starts. That matters more than it sounds: without preflight, a typo in an operator argument surfaces after the job has already read part of the corpus.
Execution is where the design gets opinionated. The README describes scale in terms of Ray nodes and cores, and v1.6.0 added cluster-aware partitioning, where automatic partition counts use live Ray cluster resources. Manual partition.size targets split data at row boundaries, including inputs with fewer blocks than partitions. If you are not running Ray, you are using a subset of the intended deployment shape.
Installing Data-Juicer and running a first recipe
The README gives a two-line quick start. The package on PyPI is named py-data-juicer, and the console entry point is dj-process.
uv pip install py-data-juicer
dj-process --config demos/process_simple/process.yamlThe first line installs the package. The second runs the bundled demo recipe at demos/process_simple/process.yaml, which is the smallest end-to-end example in the repository. The demo directory is part of the repository, so if you installed from PyPI without cloning, point --config at a recipe path you control instead.
If you would rather compose the same idea in Python, the README shows this example:
from data_juicer.core.data import NestedDataset
from data_juicer.ops.filter import TextLengthFilter
from data_juicer.ops.mapper import WhitespaceNormalizationMapper
ds = NestedDataset.from_dict({
"text": ["Short", "This passes the filter.", "Text with spaces"]
})
res_ds = ds.process([
TextLengthFilter(min_len=10),
WhitespaceNormalizationMapper()
])
for s in res_ds:
print(s)Three strings go in. TextLengthFilter(min_len=10) removes the first one, and WhitespaceNormalizationMapper collapses the repeated spaces in the third. Iterating over res_ds prints what survived. That is the fastest way to check that an OP does what you expect before you put it in a recipe.
For a container, the repository ships a Dockerfile built from nvidia/cuda:12.6.3-cudnn-devel-ubuntu24.04, installing Python 3.11 and uv, and the README points at the datajuicer/data-juicer image on Docker Hub. The image is described as installing data-juicer in editable mode, which is convenient for development and not what you want for a pinned production run.
Where Data-Juicer is the wrong tool
The dependency footprint is the first real constraint. The core install pulls datasets, pandas, numpy, pydantic, streamlit, Pillow, and a list of compression and parsing libraries. The optional generic extra pins torch==2.8.0, transformers==4.57.1 and vllm==0.11.0. If your project already pins a different PyTorch or vLLM version, installing the generic extra will fight your environment. There is no documented path in the README for running the GPU-backed OPs against a different torch build.
Scale is the second constraint, and it cuts the other way. The headline numbers in the README are cluster numbers: 70B samples in 2h on 50 Ray nodes, 5TB deduplicated in 2.8h on 1280 cores. Nothing in the README describes the single-machine experience for a 10,000-row file, and adopting Ray to clean a small file adds a scheduler and a partitioning model you do not need. For a corpus that fits in memory, a script with pandas and a few regexes is less machinery for the same result.
The third constraint is the operator surface itself. With 200+ operators, the hard part stops being writing the transform and becomes choosing the right one and configuring it correctly. Release v1.6.0 documents 28 previously undocumented operators, which tells you the documentation was lagging the code. Config validation reduces the cost of a wrong argument, but it does not tell you whether the operator's semantics match your intent. That is still on you to check on a sample.
Finally, the export path has a boundary. v1.6.0 unified local, S3 and HDFS export under a shared filesystem dispatch. If your data lives somewhere else, the README does not describe a supported export target for it.
How Data-Juicer differs from writing your own pipeline
The obvious alternative is not another framework. It is the script you would write yourself, using pandas or Hugging Face datasets plus a handful of functions. That approach has real advantages: no Ray, no YAML schema, no dependency pins, and total control over the semantics of each step.
The difference is in what you get back. A hand-written script encodes the pipeline in control flow, so changing the order of steps means editing code, and reproducing an old corpus means checking out an old commit of the script and hoping its dependencies still resolve. Data-Juicer encodes the pipeline in a recipe file, so the sequence of OPs is data you can diff, fork and share. The project's own framing is that recipes can be versioned and shared like code, and the companion data-juicer-hub repository is described as holding 50+ recipes.
The second difference is the operator catalogue. Writing a MinHash deduplicator, a token-count filter with bounded tokenizer batches, or an image OHEM selector yourself is a project. Here they are OPs you configure. The cost is that you inherit their semantics and their bugs. v1.6.0 lists fixes for fused-filter cache isolation, MinHash state reuse and empty inputs, and deduplicator execution-mode declarations. Those are the kinds of defects that sit quietly in a deduplication step and change your corpus without an error.
The third difference is the natural-language layer. The Juicer model, described as a 35B-A3B model available on HuggingFace and ModelScope, turns cleaning instructions and filtering rules into structured outputs. That is a different way to author a pipeline than editing YAML by hand, and it is the part of the project with the least independent track record.
Maintenance, licence and upgrade cost
The repository is not archived, and the last push was on 2026-09-09. Releases have been frequent and substantive: v1.6.0 on 2026-09-09, v1.5.5 on 2026-08-07, v1.5.4 on 2026-07-23. Each carries named features and a list of fixes rather than only version bumps.
Upgrading is not free. v1.6.0 changed reader defaults so they apply consistently across execution and analysis, which is the kind of change that can shift results without raising an error. It also added config validation, so a recipe that previously ran may now fail preflight because an operator argument was always wrong. Budget time to re-run a sample through the pipeline after each upgrade and compare record counts, not just exit codes.
The licence is Apache-2.0, declared in both the LICENSE file and pyproject.toml. That is a permissive licence with an explicit patent grant, and it does not impose copyleft obligations on your own code. It says nothing about the licences of the datasets you process or of the optional dependencies you install, and those are separate questions. This is a description of what the repository declares, not legal advice.
One governance detail worth noting: the project is authored by the SysML Team of Alibaba Tongyi Lab, and the README states that Alibaba Cloud PAI has integrated Data-Juicer into its data processing products. The open source project and the hosted product are not the same thing, and the README does not describe a support relationship between them.
Editorial conclusion
Adopt Data-Juicer if you already have a data corpus and need repeatable, versioned cleaning and filtering steps rather than a one-off script, and if you are willing to run Ray for anything large. Do not adopt it if your pipeline is a single regex pass over a few thousand rows, or if you cannot take a dependency on the pinned generic extras such as torch 2.8.0 and vllm 0.11.0. Before committing, verify two things against your own data: that the operator you intend to use behaves as its documentation describes on a small sample, and that your storage backend is one of the ones the export path actually dispatches to (local, S3, HDFS).
Frequently asked questions
What is Data-Juicer and what does it do?
It is a Python data processing system for foundation models, built around 200+ operators that clean, synthesize and analyze text, image, audio, video and multimodal data. Operators are composed into YAML recipes that can be versioned and shared.
What is an example of data processing with Data-Juicer?
The README shows building a NestedDataset from a dict of text, then calling process() with TextLengthFilter(min_len=10) to drop short records and WhitespaceNormalizationMapper() to collapse repeated spaces. The processed dataset is what remains after the operators run in order.
What tools are used for data processing in Data-Juicer?
The core install depends on datasets, pandas, numpy, pydantic, streamlit and several compression and parsing libraries. The optional generic extra adds torch, transformers and vllm for model-backed operators, and execution at scale uses Ray.
Can AI be used for data processing in Data-Juicer?
The project ships Juicer, described as a natural-language data-refinement model that turns cleaning instructions, filtering rules and semantic-tagging requirements into structured outputs. It is offered on HuggingFace and ModelScope as Juicer-35B-A3B.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/datajuicer-data-juicer)