Model or dataset
data-prep-kit/data-prep-kit avatar
data-prep-kit/data-prep-kit

Data-Prep-Kit: a transform toolkit for LLM data preparation

Open source project for data preparation for GenAI applications

964 stars255 forksHTMLApache-2.0

At a glance

What is it?
Data-Prep-Kit packages ingestion, deduplication and filtering transforms for LLM training and RAG data into Python and Ray runtimes. It suits teams that need repeatable Parquet-based pipelines more than one-off cleaning scripts.
Who is it for?
Adopt Data-Prep-Kit if you already store training or RAG data as Parquet, NDJSON, JSONL or ZIP and want dedup, filtering and profiling as repeatable transforms instead of ad hoc scripts. Do not adopt it if your data lives in a warehouse and you only need SQL, or if you need a visual data preparation tool.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly HTML, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Data-Prep-Kit actually solves

Most teams preparing data for a model end up writing the same scripts twice: once to parse PDFs, HTML or ZIP archives into records, and again to deduplicate and filter those records before training or retrieval. Data-Prep-Kit turns those steps into a catalog of modules that share input and output conventions. The README describes the goal plainly: developers use it to "cleanse, transform, and enrich use case-specific unstructured data to pre-train LLMs, fine-tune LLMs, instruct-tune LLMs, or build Retrieval Augmented Generation (RAG) applications".

The target user is an engineer who has raw files and needs a repeatable pipeline, not a notebook that runs once. The kit supports Natural Language, Code and Image modalities today, and it ships recipes under examples/ for end-to-end cases such as PDF processing and RAG. The unit of work is a transform, and transforms compose into pipelines.

How transforms, runtimes and pipelines fit together

The mechanism is a transform that reads files, applies one operation, and writes files. The README states that the kit provides a framework for developing custom transforms for processing Parquet files as well as ZIP, NDJSON and JSONL formats. That file-format boundary is the contract: ingestion transforms such as code2parquet, docling2parquet, html2parquet and web2parquet convert source material into Parquet, and universal transforms such as ededup, fdedup, doc_id, filter and profiler operate on those records.

Runtime choice is explicit. The supported-transform matrix marks each module as Python-only, Ray, or both. Code to Parquet, Docling to Parquet, HTML to Parquet, exact dedup, fuzzy dedup, unique ID annotation, filter and profiler all carry both Python and Ray checkmarks in the table shown in the README, while Web to Parquet is listed as Python-only. That asymmetry matters: a pipeline that mixes web2parquet with a Ray runtime cannot simply swap execution engines for every stage.

Scaling is handled at the deployment layer rather than inside the transform. The README says the kit provides examples of how a single transform can be deployed on Kubernetes clusters as a Python or a Ray job, and that when multiple transforms are deployed in sequence it uses Tekton pipelines. So the same transform code runs on a laptop or in a cluster; what changes is the orchestrator around it.

Installing Data-Prep-Kit and running a first transform

The README points to PyPI and supports Python 3.10 through 3.13. The install command uses uv and the extras syntax to pull every transform:

bash
pip install uv
uv pip install 'data-prep-toolkit-transforms[all]'

After that, the README directs readers to the quick-start guide under doc/quick-start/quick-start.md for creating a virtual environment. For a first run without any local setup, the README offers a Google Colab notebook, examples/notebooks/Run_your_first_transform_colab.ipynb, described as a simple transform that extracts content from PDF files. The README notes the same notebook can be downloaded and run locally without cloning the repository.

Once a single transform works, the next step in the README is composing transforms into a pipeline for real use cases, using the recipes in the examples/ directory. Custom transforms are covered by a developer tutorial at doc/quick-start/contribute-your-own-transform.md, and advanced topics such as running transforms from the command line and scaling are in ADVANCED.md. The README also notes that all transforms include small sample data files for testing, and that downloading real HuggingFace data for tests is covered in ADVANCED.md.

Where the toolkit stops short

The first limitation is visible in the README itself: the matrix lists Web to Parquet as Python-only, so a team that standardizes on Ray for scale cannot run that ingestion step the same way as the rest. Mixed-runtime pipelines are a real constraint, not a footnote.

Second, the supported modality list is described as what the kit supports "today": Natural Language, Code and Image. If your data is audio, video or structured tables, nothing in the README suggests a transform for it.

Third, the project is a library of transforms plus deployment examples, not a managed service. The README documents PyPI installation, Colab, Kubernetes job examples and Tekton pipelines. It does not document a hosted control plane, a scheduler UI, or automatic retries for failed pipeline stages. Teams that want a service they can point at a bucket and forget will be building that layer themselves. The README is also silent on rollback behaviour for partially completed pipelines, which is the kind of thing you discover only when a long dedup job fails halfway.

How it compares with notebook-based cleaning and full platforms

The closest thing most teams already have is a set of pandas or Spark scripts in notebooks. The difference in approach is where the logic lives. A notebook script encodes parsing, dedup and filtering inline, so reuse means copying cells. Data-Prep-Kit moves each step into a named transform with a declared input and output format, and the README's recipes folder shows those transforms assembled into pipelines for fine-tuning and RAG. That structure is what allows the same transform to be submitted as a Python job or a Ray job on Kubernetes, and to be chained through Tekton.

The trade-off is overhead. A one-off CSV cleanup does not need a transform framework, a Parquet contract, or a pipeline orchestrator. Data-Prep-Kit is the wrong tool when the data fits in memory and the cleaning logic will never run again. It is also a poor fit when the real problem is querying a warehouse; the kit works on files, and the README's format list (Parquet, ZIP, NDJSON, JSONL) reflects that.

Maintenance, releases and what the licence allows

The repository is not archived, and the last push was on 2026-09-08, which is recent. Recent releases are v1.1.8 on 2026-06-22, v1.1.7 on 2026-02-11 and v1.1.6 on 2025-11-14. That cadence, roughly one minor release every few months, is what a team pinning a version should plan around: budget for reading release-notes.md and re-running pipeline tests when you upgrade, especially if you have written custom transforms against the framework.

The licence is Apache-2.0, per the repository metadata and the badge in the README. Apache-2.0 permits commercial use and modification and includes an explicit patent grant, which matters if you embed the transforms in a product. It also requires that you keep the licence and notices when redistributing. That is a description of the licence text, not legal advice; have counsel review redistribution plans. The project is listed under LF AI & Data, and the README links to OpenSSF Best Practices, which indicates a governance structure beyond a single maintainer, though the README does not describe how decisions are made.

Editorial conclusion

Adopt Data-Prep-Kit if you already store training or RAG data as Parquet, NDJSON, JSONL or ZIP and want dedup, filtering and profiling as repeatable transforms instead of ad hoc scripts. Do not adopt it if your data lives in a warehouse and you only need SQL, or if you need a visual data preparation tool. Before committing, verify that the transform you need supports your runtime (Python or Ray), check whether your deployment target is covered by the Tekton and Kubernetes examples, and read the release notes for the version you pin.

Frequently asked questions

What is Data-Prep-Kit?

It is an open source project for preparing unstructured data for GenAI applications. The README describes it as a kit for cleansing, transforming and enriching use case-specific data to pre-train, fine-tune or instruct-tune LLMs, or to build RAG applications.

What does data preparation mean in the context of Data-Prep-Kit?

In this project it means running transforms over files: ingestion steps such as HTML to Parquet or Docling to Parquet, then universal steps such as exact dedup, fuzzy dedup, unique ID annotation, filter and profiler. The transforms can be combined into pipelines, as shown in the examples folder.

Can you give an example of data preparation with Data-Prep-Kit?

The README points to examples/notebooks/Run_your_first_transform_colab.ipynb, a Colab-friendly notebook that runs a simple transform to extract content from PDF files. The examples folder also contains end-to-end recipes for PDF processing and RAG.

Official sources

  1. data-prep-kit/data-prep-kit on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/data-prep-kit-data-prep-kit.svg)](https://hysenlabs.com/projects/data-prep-kit-data-prep-kit)