Model or dataset
OpenDCAI/DataFlow avatar
OpenDCAI/DataFlow

OpenDCAI/DataFlow: LLM Data Preparation Through Operators and Pipelines

Easy Data Preparation with latest LLMs-based Operators and Pipelines.

8,211 stars1,179 forksPythonApache-2.0

At a glance

What is it?
DataFlow is an Apache-2.0 Python package that turns PDF, plain-text and low-quality QA sources into training data using composable operators. The design is sound for pipeline reuse; the documentation and packaging details are the parts to check before you commit.
Who is it for?
Adopt DataFlow if you are preparing domain corpora for pre-training, supervised fine-tuning, RL training or RAG and you want the cleaning steps expressed as reusable operators rather than one-off scripts. Skip it if your data work is small enough to do with pandas, or if you need a stable 1.x API surface, since the package still classifies itself as Alpha.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 10 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem DataFlow targets: raw PDFs and low-quality QA into training data

Most teams that fine-tune a model spend more time on the corpus than on the training loop. The raw material arrives as PDFs, plain text dumps and question-answer pairs of uneven quality, and the cleaning work is usually written as throwaway scripts that nobody can rerun six months later. DataFlow takes that work and expresses it as a pipeline of operators, so the same cleaning sequence can be applied to a new corpus, shared with another team, or recombined into a different pipeline.

The README frames the audience directly: the system is for generating, refining, evaluating and filtering data for AI, with the stated goal of improving LLM performance in specific domains through pre-training, supervised fine-tuning, RL training or a RAG system. The domains named are healthcare, finance, legal and academic research. That list matters, because those are the settings where a general-purpose cleaning script tends to fail: a legal PDF and a clinical note need different extraction and different filtering, and a pipeline abstraction lets you swap the operator rather than the whole script.

The second audience is the one building tooling on top. The README describes an operator-based design whose purpose is to make the cleaning workflow reproducible, reusable and shareable, and it positions the project as core infrastructure for the Data-Centric AI community. If you are writing internal data tooling, that is the part worth evaluating, not the individual cleaning steps.

How the operator and pipeline model actually fits together

The unit of work is the operator. An operator is a step that does one thing to a dataset, and the README's phrasing is that the design turns the cleaning workflow into a reproducible, reusable and shareable pipeline. A pipeline is a sequence of those operators, which is why the project can claim the same cleaning sequence is reusable across domains: the sequence is data, not code.

On top of that sits the DataFlow-agent, which the README describes as capable of dynamically assembling new pipelines by recombining existing operators or creating new ones on demand. That is a different proposition from a fixed pipeline library. A fixed library gives you the pipelines the maintainers wrote; the agent is meant to compose one for your case. Treat the two as separate bets. The operator abstraction is the stable part, and the agent is the part whose output you should inspect before you train on it.

The repository layout supports the description. The package lives under dataflow/, tests under test/, and the project ships a CLI entry point, since pyproject.toml declares `dataflow = "dataflow.cli:app"` under [project.scripts]. The Dockerfile shows the intended runtime more concretely than the README does: it builds on nvidia/cuda:12.4.1-cudnn-runtime-ubuntu22.04, creates a virtualenv at /opt/venv, copies the source to /app/DataFlow, and installs with `pip install -e ".[vllm]"`. The vllm extra is the one the Dockerfile uses, and requirements.txt also lists sglang-related and distributed dependencies, so the inference backend is a choice you make at install time rather than something the package picks for you.

Installing open-dataflow and running a first pipeline

The package is published on PyPI as open-dataflow, which is the name you install even though the import package and the CLI are both called dataflow. The README points at the documentation site at OpenDCAI.github.io/DataFlow-Doc/ for the full walkthrough, and pyproject.toml shows the console script entry point.

A plain install pulls the dependency set from requirements.txt, which is long: torch, transformers, datasets, pyarrow pinned to 20.0.0, pypdf, and a knowledge-base cleaning group that includes chonkie, trafilatura and pymupdf.

bash
pip install open-dataflow

If you are running against a local vLLM server, the Dockerfile shows the extra the project itself uses. The same extra name is what you would pass on a normal machine, though the CUDA base image in that Dockerfile is not something a pip install reproduces.

bash
pip install -e ".[vllm]"

The README also documents a WebUI that starts from the command line, added in the 2026-02-02 release note. The documented command is a single word:

bash
dataflow webui

That launches the visual pipeline builder described in the WebUI docs section of the README, where you build and run pipelines through a browser instead of writing the sequence yourself. For a first real use, the WebUI is the lower-risk entry point: you can see which operators exist and what each one expects before you commit to a pipeline definition. The Colab notebook linked from the README is the other zero-setup path, and the repository also publishes a Docker image on Docker Hub under molyheci/dataflow.

One packaging detail to watch: pyproject.toml declares requires-python as ">=3.7, <4" while the classifiers only list 3.10 through 3.12, and requirements.txt pins pyarrow==20.0.0 with a comment that a larger version bugs on Python 3.10. The metadata and the classifiers disagree, so pick a version from the classifier list rather than the range.

Where DataFlow is the wrong tool, and what the docs leave open

The clearest limitation is stated by the project itself. pyproject.toml carries the classifier "Development Status :: 3 - Alpha". That is the maintainers' own label, and it should shape how you depend on the API. If you need a stable interface with deprecation guarantees, this is not that yet.

The dependency footprint is the second constraint. requirements.txt pulls torch, torchvision and torchaudio, plus transformers, accelerate and a separate group for the data agent that includes fastapi, uvicorn and cloudpickle. The Dockerfile installs ffmpeg, libgl1 and libglib2.0-0 at the system level. If your cleaning job is a few thousand rows of text, you are paying for an inference stack to do work that a pandas script would finish first.

The third issue is documentation depth, and it is the one I would weigh most. The README explains what the system is for and links to the documentation site, but it does not enumerate the operators, show a pipeline definition, or document how a pipeline is serialized and reloaded. The README does not document rollback or versioning of a saved pipeline either. The repository does ship DataFlow-Skills and DataFlow-WebUI as separate projects, and the 2026-07-18 release note says the WebUI is where you can use DataFlow through either the visual interface or MCP, so the practical documentation is spread across repositories rather than concentrated in one place. Budget time for reading operator source, not just the README.

Finally, the name is a real cost. Searching for "dataflow" returns Power BI dataflows, Google Cloud Dataflow, and an unrelated verification service of the same name. If you are filing issues or looking for prior art, search for OpenDCAI or open-dataflow instead.

How DataFlow differs from Power BI dataflows and Apache Beam

The name collision hides a genuine difference in approach, and it is worth being precise about it.

Power BI dataflows are a Microsoft Fabric feature for reusable tables inside the Power BI ecosystem. They exist to serve reporting, and their operators are Power Query transformations. DataFlow has nothing to do with that pipeline: its operators act on text and documents for the purpose of producing LLM training data, and its output is a dataset, not a semantic model. If you arrived here from a search about Power BI or about dataflow gen2, you are in the wrong repository.

Apache Beam is the closer comparison and the more useful one. Beam gives you a programming model over a runner, so a pipeline is code that the runner executes across a cluster, and the abstraction is the transform. DataFlow's abstraction is the operator, and the pipeline is a composition of operators that the README describes as reproducible and shareable. The difference shows up in what you get for free. Beam gives you a mature runner and a well-defined windowing and trigger model; DataFlow gives you operators aimed at LLM data specifically, such as the knowledge-base cleaning group in requirements.txt (chonkie, trafilatura, pymupdf) and the speech group (librosa, soundfile). Choosing Beam means writing the document extraction and quality filtering yourself. Choosing DataFlow means accepting an Alpha API in exchange for those steps already existing.

There is also DataFlow-Harness, released 2026-07-18, which the README says lets coding agents build DataFlow pipelines for you, with entry points in DataFlow-WebUI. That is an adjacent tool rather than an alternative, and it inherits the same caveat as the DataFlow-agent: generated pipelines need review.

Maintenance, licensing and the upgrade path

The repository is not archived, and the last push was on 2026-09-10, which is recent. The release cadence visible in the release notes is roughly monthly to quarterly: v1.0.10 on 2026-03-26, v1.0.9 on 2026-02-27, and v1.0.8 on 2025-12-19. The gap between the last release and the last push is several months, so code is moving between tagged releases. If you pin, pin to a tag rather than to main.

The licence is Apache-2.0, declared both in the repository and in pyproject.toml. That is a permissive licence, and it is the reason the project can be used in commercial data pipelines. One inconsistency to note rather than act on: pyproject.toml also carries the classifier "License :: Free For Educational Use", which does not describe Apache-2.0 and appears to be a leftover. The LICENSE file at the repository root is the reference point. I am not a lawyer and this is not legal advice; if the classifier discrepancy matters to your organisation, have counsel read the LICENSE file.

Upgrade cost is dominated by the dependency pins rather than by the DataFlow API. requirements.txt pins pyarrow==20.0.0 with an explicit note that a larger version bugs on Python 3.10, and pyproject.toml declares an optional test group that pins setuptools<=81.0.0. Those pins will conflict with other packages in a shared environment before DataFlow's own operators change. Isolate it in a virtualenv or use the published Docker image.

Editorial conclusion

Adopt DataFlow if you are preparing domain corpora for pre-training, supervised fine-tuning, RL training or RAG and you want the cleaning steps expressed as reusable operators rather than one-off scripts. Skip it if your data work is small enough to do with pandas, or if you need a stable 1.x API surface, since the package still classifies itself as Alpha. Before you build on it, verify the extra name that matches your inference backend, confirm the Python version you intend to run, and read the operator source for the steps you plan to depend on, because the README describes the system rather than the API.

Frequently asked questions

Is DataFlow free to use?

The repository and pyproject.toml both declare Apache-2.0, a permissive licence, so the package itself is free to use. Note that pyproject.toml also carries a "License :: Free For Educational Use" classifier that does not match Apache-2.0; the LICENSE file at the repository root is the reference.

How much does DataFlow cost?

There is no price for DataFlow itself, since it is distributed under Apache-2.0 on PyPI as open-dataflow. Your real cost is the runtime: requirements.txt pulls torch, transformers and accelerate, and the Dockerfile installs CUDA, so inference hardware or API spend is the expense.

How to use DataFlow?

Install the package as open-dataflow, then either run the documented `dataflow webui` command to build pipelines in the browser or compose operators into a pipeline in Python. The README points to the documentation site at OpenDCAI.github.io/DataFlow-Doc/ and to a Colab notebook for the full walkthrough.

Official sources

  1. License: Apache-2.0
  2. OpenDCAI/DataFlow on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/opendcai-dataflow.svg)](https://hysenlabs.com/projects/opendcai-dataflow)