Model or dataset
OpenDCAI/DataFlex avatar
OpenDCAI/DataFlex

OpenDCAI/DataFlex: dynamic data selection, mixture and reweighting inside the LLaMA-Factory training loop

Data-centric LLM training with dynamic sample selection, domain mixture optimization, and example reweighting inside the LLaMA-Factory training loop.

2,877 stars405 forksPythonApache-2.0

At a glance

What is it?
DataFlex is a Python framework layered on LLaMA-Factory that schedules training samples, domain ratios and per-sample weights while the model is still training. It is research infrastructure, and its install path is opinionated about torch and DeepSpeed versions.
Who is it for?
Adopt DataFlex if you already train with LLaMA-Factory and want LESS, DOREMI, ODM or the loss-based selectors available behind one YAML config instead of stitching together separate research repositories. Do not adopt it if you are not on LLaMA-Factory, or if you need a stable API: pyproject.toml classifies the project as Beta and the README states your YAML must carry DataFlex-specific parameters that vanilla configs do not have.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem DataFlex targets: data scheduling as a training-time decision

Most training pipelines fix the dataset before the first step. You pick a mixture, filter or sample the corpus, then hand a static file list to the trainer. DataFlex takes the opposite position: which samples the model sees, in what proportion per domain, and with what weight in the loss should be decided while optimization is running. The README frames this as Data Select, Mix, Reweight, and describes the project as a dynamic training framework built on top of LLaMA-Factory that integrates several difficult-to-reproduce repositories into one framework.

The intended audience is narrow and identifiable. It is researchers and engineers who already run LLaMA-Factory and want to compare selection strategies under one harness, plus people reproducing papers such as LESS, NICE, TSDS, DOREMI or ODM. The README's own tables list those methods with a Requires Model-in-the-Loop column, which is the real dividing line. Gradient-based and loss-based selectors need the model to produce signal during training; NEAR and TSDS do not. If your method needs no model feedback, DataFlex is a heavier wrapper than the problem requires.

How the selector, mixer and weighter hook into the LLaMA-Factory loop

DataFlex does not replace the trainer. It sits inside it. The README states the launch command is similar to LLaMA-Factory and that the framework is a drop-in replacement, with the important caveat that your .yaml config must also include DataFlex-specific parameters. So the dataflow is: LLaMA-Factory parses the YAML and drives the training loop as usual, while DataFlex reads its extra keys and decides, per step or per epoch, which samples enter the batch, how domains are mixed, and how gradients are scaled.

The three algorithm families map onto three intervention points. Data Selection picks training samples according to a strategy, for example focusing on hard examples; the README groups the methods into gradient-based (LESS, NICE), loss-based (Loss, Delta Loss), distribution-based (NEAR, TSDS) and no-selection (Static, Random). Data Mixture adjusts the ratio of data from different domains during training, with DOREMI as an offline method and ODM as online. Data Reweighting adjusts sample weights during backpropagation, with Loss Reweighting and Joint-Update-Aware Reweighting, the latter described as batch-aware.

The repository layout supports that split: src/ holds the package, examples/ is divided into train_lora, train_full, deepspeed, accelerate, merge_lora and test, and skills/ contains two documents, how_to_use.md and how_to_add_algorithm.md. The second one describes a registry system and base class interfaces for adding selectors, mixers or weighters, which is the extension point to read before writing your own method. A March 2026 news entry states that gradient computation under DeepSpeed ZeRO-3 is now supported, which matters because the gradient-based selectors need that path to work at scale.

Installing DataFlex and running a first LESS training job

The README is explicit that torch should be installed first, so pip does not resolve a different build and then have to replace it. This block pins the CUDA 12.4 wheels and then prints the version and whether CUDA is visible:

bash
pip install --index-url https://download.pytorch.org/whl/cu124 \
    torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0
python -c "import torch; print(torch.__version__, torch.cuda.is_available())"

With torch in place, the package itself is a single pip install. The alternative for development is a source checkout with an editable install:

bash
pip install dataflex
# or, from source:
git clone https://github.com/OpenDCAI/DataFlex.git
cd DataFlex
pip install -e .

Every config under examples/deepspeed needs DeepSpeed, which ships as an optional extra. The README notes the extra stays below 0.17 on purpose because 0.17+ fails to import on the pinned torch. The LESS selector needs TRAK, a separate and heavy extra:

bash
pip install "dataflex[deepspeed]"   # from source: pip install -e ".[deepspeed]"
pip install dataflex[less]

The entry point is the dataflex-cli script declared in pyproject.toml. The README's example launches a LoRA training run driven by a LESS selector config:

bash
dataflex-cli train examples/train_lora/selectors/less.yaml

What you should see is a LLaMA-Factory style training log, with the difference that the config file supplies the DataFlex-specific keys. The README does not show a complete example YAML inline, and it points to DataFlex-Doc for the parameter reference. Treat that documentation as required reading before your first run rather than optional.

Version constraints and the DeepSpeed cap are the real adoption cost

The dependency window is the least flexible part of this project. requirements.txt pins transformers to >=4.55.0, !=4.57.0, <=5.6.0, notes that Qwen3.5 and Gemma 4 need 5.5.0 or newer, and recommends 5.3.0+ if you need train_from_scratch under DeepSpeed ZeRO-3. The README repeats the same guidance. That means a team on an older transformers stack cannot simply add DataFlex; they upgrade, and the upgrade may touch their other tooling.

The DeepSpeed situation is a deliberate trade-off with a cost. The extra is capped at >=0.15.0,<0.17.0 because 0.17+ cannot import on the torch 2.6 build the README pins, and pyproject.toml carries the same comment. The README does say that on a newer torch you are free to lift that cap, so the constraint is coupled to the torch choice rather than permanent. Still, anyone who needs a DeepSpeed feature added after 0.16 has to move torch first and validate the combination themselves.

There is a second class of limitation in the algorithms themselves. The README's tables use a warning symbol for the official repositories of LESS, NICE, DOREMI and ODM, and a cross where no official repository exists, which is the case for Loss, Delta Loss, NEAR, Static, Random and both reweighting methods. DataFlex's value proposition is that it reimplements these reproducibly in one place, but a reimplementation is not the reference code. If your work depends on matching a published number exactly, verify against the original where one exists. The README does not document rollback or how to revert a run to static training mid-experiment beyond switching the selector, so plan your comparisons as separate runs.

Where DataFlex is the wrong tool

If you are not on LLaMA-Factory, this is the wrong layer. DataFlex is built on it, requires llamafactory>=0.9.5, and the README describes full compatibility as a drop-in replacement rather than a standalone trainer. A team using a custom PyTorch loop or a different fine-tuning stack would be adopting LLaMA-Factory and DataFlex together, which is a much larger change than adding a library.

Small runs are also a poor fit. The model-in-the-loop methods need the model to produce loss or gradient signal that the scheduler consumes, and the gradient-based ones need that signal to flow under the parallelism you have configured. If you are fine-tuning a small model on a single GPU with a fixed, already-curated dataset, Static or Random selection is what you would end up with anyway, and the extra configuration surface buys nothing.

Finally, the project is classified as Beta in pyproject.toml, and the README states plainly that unlike vanilla LLaMA-Factory your YAML must include DataFlex-specific parameters. Configs are therefore not portable in either direction. If your workflow depends on handing the same YAML to a colleague running stock LLaMA-Factory, that will not work without edits.

How DataFlex differs from plain LLaMA-Factory and from offline curation

The honest alternative is LLaMA-Factory without DataFlex. The difference is not the training loop, which is shared, but when the data decision happens. LLaMA-Factory trains on the dataset you give it; the dataset is the input. DataFlex treats the dataset as something the training loop queries, with a selector, a mixer and a weighter deciding what that query returns. The README's own framing, Right in the LLM Training Loop, is the distinction.

The second alternative is offline curation: run a scoring pass, keep the top-k samples, then train on the frozen result. That is cheaper and easier to reason about, and it is what the Static and Random entries in the selection table amount to. The trade-off is that offline scores are computed by a model that is not the model being trained, so the selection cannot adapt as the model changes. DataFlex's model-in-the-loop methods exist precisely to respond to that drift, at the cost of doing extra computation during training and of depending on gradient flow under your chosen parallelism.

A third comparison is between DataFlex's own methods. NEAR and TSDS need no model feedback and are therefore the cheapest to run and the easiest to reproduce; LESS and NICE are gradient-based and the heaviest. DOREMI is offline mixture, ODM is online. Picking within DataFlex is a budget decision as much as a quality one, and the Requires Model-in-the-Loop column in the README is the table to read first.

Licence, maintenance and upgrade cost

DataFlex is Apache-2.0, both in the repository LICENSE and in the pyproject.toml license field, and the PyPI classifiers list it as an OSI approved Apache Software License. Apache-2.0 is permissive and includes an explicit patent grant, but the practical implication here is dependency licensing rather than DataFlex's own terms. The optional less extra pulls in TRAK, and the metrics extra pulls in nltk, jieba and rouge-related packages, each with its own terms. If you ship a model trained with a LESS selector, check those upstream licences rather than assuming DataFlex's licence covers the whole stack. This is a description of the declared metadata, not legal advice.

Maintenance looks current: the last push was on 2026-09-10, the repository is not archived, and the only release listed is v1.0.0 from 2026-04-16. The news entries in the README record DeepSpeed ZeRO-3 gradient support in March 2026 and the first release in December 2025. The upgrade cost is dominated by the pins rather than by API churn: transformers is bounded at both ends, accelerate is bounded below 1.12.0, tyro is pinned to exactly 0.8.14, and deepspeed is capped below 0.17.0. Every one of those ceilings is a place where a future upgrade of an unrelated dependency can collide with DataFlex, and the README's own advice about lifting the DeepSpeed cap on newer torch shows the project expects users to make that call themselves.

Editorial conclusion

Adopt DataFlex if you already train with LLaMA-Factory and want LESS, DOREMI, ODM or the loss-based selectors available behind one YAML config instead of stitching together separate research repositories. Do not adopt it if you are not on LLaMA-Factory, or if you need a stable API: pyproject.toml classifies the project as Beta and the README states your YAML must carry DataFlex-specific parameters that vanilla configs do not have. Before committing, verify that the pinned torch build imports on your GPU node, that the deepspeed extra resolves below 0.17, and that the selector you want is not one of the entries the README marks with a warning or a cross.

Frequently asked questions

What does data software do?

In DataFlex's case the software schedules training data rather than processing it offline: it selects samples, adjusts the ratio between domains, and reweights samples during backpropagation, all inside the LLaMA-Factory training loop.

What are examples of data software?

DataFlex is one: a data-centric LLM training system built on LLaMA-Factory, licensed Apache-2.0 and installed with pip install dataflex. The README groups its methods into Data Selection, Data Mixture and Data Reweighting.

What is DataFlex?

DataFlex is a data-centric LLM training system released by OpenDCAI, licensed Apache-2.0 and installed with pip install dataflex. It requires Python 3.11+ and LlamaFactory 0.9.5+, and its CLI entry point is dataflex-cli.

Official sources

  1. License: Apache-2.0
  2. OpenDCAI/DataFlex on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/opendcai-dataflex.svg)](https://hysenlabs.com/projects/opendcai-dataflex)