Model or dataset
NVIDIA-NeMo/Curator avatar
NVIDIA-NeMo/Curator

NVIDIA NeMo Curator: GPU data pipelines for LLM training sets

Scalable data pre processing and curation toolkit for LLMs

1,788 stars340 forksPythonApache-2.0

At a glance

What is it?
NeMo Curator is NVIDIA's Apache-2.0 toolkit for building repeatable, GPU-accelerated pipelines that load, filter, deduplicate and transform text, image, video and audio data. It is built for data teams who have outgrown ad-hoc notebooks, and it expects a Ray cluster and NVIDIA hardware to earn its keep.
Who is it for?
Adopt NeMo Curator if you already run NVIDIA GPUs and need deduplication, quality classification or embedding stages over datasets too large for a single machine, and you want the pipeline to be reproducible rather than a notebook someone edited last quarter. Skip it if your corpus fits in pandas on a laptop, if you have no NVIDIA hardware, or if you only need a one-off regex cleanup.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem NeMo Curator solves, and who it is actually for

Most teams that train or fine-tune a language model eventually hit the same wall. The data cleaning code starts as a notebook, grows a few helper scripts, and then nobody can reproduce the corpus that produced last month's checkpoint. NeMo Curator is aimed squarely at that situation. The README describes it as a way to build repeatable, GPU-accelerated pipelines that load, filter, deduplicate and transform large text, image, video and audio datasets for AI training, and it states that the same pipeline can run on a laptop or across a multi-node Ray cluster.

The intended audience is narrow and specific. The README lists the conditions under which you should reach for it: you need repeatable curation pipelines rather than one-off notebooks or ad-hoc scripts; you need GPU and distributed execution for data-heavy stages such as dedupe, classification, embedding and inference; you need modality-aware building blocks; or you want recipes that map to NVIDIA training workflows like Nemotron and Nemotron-CC. Read that list as a filter, not marketing. If your corpus is a few gigabytes of English text and your cleaning rules are a handful of regular expressions, the machinery here is heavier than the job. The project's own framing is about scale, and the value proposition only holds when the data volume justifies the setup cost.

How the pipeline architecture works: Ray, Xenna and modality stages

The architecture changed direction in 2026. According to the README's updates section, NeMo Curator 26.02 introduced a Ray-based pipeline architecture covering all modalities, and 26.04 followed with a Cosmos-Xenna 0.2.0 upgrade, a simplified Resources API and a Ray runtime upgrade. The consequence for anyone reading older tutorials is that the pipeline abstraction, not just the surface API, moved. Code written against the pre-26.02 layout will not map cleanly onto the current one.

The data flow the README describes is a staged pipeline: load, then filter, deduplicate and transform, with GPU-accelerated stages for the expensive parts. The repository layout reflects that split. The nemo_curator/ package holds the stages, and pyproject.toml declares package data for one of them specifically, nemo_curator.stages.text.experimental.translation, which ships prompt YAML files. That is a useful signal about how the text stages work: some of them are prompt-driven rather than rule-driven.

On the compute side, the README states that NeMo Curator draws on NVIDIA RAPIDS (cuDF, cuML, cuGraph) and that deduplication, classification and embedding stages are the ones meant to run on GPU. The distribution layer is Ray. The README also points to a Slurm deployment guide for multi-node Ray pipelines on HPC clusters, which tells you the target environment is not only cloud Kubernetes but also traditional research clusters. There is also an inference server component described as an OpenAI-compatible LLM endpoint you can spin up inside your pipeline for synthetic data generation, classification and synthetic data workflows. The practical reading: the same job graph can execute locally for development and on a cluster for the real run, which is the property that makes the pipeline reproducible in the first place.

Installing NeMo Curator and running a first text pipeline

The README gives three installation paths and states that the project uses uv for installation. Install uv first if you do not have it:

bash
curl -LsSf https://astral.sh/uv/install.sh | sh

Path A is the CPU smoke test. It needs no GPU, and its purpose is to confirm that the package imports and reports a version. The README gives exactly this sequence:

bash
uv venv && source .venv/bin/activate
uv pip install "nemo-curator[text_cpu]"
python -c "import nemo_curator; print(nemo_curator.__version__)"

If that prints a version string, the environment is sound and you can stop here while you evaluate.

Path B is the GPU text pipeline, and this is where the README is unusually explicit about a real constraint. Prerequisites listed are CUDA 12 toolkit, an NVIDIA driver supporting CUDA 12, Linux x86_64, roughly 16 GB of GPU memory, and network access to Hugging Face. The README states that standard pip install is not supported for the text_cuda12 extra, because vLLM and RAPIDS declare incompatible Numba requirements. The supported route applies a tested override file:

bash
uv venv && source .venv/bin/activate
curl -O https://raw.githubusercontent.com/NVIDIA-NeMo/Curator/main/requirements/text_cuda12-overrides.txt
uv pip install \
  --override text_cuda12-overrides.txt \
  --torch-backend cu129 \
  --extra-index-url https://wheels.vllm.ai/0.22.0/cu129 \
  "nemo-curator[text_cuda12]"
python tutorials/quickstart.py

The README notes that any uv pip install command including text_cuda12, including nemo-curator[all], needs the override file, and that from a source checkout, uv sync --extra text_cuda12 and uv sync --extra all apply the project override automatically. That last detail is the one to remember: working from a checkout removes an entire class of dependency-resolution failure.

The quickstart script itself, per the README, starts Ray, downloads a Hugging Face model and runs a sentiment classification pipeline on GPU. Expect the model download to dominate the first run.

Path C is Docker, which the README recommends for video and audio because those pipelines depend on system codec libraries and the published container ships them preconfigured. The container is published on NGC as nemo-curator, and the README directs readers to the installation guide for setup. If your work is video or audio, start there rather than fighting codec packages on a bare host.

Where NeMo Curator is the wrong tool

The dependency situation is the first real limitation, and it is not incidental. The README states plainly that text_cuda12 cannot be installed with plain pip because two of its dependencies declare incompatible Numba requirements. A project that needs an override file to install has accepted a maintenance burden on behalf of its users, and that burden lands on you every time you upgrade. The README does not describe what happens when the override drifts out of sync with a new vLLM or RAPIDS release, and it does not document a rollback procedure for a pipeline that produces a bad corpus. Both gaps matter in production.

The second limitation is hardware. The GPU path assumes an NVIDIA GPU with roughly 16 GB of memory, a CUDA 12 driver and Linux x86_64. There is a CPU path, and the README presents it as a smoke test for verifying your environment, not as a way to curate a large corpus. Teams on Apple silicon, on AMD accelerators, or on Windows without WSL are outside the documented envelope.

The third is scope fit. The README's own criteria for choosing NeMo Curator are repeatability, GPU and distributed execution, modality-aware building blocks, and alignment with NVIDIA training recipes. A team that wants a quick regex pass over a few hundred megabytes of text satisfies none of those, and would pay the Ray and RAPIDS setup cost for nothing. The honest summary is that this is infrastructure for data volumes where a single machine is the bottleneck. Below that threshold, ordinary Python and pandas are faster to write and easier to debug.

One more thing worth flagging: the README's quickstart for video and audio routes through Docker rather than a pip install, which means those modalities have a heavier floor for experimentation than text does.

How it compares to Hugging Face Datasets and Spark-based cleaning

The obvious alternative for text work is Hugging Face Datasets with a map function. The difference in approach is where the compute happens and what the unit of work is. Datasets gives you an Arrow-backed table and a Python function applied per row or per batch, running on CPU unless you write the GPU code yourself. NeMo Curator gives you prebuilt stages, deduplication and classification among them, that are written to execute on GPU through RAPIDS and to be distributed through Ray. If your cleaning logic is bespoke, Datasets is the more direct tool. If your logic is one of the standard curation operations the README lists, NeMo Curator saves you from writing and tuning that operation yourself.

The second alternative is a Spark or Dask job with hand-written cleaning logic. That approach is portable across hardware and has a much larger hiring pool. What it does not give you is GPU execution for embedding generation or classifier inference, which are the stages where NeMo Curator's README claims its advantage. The trade is explicit: you accept NVIDIA-specific dependencies and the override-file installation in exchange for GPU-accelerated stages and recipes that the README says map to Nemotron and Nemotron-CC pipelines.

A third point of comparison is the pipeline abstraction itself. A notebook plus a shell script can produce the same corpus, and for a one-time job it is less work. NeMo Curator's argument is that the corpus needs to be reproducible and the pipeline needs to scale without a rewrite. That argument holds for teams that retrain regularly. It holds less well for a team that curates once and moves on.

Maintenance, releases and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-09-09. The release cadence visible in the release list is roughly quarterly: v1.1.0 on 2026-02-23, v1.2.0 on 2026-05-14, v1.3.0 on 2026-07-27. The README's updates section also references versioned trains, 26.02 and 26.04, with 26.04 described as bringing a Cosmos-Xenna 0.2.0 upgrade, a simplified Resources API and a Ray runtime upgrade. Treat that as the upgrade cost signal: the pipeline architecture was reworked in 2026, so a codebase written against an earlier layout will need attention, and the release notes page is the place to check before bumping.

Licensing is Apache-2.0, declared in pyproject.toml and in the repository's LICENSE file. That is a permissive licence, and it is the same licence the README badge points at. The practical implication for most teams is that internal use and modification carry few obligations beyond preserving notices. This is not legal advice, and the licence file is the authoritative text.

Two operational details are worth noting because they affect upgrade work. First, the project publishes a Fern-based documentation site with its own Makefile targets, and the Makefile comments state that CI substitutes version variables in the MDX automatically, with DOCS_VERSION defaulting to 26.04. That means the docs you read are version-scoped, and a page describing one train may not describe another. Second, the README directs bug reports and feature requests to GitHub issues, and the repository carries CONTRIBUTING.md, CODE_OF_CONDUCT.md and SECURITY.md, so there is a defined path for reporting problems rather than a dead-end inbox.

Editorial conclusion

Adopt NeMo Curator if you already run NVIDIA GPUs and need deduplication, quality classification or embedding stages over datasets too large for a single machine, and you want the pipeline to be reproducible rather than a notebook someone edited last quarter. Skip it if your corpus fits in pandas on a laptop, if you have no NVIDIA hardware, or if you only need a one-off regex cleanup. Before committing, verify three things: that your Python version falls inside the >=3.11,<3.14 range declared in pyproject.toml, that your CUDA 12 install succeeds through the uv override path rather than plain pip, and that you can actually run the bundled tutorials/quickstart.py end to end on your GPU.

Frequently asked questions

How do I install NVIDIA NeMo Curator for a GPU text pipeline?

The README states that NeMo Curator uses uv for installation, and that the text_cuda12 extra cannot be installed with plain pip because vLLM and RAPIDS declare incompatible Numba requirements. The supported command downloads requirements/text_cuda12-overrides.txt and passes it with --override to uv pip install, alongside --torch-backend cu129 and the vLLM wheel index. From a source checkout, uv sync --extra text_cuda12 applies the project override automatically.

What Python versions does NeMo Curator support?

The pyproject.toml declares requires-python as >=3.11,<3.14, so Python 3.11 through 3.13 are inside the supported range. The package classifiers list Python 3.11 explicitly.

Does NeMo Curator require an NVIDIA GPU?

Not for the CPU smoke test. The README offers a text_cpu extra that installs without a GPU and is presented as a way to verify your environment and run a tiny text pipeline. The GPU path lists CUDA 12 toolkit, a CUDA 12 driver, Linux x86_64 and roughly 16 GB of GPU memory as prerequisites.

Which modalities can NeMo Curator process?

The README lists text, image, video and audio, each with its own guide and common operations. Text covers deduplication, classification, quality filtering and language detection; image covers aesthetic filtering, NSFW detection, embedding generation and deduplication; video covers scene detection, clip extraction, motion filtering and deduplication; audio covers ASR transcription, quality assessment and WER filtering.

Why does the README recommend Docker for video and audio pipelines?

The README states that video and audio pipelines depend on system codec libraries, and that the published container ships them preconfigured. The container is published on NGC as nemo-curator, and the README points to the installation guide for setup instructions.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. NVIDIA-NeMo/Curator on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nvidia-nemo-curator.svg)](https://hysenlabs.com/projects/nvidia-nemo-curator)