# NeMo Data Designer: declarative synthetic datasets for LLM and multimodal pipelines

> NeMo Data Designer builds synthetic datasets as a column graph: samplers, LLM columns and validators, previewed before a full run. It fits when you need controlled distributions and correlated fields, and is the wrong tool when a single prompt call would do.

**NVIDIA-NeMo/DataDesigner** — 🎨 NeMo Data Designer: Generate high-quality synthetic data from scratch or from seed data.

- Repository: https://github.com/NVIDIA-NeMo/DataDesigner
- Website: https://docs.nvidia.com/nemo/datadesigner
- Stars: 2,285 · Forks: 211
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/nvidia-nemo-datadesigner

## The problem NeMo Data Designer solves, and who it is for

Prompting a model for a few hundred JSON records is easy. Producing a dataset where a category column is drawn from a fixed distribution, a free-text column is conditioned on that category, and every row passes a check before it is written out is a different job. That second job is what NeMo Data Designer targets.

The README frames the gap directly: the library helps you create datasets "that go beyond simple LLM prompting", naming diverse statistical distributions, correlations between fields and validated outputs. The intended user is an engineer building training or evaluation data for language and multimodal models, not an analyst assembling a spreadsheet. The topics attached to the repository (synthetic-data, data-augmentation, multimodal, tool-use) describe the same audience.

It is a Python library rather than an application. There is no web UI described in the README; the surface is a config builder plus a DataDesigner object, and a CLI for provider and model configuration. If your team does not write Python, this is not the tool for you.

## How the column graph works: samplers, LLM columns and validators

The mechanism is a config builder that accumulates column definitions, then hands the whole config to a generator. The README's quick start shows two column types in sequence: a SamplerColumnConfig that draws from a categorical list, and an LLMTextColumnConfig whose prompt template references the first column by name.

That reference is the dependency mechanism. The prompt string contains a placeholder such as a product category, and the engine resolves it per row before calling the model. The README calls this dependency-aware generation and lists it as a way to control relationships between fields. Practically, it means field order in the builder is not cosmetic: a column that references another column depends on it.

The README also lists validators in several forms (Python, SQL, custom validators, LLM judges) and separate concepts for column types, model configuration and person sampling in the docs. The repository layout supports the layered description: packages/data-designer-config, packages/data-designer-engine and packages/data-designer are separate workspace members, so configuration, execution and the user-facing interface are distinct packages under one data_designer namespace.

What the README does not spell out is the scheduling behaviour of that graph: whether independent columns are generated in parallel, how retries interact with validators, or what happens to a row that fails validation. Those answers live in the docs pages it links to, not in the README text.

## Install NeMo Data Designer and generate a first preview

The README gives two install paths. The published package installs with pip, and a source install clones the repository and runs make install. The source path assumes the uv workspace described in pyproject.toml, which requires uv 0.7.10 or newer and installs the config, engine and interface packages in editable mode.

```bash
pip install data-designer
```

After installing, set a key for at least one supported provider. The README names NVIDIA Build API, OpenAI and OpenRouter, and shows the corresponding environment variables.

```bash
export NVIDIA_API_KEY="your-api-key-here"
```

With a key in place, the config builder is where the dataset is defined. This is the README's own example, trimmed to the two columns it adds: a category sampler and a text column whose prompt references the sampled category.

```python
import data_designer.config as dd
from data_designer.interface import DataDesigner

data_designer = DataDesigner()
config_builder = dd.DataDesignerConfigBuilder()

config_builder.add_column(
    dd.SamplerColumnConfig(
        name="product_category",
        sampler_type=dd.SamplerType.CATEGORY,
        params=dd.CategorySamplerParams(
            values=["Electronics", "Clothing", "Home & Kitchen", "Books"],
        ),
    )
)
```

The second column uses model_alias="nvidia-text" and a prompt containing a product category placeholder. That alias is what ties a column to a configured model, so a preview will fail if the alias is not registered for your provider.

```python
config_builder.add_column(
    dd.LLMTextColumnConfig(
        name="review",
        model_alias="nvidia-text",
        prompt="Write a brief product review for a {{ product_category }} item you recently purchased.",
    )
)

preview = data_designer.preview(config_builder=config_builder)
preview.display_sample_record()
```

The expected result is a small sample record printed in your notebook or terminal, with the sampled category and the generated review side by side. Preview is the cheap step; the README positions it as the way to check a schema before committing to a large run. If the alias or key is wrong, this is where you find out, not after a long generation job.

## Where NeMo Data Designer is the wrong choice

The strongest limitation is structural: the library is a generation framework, and generation means model calls. The quick start requires an API key from a hosted provider before anything runs. If your environment has no outbound access to those providers, or if the data cannot leave your network, the default path in the README does not apply and you are in the custom model configuration docs instead.

Cost scales with columns, not rows. Each LLM column is a per-row model call, so a design with three text columns over a large row count is three times the inference spend of one, and preview tells you nothing about the total. The README describes monitoring large runs but does not give a cost model.

There is also a governance detail worth reading before adoption. The README states that Data Designer collects telemetry to see which models are most popular for synthetic data generation, and that the aggregate is shared with the community. It documents NEMO_TELEMETRY_ENABLED=false as the switch. For a library that may run inside a regulated pipeline, that default deserves a decision rather than an assumption.

Finally, the repository is candid about being a moving target. The workspace pins nbconvert to a specific git revision to work around an end-of-life dependency upstream, and the docs note that Fern prose under fern/ is the source of truth while generated notebooks are not. Neither is a defect, but both mean you should pin the version you validate against.

## NeMo Data Designer compared with NeMo Curator and plain prompt scripts

The related searches pair this project with NeMo Curator, and the two solve different halves of the data problem. Curator is a curation pipeline: it filters, deduplicates and reshapes data you already have. Data Designer is a generation pipeline: it produces records that did not exist, from samplers, models or seed data. If your corpus is large and messy, Curator is the relevant tool. If your corpus is too small or too uniform, Data Designer is.

The other alternative is the one most teams already have: a Python script that loops over prompts and writes JSONL. The difference is not the model call, it is everything around it. A prompt script has no column abstraction, no dependency resolution between fields, no validator stage, and no preview step that renders a sample record from the config. It also has no resume story, while the README lists preview, resume and monitoring as generation features.

The honest trade-off: a prompt script is a few dozen lines and has no provider configuration layer, no telemetry, and no workspace of three packages. Data Designer earns its complexity when you need distributions and validation. Below that threshold, the script wins.

## Maintenance, licensing and what a dependency policy file implies

The last push to the default branch was on 2026-09-09, and the most recent release listed is v0.9.2 on 2026-09-03, following v0.9.1 and v0.9.0 in August 2026. The repository is not archived. That is a release cadence measured in weeks, and the 0.9.x version numbers are a reminder that the API is still pre-1.0.

The project is Apache-2.0, and the README repeats that the license text is in LICENSE. Two things follow. First, the usual patent grant and attribution terms of Apache-2.0 apply, and you should read the file rather than this paragraph. Second, the install note in the README warns that the project downloads and installs additional third-party open source software and tells you to review those licenses before use. The presence of dependency-license-policy.toml at the repository root suggests the maintainers track that themselves, but it governs their build, not your deployment. If your organization runs a license scan, expect the transitive set, not just the three workspace packages.

Upgrade cost is bounded by the pre-1.0 status. The workspace pins a git revision of nbconvert, so a source install can move under you when that pin changes. Pin the released data-designer version in your own lockfile and read VERSIONING.md before a jump.

## Conclusion

Adopt NeMo Data Designer if your synthetic data needs controlled distributions, correlated fields or validator-gated output, and you accept that every generation path depends on a configured model provider. Skip it if a single templated prompt call already produces the records you need, because the config builder and provider setup cost more than they return at that scale. Before committing, verify three things on the version you install: which sampler types and validator kinds the docs list for your release, that your chosen provider key works through the data-designer config models flow, and whether the telemetry described in the README is acceptable for your environment or must be switched off.

## FAQ

### What is NeMo Data Designer?

It is an NVIDIA Python library for generating synthetic datasets from scratch or from seed data. The README describes it as going beyond simple LLM prompting, with samplers, LLM columns, validators and dependency-aware generation between fields.

### What does a data designer do in this project?

In NeMo Data Designer the work is defining a column graph in code: a config builder accumulates sampler, LLM and validator columns, and the DataDesigner object previews or runs that config. The role is closer to pipeline authoring than to visual design.

### How is NeMo Data Designer different from a data engineer's work?

The library is a generation framework, not a curation or warehouse tool. It produces records from samplers, models or seed data, while the README's related concepts for validation and model configuration stay inside the same config builder.

### Which AI data generator is the best?

The README does not rank generators against each other. It positions NeMo Data Designer specifically around statistical distributions, correlations between fields and validated outputs, and lists NVIDIA Build API, OpenAI and OpenRouter as the default providers.

## Sources

- [License: Apache-2.0](https://github.com/NVIDIA-NeMo/DataDesigner/blob/main/LICENSE)
- [NVIDIA-NeMo/DataDesigner on GitHub](https://github.com/NVIDIA-NeMo/DataDesigner)
- [Project website](https://docs.nvidia.com/nemo/datadesigner)
- [README](https://github.com/NVIDIA-NeMo/DataDesigner/blob/main/README.md)
- [Releases](https://github.com/NVIDIA-NeMo/DataDesigner/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nvidia-nemo-datadesigner
