NeMo Data Designer: declarative synthetic datasets for LLM and multimodal pipelines
🎨 NeMo Data Designer: Generate high-quality synthetic data from scratch or from seed data.
At a glance
- What is it?
- NeMo Data Designer builds synthetic datasets as a column graph: samplers, LLM columns and validators, previewed before a full run. It fits when you need controlled distributions and correlated fields, and is the wrong tool when a single prompt call would do.
- Who is it for?
- Adopt NeMo Data Designer if your synthetic data needs controlled distributions, correlated fields or validator-gated output, and you accept that every generation path depends on a configured model provider. Skip it if a single templated prompt call already produces the records you need, because the config builder and provider setup cost more than they return at that scale.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem NeMo Data Designer solves, and who it is for
Prompting a model for a few hundred JSON records is easy. Producing a dataset where a category column is drawn from a fixed distribution, a free-text column is conditioned on that category, and every row passes a check before it is written out is a different job. That second job is what NeMo Data Designer targets.
The README frames the gap directly: the library helps you create datasets "that go beyond simple LLM prompting", naming diverse statistical distributions, correlations between fields and validated outputs. The intended user is an engineer building training or evaluation data for language and multimodal models, not an analyst assembling a spreadsheet. The topics attached to the repository (synthetic-data, data-augmentation, multimodal, tool-use) describe the same audience.
It is a Python library rather than an application. There is no web UI described in the README; the surface is a config builder plus a DataDesigner object, and a CLI for provider and model configuration. If your team does not write Python, this is not the tool for you.
How the column graph works: samplers, LLM columns and validators
The mechanism is a config builder that accumulates column definitions, then hands the whole config to a generator. The README's quick start shows two column types in sequence: a SamplerColumnConfig that draws from a categorical list, and an LLMTextColumnConfig whose prompt template references the first column by name.
That reference is the dependency mechanism. The prompt string contains a placeholder such as a product category, and the engine resolves it per row before calling the model. The README calls this dependency-aware generation and lists it as a way to control relationships between fields. Practically, it means field order in the builder is not cosmetic: a column that references another column depends on it.
The README also lists validators in several forms (Python, SQL, custom validators, LLM judges) and separate concepts for column types, model configuration and person sampling in the docs. The repository layout supports the layered description: packages/data-designer-config, packages/data-designer-engine and packages/data-designer are separate workspace members, so configuration, execution and the user-facing interface are distinct packages under one data_designer namespace.
What the README does not spell out is the scheduling behaviour of that graph: whether independent columns are generated in parallel, how retries interact with validators, or what happens to a row that fails validation. Those answers live in the docs pages it links to, not in the README text.
Install NeMo Data Designer and generate a first preview
The README gives two install paths. The published package installs with pip, and a source install clones the repository and runs make install. The source path assumes the uv workspace described in pyproject.toml, which requires uv 0.7.10 or newer and installs the config, engine and interface packages in editable mode.
pip install data-designerAfter installing, set a key for at least one supported provider. The README names NVIDIA Build API, OpenAI and OpenRouter, and shows the corresponding environment variables.
export NVIDIA_API_KEY="your-api-key-here"With a key in place, the config builder is where the dataset is defined. This is the README's own example, trimmed to the two columns it adds: a category sampler and a text column whose prompt references the sampled category.
import data_designer.config as dd
from data_designer.interface import DataDesigner
data_designer = DataDesigner()
config_builder = dd.DataDesignerConfigBuilder()
config_builder.add_column(
dd.SamplerColumnConfig(
name="product_category",
sampler_type=dd.SamplerType.CATEGORY,
params=dd.CategorySamplerParams(
values=["Electronics", "Clothing", "Home & Kitchen", "Books"],
),
)
)The second column uses model_alias="nvidia-text" and a prompt containing a product category placeholder. That alias is what ties a column to a configured model, so a preview will fail if the alias is not registered for your provider.
config_builder.add_column(
dd.LLMTextColumnConfig(
name="review",
model_alias="nvidia-text",
prompt="Write a brief product review for a {{ product_category }} item you recently purchased.",
)
)
preview = data_designer.preview(config_builder=config_builder)
preview.display_sample_record()The expected result is a small sample record printed in your notebook or terminal, with the sampled category and the generated review side by side. Preview is the cheap step; the README positions it as the way to check a schema before committing to a large run. If the alias or key is wrong, this is where you find out, not after a long generation job.
Where NeMo Data Designer is the wrong choice
The strongest limitation is structural: the library is a generation framework, and generation means model calls. The quick start requires an API key from a hosted provider before anything runs. If your environment has no outbound access to those providers, or if the data cannot leave your network, the default path in the README does not apply and you are in the custom model configuration docs instead.
Cost scales with columns, not rows. Each LLM column is a per-row model call, so a design with three text columns over a large row count is three times the inference spend of one, and preview tells you nothing about the total. The README describes monitoring large runs but does not give a cost model.
There is also a governance detail worth reading before adoption. The README states that Data Designer collects telemetry to see which models are most popular for synthetic data generation, and that the aggregate is shared with the community. It documents NEMO_TELEMETRY_ENABLED=false as the switch. For a library that may run inside a regulated pipeline, that default deserves a decision rather than an assumption.
Finally, the repository is candid about being a moving target. The workspace pins nbconvert to a specific git revision to work around an end-of-life dependency upstream, and the docs note that Fern prose under fern/ is the source of truth while generated notebooks are not. Neither is a defect, but both mean you should pin the version you validate against.
NeMo Data Designer compared with NeMo Curator and plain prompt scripts
The related searches pair this project with NeMo Curator, and the two solve different halves of the data problem. Curator is a curation pipeline: it filters, deduplicates and reshapes data you already have. Data Designer is a generation pipeline: it produces records that did not exist, from samplers, models or seed data. If your corpus is large and messy, Curator is the relevant tool. If your corpus is too small or too uniform, Data Designer is.
The other alternative is the one most teams already have: a Python script that loops over prompts and writes JSONL. The difference is not the model call, it is everything around it. A prompt script has no column abstraction, no dependency resolution between fields, no validator stage, and no preview step that renders a sample record from the config. It also has no resume story, while the README lists preview, resume and monitoring as generation features.
The honest trade-off: a prompt script is a few dozen lines and has no provider configuration layer, no telemetry, and no workspace of three packages. Data Designer earns its complexity when you need distributions and validation. Below that threshold, the script wins.
Maintenance, licensing and what a dependency policy file implies
The last push to the default branch was on 2026-09-09, and the most recent release listed is v0.9.2 on 2026-09-03, following v0.9.1 and v0.9.0 in August 2026. The repository is not archived. That is a release cadence measured in weeks, and the 0.9.x version numbers are a reminder that the API is still pre-1.0.
The project is Apache-2.0, and the README repeats that the license text is in LICENSE. Two things follow. First, the usual patent grant and attribution terms of Apache-2.0 apply, and you should read the file rather than this paragraph. Second, the install note in the README warns that the project downloads and installs additional third-party open source software and tells you to review those licenses before use. The presence of dependency-license-policy.toml at the repository root suggests the maintainers track that themselves, but it governs their build, not your deployment. If your organization runs a license scan, expect the transitive set, not just the three workspace packages.
Upgrade cost is bounded by the pre-1.0 status. The workspace pins a git revision of nbconvert, so a source install can move under you when that pin changes. Pin the released data-designer version in your own lockfile and read VERSIONING.md before a jump.
Editorial conclusion
Adopt NeMo Data Designer if your synthetic data needs controlled distributions, correlated fields or validator-gated output, and you accept that every generation path depends on a configured model provider. Skip it if a single templated prompt call already produces the records you need, because the config builder and provider setup cost more than they return at that scale. Before committing, verify three things on the version you install: which sampler types and validator kinds the docs list for your release, that your chosen provider key works through the data-designer config models flow, and whether the telemetry described in the README is acceptable for your environment or must be switched off.
Frequently asked questions
What is NeMo Data Designer?
It is an NVIDIA Python library for generating synthetic datasets from scratch or from seed data. The README describes it as going beyond simple LLM prompting, with samplers, LLM columns, validators and dependency-aware generation between fields.
What does a data designer do in this project?
In NeMo Data Designer the work is defining a column graph in code: a config builder accumulates sampler, LLM and validator columns, and the DataDesigner object previews or runs that config. The role is closer to pipeline authoring than to visual design.
How is NeMo Data Designer different from a data engineer's work?
The library is a generation framework, not a curation or warehouse tool. It produces records from samplers, models or seed data, while the README's related concepts for validation and model configuration stay inside the same config builder.
Which AI data generator is the best?
The README does not rank generators against each other. It positions NeMo Data Designer specifically around statistical distributions, correlations between fields and validated outputs, and lists NVIDIA Build API, OpenAI and OpenRouter as the default providers.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvidia-nemo-datadesigner)