# SDG (sdgx): a Python framework for structured tabular synthetic data

> SDG, published on PyPI as sdgx, generates tabular synthetic data with GAN, statistical and LLM models. The README claims a CTGAN that handles billions of rows; the docs are thinner on evaluation and rollback.

**hitsz-ids/synthetic-data-generator** — SDG is a specialized framework designed to generate high-quality structured tabular data.

- Repository: https://github.com/hitsz-ids/synthetic-data-generator
- Stars: 2,439 · Forks: 390
- Language: Python
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/hitsz-ids-synthetic-data-generator

## What sdgx solves, and who it is actually for

Most teams that need synthetic tabular data start with a script that samples columns independently. That produces a CSV that looks plausible column by column and useless row by row, because the joint distribution is gone. SDG exists to keep the joint structure: it takes a pandas DataFrame, infers a metadata description of the columns, and hands that description to a model that learns relationships between columns before generating rows.

The README frames the payoff in regulatory terms, stating that synthetic data "does not contain any sensitive information" and is therefore exempt from regulations such as GDPR and ADPPA. That is the project's framing, not a legal finding, and it is the single claim in the README most worth treating with suspicion. A generator that memorises rare rows can reproduce real individuals; the README does not document a privacy metric or a differential privacy mode. Treat sdgx as a way to produce shareable test data and training data, not as a compliance control.

The intended user is a Python data engineer or ML engineer who already works in pandas and wants the model choice to be a parameter rather than a rewrite. The repository ships notebooks under example/ for CTGAN, metadata, and two LLM flows, which is the fastest way to see the intended shape of a workflow.

## The metadata layer is the real architecture

Strip away the model names and SDG is a metadata pipeline. A Data Processor module, merged in May 2024 according to the README news entries, sits between the raw DataFrame and the model. It converts column formats before training (the README names Datetime columns specifically, so they are not treated as discrete categories), and converts generated values back into the original format afterwards. It also handles null values and is described as supporting a plug-in system, which is why pluggy appears in the dependency list.

Metadata inference lives in sdgx.data_models.metadata. A February 2024 entry says it supports single-table and multi-table metadata, multiple data types, and automatic type inference. That inference step is where most practical failures will surface: a numeric column with sentinel values, a high-cardinality ID column, or a free-text field will each be typed by the inference rules, and the model that gets selected follows from that typing. The README does not document the inference rules themselves, so the honest position is that you will learn the mapping by reading the source or by inspecting the metadata object in a notebook.

The second architectural thread is the column-relationship work from November 2024: automatic detection of relationships between columns, plus the ability to specify them explicitly. This matters because independent per-column generation is exactly the failure mode described above, and explicit relationship specification is the escape hatch when detection gets it wrong.

Models are pluggable and split by family: CTGAN and GaussianCopula under the statistical and GAN side, and sdgx.models.LLM.single_table.gpt.SingleTableGPTModel on the LLM side. The LLM model is the one that breaks the usual assumption, because the README states it can generate synthetic data with no training data at all, working from metadata alone. That is a genuinely different capability from a GAN, which needs rows to learn from.

## Installing sdgx and generating a first table

The package name on PyPI is sdgx, not synthetic-data-generator. The console entry point defined in pyproject.toml is also sdgx, mapping to sdgx.cli.main:cli. Python 3.9 or newer is required, and the dependency list pins numpy<2 alongside scikit-learn>=0.24,<2 and torch>=2. That numpy pin is the first thing to check, because a project already on numpy 2.x will conflict.

```bash
pip install sdgx
```

After installation, `sdgx --help` should list the CLI commands exposed by the click application. The README does not enumerate the subcommands, so treat the CLI as a convenience wrapper and the Python API as the documented path. The repository's own starting points are the notebooks in example/, including example/sdgx_example_ctgan.ipynb and example/sdgx_example_metadata.ipynb.

The metadata model can be exercised directly, which is worth doing before any training run because it tells you how SDG has typed your columns:

```python
from sdgx.data_models.metadata import Metadata
```

The exact constructor signature is not given in the README, so confirm it against the API docs at synthetic-data-generator.readthedocs.io before relying on it. For LLM-backed generation, openai>=1.10.0 and python-dotenv are already dependencies, which is a strong hint that credentials are read from a .env file, though the README does not spell out the variable name. The two LLM notebooks in example/ are the place to look for the working configuration.

## The billion-row CTGAN claim needs a closer reading

The headline from the December 2023 v0.1.0 release is a CTGAN model supporting billions of data records, backed by a benchmark against SDV hosted under benchmarks/ in the repository. The README summarises the result as lower memory consumption and avoiding a crash during training. Read that carefully: the stated advantage is that SDG did not fall over where the comparison did. That is a robustness claim about memory behaviour, not a claim about the statistical fidelity of the generated rows.

Fidelity is a separate question and the README does not answer it. table-evaluator is in the dependency list, which suggests evaluation is possible, but no metric, threshold or recommended workflow is documented. If your acceptance criterion is that a model trained on synthetic data performs within some margin of one trained on real data, you will have to build that comparison yourself.

The November 2024 note about GaussianCopula is more concrete and more useful: memory usage for discrete data was reduced enough to train on thousands of categorical entries in a 2C4G environment. Two cores and four gigabytes is a modest box, and that is a meaningful statement about where the copula path can run. It is also the strongest argument for choosing GaussianCopula over CTGAN on small categorical tables, where a GAN's overhead buys little.

## Where sdgx is the wrong tool

The scope is single-table. The README describes metadata support for multiple tables, and the roadmap file exists, but the model classes named in the README are single-table models, with the LLM one explicitly namespaced under single_table. If your data is a star schema and you need foreign keys to line up across generated tables, SDG does not document that guarantee, and generating each table independently will produce orphans.

Time series is the second gap. The Data Processor converts Datetime columns so they are not treated as categorical, which preserves the format, but nothing in the README describes temporal ordering, lag structure or sequential dependencies between rows. A model that treats rows as exchangeable will produce timestamps that are individually plausible and collectively wrong.

The third limit is maturity. The most recent release listed is 0.2.4 from 2024-12-03, and the version scheme is still pre-1.0. The last push to the repository was on 2026-08-31, so work is ongoing, but a moving pre-1.0 API means pinning a version and reading the changelog before upgrading. If you need a frozen interface with a long support window, this is not that.

Finally, the LLM path inherits every property of the model behind it: cost per generation, non-determinism, and the fact that your metadata description is sent to a third-party API. For a schema with sensitive column names, that is a data-egress decision, not just a modelling one.

## How sdgx differs from SDV and from an LLM-only generator

SDV is the obvious comparison, and the repository makes it directly by benchmarking against it. The architectural difference is where the flexibility sits. SDV centres on a metadata object that drives a family of synthesizers through a constrained, opinionated API, and it has long covered multi-table and sequential data as first-class cases. SDG instead treats the model as a plug-in behind its own Data Processor and metadata layer, which is why GaussianCopula could be added in a single pull request and why an LLM model sits alongside a GAN under the same metadata description.

That plug-in design is the bet: if you expect to swap generators as your data changes, or you want to compare a copula, a GAN and an LLM on the same table with the same preprocessing, SDG's structure is the reason to pick it. If you want multi-table synthesis with referential integrity and a stable API today, SDV covers ground that SDG's README does not claim.

The LLM-only route is the other alternative, and it is not really a competitor so much as a different cost model. Calling a model per row gives you no training step and, per the README, no training data requirement at all, but you pay per generation, you get no distributional guarantee, and you send your schema out. SDG's value here is that it puts that option behind the same interface as the trained models, so you can start with metadata-only generation to unblock a pipeline and move to CTGAN once you have real rows.

## Maintenance, licensing and what an upgrade costs

The licence is Apache-2.0, declared in pyproject.toml as "Apache Software License 2.0" and shipped as a LICENSE file with a separate NOTICE file. Apache-2.0 is permissive and includes an explicit patent grant, which is usually the reason teams pick it over MIT for infrastructure code. The NOTICE file means attribution obligations travel with redistribution, so if you vendor sdgx into a product, keep that file. This is a description of the licence text, not legal advice.

The upgrade cost is dominated by two things. First, the dependency pins: numpy<2 and scikit-learn>=0.24,<2 constrain the rest of your environment, and a future release that lifts the numpy pin is the upgrade that will actually hurt. Second, the pre-1.0 version number: 0.2.2, 0.2.3 and 0.2.4 all landed within a month of each other in late 2024, which is a fast patch cadence on a young API. Pin the version in your requirements file and read the release notes before moving.

On maintenance, the repository is not archived and the last push was on 2026-08-31. The release history shown stops at 2024-12-03, so the gap between commits and tagged releases is worth checking in the releases page before you assume a fix has shipped. The ROADMAP.md file at the repository root is the place to look for what the maintainers consider unfinished.

## Conclusion

Adopt sdgx if you have a pandas-shaped table and want to compare a GAN, a copula and an LLM generator behind one metadata layer without leaving Python. Do not adopt it if you need multi-table referential integrity out of the box, a stable 1.0 API, or an evaluation report you can hand to a privacy reviewer: the README points at a benchmark page and table-evaluator but does not document a privacy guarantee. Before committing, check the pyproject pin numpy<2 against your existing environment, confirm which model class your data type maps to under sdgx.data_models.metadata, and read the ROADMAP to see whether the multi-table work you need is scheduled or still a proposal.

## FAQ

### What does synthetic data generation mean?

It means producing data that has the statistical characteristics of a real dataset without containing the original records. SDG's README describes the output as retaining the essential characteristics of the original data while carrying no sensitive information.

### Which AI is best for synthetic data generation?

SDG does not rank its models. It ships a GAN (CTGAN), a statistical model (GaussianCopula) and an LLM model under sdgx.models.LLM.single_table.gpt, and the README notes the LLM one can generate from metadata with no training data, which the others cannot.

### Can you give me an example of synthetic data?

The repository provides worked examples rather than a sample file: the notebooks in example/ cover CTGAN, metadata, and two LLM flows, and the README links Colab notebooks for LLM data synthesis, off-table inference and the billion-level CTGAN.

### How to generate a synthetic dataset with SDG?

Install the sdgx package, build a metadata description from your pandas DataFrame with sdgx.data_models.metadata, then pass that metadata to one of the model classes such as CTGAN or SingleTableGPTModel. The example/ notebooks show the intended sequence.

### What is synthetic data generator?

In this context it is the project's own name: SDG, the Synthetic Data Generator from hitsz-ids, a Python framework for generating structured tabular data. The installable package is called sdgx.

## Sources

- [hitsz-ids/synthetic-data-generator on GitHub](https://github.com/hitsz-ids/synthetic-data-generator)
- [Issues](https://github.com/hitsz-ids/synthetic-data-generator/issues)
- [License: Apache-2.0](https://github.com/hitsz-ids/synthetic-data-generator/blob/main/LICENSE)
- [README](https://github.com/hitsz-ids/synthetic-data-generator/blob/main/README.md)
- [Releases](https://github.com/hitsz-ids/synthetic-data-generator/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/hitsz-ids-synthetic-data-generator
