Library / SDK
sdv-dev/SDGym avatar
sdv-dev/SDGym

SDGym benchmarks tabular synthesizers by listing them, not the datasets

Benchmarking synthetic data generation methods.

311 stars68 forksPythonNOASSERTION

At a glance

What is it?
SDGym wraps the Synthetic Data Vault ecosystem into a harness that trains and samples a list of synthesizers across bundled datasets and reports performance, memory, quality, and privacy. Its interface is small, but its packaging carries real inconsistencies, from a Docker build that copies two files the root does not have to a Makefile target that repeats a cleanup step.
Who is it for?
SDGym earns a place when you need one repeatable comparison across several tabular synthesizers on the same datasets and you already use SDV, because the harness supplies the datasets, the baseline, and the metrics that a hand-rolled loop would not.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

benchmark_single_table takes a synthesizer list, not a dataset list

The whole entry point is one function and one argument. You pass a sequence of synthesizer names to `sdgym.benchmark_single_table`, and the harness handles datasets, training, sampling, and scoring. The walkthrough builds that sequence by concatenating two groups:

python
import sdgym

sdgym.benchmark_single_table(synthesizers=(sdv_synthesizers + baseline_synthesizers))

`sdv_synthesizers` holds `GaussianCopulaSynthesizer` and `CTGANSynthesizer`, which arrive from the SDV library and use different modeling techniques. `baseline_synthesizers` holds `UniformSynthesizer`, a basic synthesizer shipped inside SDGym as the reference floor. That split is the design in miniature: the point of the benchmark is contrast, so one cheap baseline always travels with the serious models. You never name a dataset in this call, and you never pass metrics or a seed. The output is described as a detailed performance, memory, and quality evaluation across the synthesizers on a variety of publicly available datasets. Privacy scores travel with the same run, since the evaluation bullet names quality and privacy metrics alongside the speed and memory numbers.

A custom synthesizer is two functions with a fixed argument contract

You can put your own model inside the same comparison by defining two functions instead of a class. The first is the training logic, called with `data` and `metadata`, and it returns a synthesizer object. The second is the sampling logic, called with `trained_synthesizer` and `num_rows`, and it returns the synthetic data itself.

python
def my_training_logic(data, metadata):
    # create an object to represent your synthesizer
    # train it using the data
    return synthesizer

def my_sampling_logic(trained_synthesizer, num_rows):
    # use the trained synthesizer to create
    # num_rows of synthetic data
    return synthetic_data

The `metadata` argument is the load-bearing one. A training function that ignores it cannot see column types, keys, or any other structural description the harness derived from the dataset, so it re-derives everything from raw values. `num_rows` decouples the synthetic row count from the training row count, which is what lets one trained model be scored at several sample sizes. Because the contract is two plain callables returning two values, anything trainable can enter the benchmark, including a model you wrote in another library.

DatasetExplorer reports megabytes and table counts before you benchmark anything

Bundled datasets are discoverable through a class in the `dataset_explorer` module. Calling `list_datasets()` on `DatasetExplorer` returns rows with three columns, `dataset_name`, `size_MB`, and `num_tables`.

python
sdgym.dataset_explorer.DatasetExplorer().list_datasets()

The sample output shows `KRK_v1` at 0.072128 MB, `adult` at 3.907448 MB, `alarm` at 4.520128 MB, and `asia` at 1.280128 MB, each with `num_tables` of 1. The sizes span more than two orders of magnitude, which is the number that matters when you plan a run: `KRK_v1` is negligible next to `alarm`, and a benchmark that sweeps every dataset at every synthesizer will be dominated by the largest ones. The uniform `num_tables` of 1 in the visible rows is consistent with the single-table entry point named `benchmark_single_table`. You can also pass your own data, which is the only lever for changing the dataset set, and the docs link a customization page for it.

Private datasets are addressed by S3 path while dependencies reach two clouds

Custom private data is not read from a local path. You point the harness at an Amazon S3 location instead, assigning a bucket URL to a folder variable.

code
my_datasets_folder = 's3://my-datasets-bucket'

That single line is the documented route for data stored on your own computer, which means a private benchmark set has to live in a bucket rather than on a mounted disk, and it means S3 credentials must be present in the environment before the run starts. The dependency list is wider than that. `boto3` and `botocore` are pinned with lower bounds of 1.28 and 1.31 for the AWS side, and `google-cloud-compute` at 1.30.0 with `google-auth` at 2.14.1 are installed alongside them, which means a Google Cloud path exists in the dependency graph even though the visible usage walkthrough shows only an S3 scheme. Anyone deciding whether they can keep a corpus on a private filesystem should read the customization docs first rather than inferring support from the package requirements. A customization page is linked for datasets and another for synthesizers, so both extension points have dedicated documentation rather than living only in the walkthrough.

The Dockerfile copies two files that are not at the repository root

The container build starts from `nvidia/cuda:11.8.0-devel-ubuntu22.04`, whose own command is `nvidia-smi`, then installs `build-essential`, `curl`, `python3.9`, `python3-pip`, and `python3-distutils`, symlinking `/usr/bin/python3.9` to `/usr/bin/python`. The copy step is where the build gets interesting:

dockerfile
COPY pyproject.toml README.md HISTORY.md MANIFEST.in LICENSE Makefile setup.cfg /SDGym/

`pyproject.toml`, `README.md`, `HISTORY.md`, `LICENSE`, and `Makefile` all appear at the top level of the repository. `MANIFEST.in` and `setup.cfg` do not, while the package metadata in `pyproject.toml` is fully declarative with a dynamic version. That makes the image build the first thing to test, because a missing copy source fails before any of the benchmark code is reachable. Two more details travel with it: installation runs `pip install . --no-binary pomegranate`, forcing a source build of that dependency, and the build then runs `make compile` with `TF_CPP_MIN_LOG_LEVEL` set to 2. The image's own command prints a usage line for `docker run -ti sdvproject/sdgym sdgym COMMAND OPTIONS`.

install-develop repeats clean-pyc and drops the compile step

The Makefile sets `help` as its default goal and generates it from itself: the help target pipes the Makefile through a small Python snippet that matches target lines carrying a `## ` comment and prints the target name beside the comment text. Add a trailing comment to a target and it appears in help, remove it and the target disappears from the listing. A second exported snippet resolves a browser command used to open locally built documentation.

The install targets form a small family. `install` depends on `clean-build`, `clean-compile`, `clean-pyc`, and `compile`, then runs `pip install .`. `install-test` runs the same chain and then `pip install .[test]`, which is how the test extra gets pulled in. `install-develop` is the odd one out: it declares `clean-build clean-pyc clean-pyc`, listing `clean-pyc` twice and omitting `compile`. The duplicated prerequisite is harmless because make runs a target once per invocation, but the missing compile step means an editable install skips the build step that the other two targets insist on. The clean family also separates `clean-build`, `clean-pyc`, `clean-coverage`, and `clean-test`, with `clean` aggregating all four.

numpy floors climb with each Python while pandas is capped below 3

The dependency block is versioned per interpreter rather than with one floor, and the pattern is systematic. Five `numpy` markers cover the supported range: 1.22.2 for Python below 3.10, 1.24.0 for 3.10 and 3.11, 1.26.0 for 3.12, 2.1.0 for 3.13, and 2.3.2 for 3.14. Four `pandas` markers cover the same span, starting at 1.4.0 below 3.11 and ending at 2.2.3 for 3.13 and 3.14, with a ceiling of `<3` on every band. `cloudpickle` splits the same way at 2.1.0 and 3.1.1 around Python 3.14. `requires-python` states the envelope as `>=3.9,<3.15`, and the classifiers name 3.9 through 3.14 individually. The practical consequence is that the newest interpreter and the newest numeric stack arrive together, so a 3.14 environment resolves to numpy 2.3.2 while a 3.9 environment is held near numpy 1.22, and a benchmark that compares models across machines is also comparing numeric library versions.

A pre-alpha classifier sits beside v0.15.1 and a license the index did not resolve

Packaging metadata and the public record disagree in three ways worth knowing before you depend on this. First, `pyproject.toml` declares `Development Status :: 2 - Pre-Alpha` while the release history shows v0.15.1 published 2026-10-01, v0.15.0 on 2026-08-31, and v0.14.4 on 2026-06-02, so the version line has moved a long way from the stability claim. Second, the same file sets `license = 'BUSL-1.1'` with `license-files = ['LICENSE']`, while the repository metadata reports the license as NOASSERTION, which is the unresolved marker rather than a statement that no license applies. The LICENSE file is the authority here, and it is present at the root. Third, the two descriptions differ: the repository says it benchmarks synthetic data generation methods, and the package says it benchmarks tabular synthetic data generators using a variety of datasets.

The last recorded push is 2026-09-30, one day before the newest tag. Alongside the standard files sit `INSTALL.md`, `DOCKER.md`, `DATASETS.md`, `HISTORY.md`, `RELEASE.md`, `CONTRIBUTING.rst`, `AUTHORS.rst`, `codecov.yml`, `latest_requirements.txt`, `static_code_analysis.txt`, `tasks.py`, `scripts/`, `tests/`, and `docs/`.

Editorial conclusion

SDGym earns a place when you need one repeatable comparison across several tabular synthesizers on the same datasets and you already use SDV, because the harness supplies the datasets, the baseline, and the metrics that a hand-rolled loop would not. Skip it if your data stays on a private machine with no S3 bucket, if you need a maintained library rather than one still classifying itself as pre-alpha, or if your legal review expects an OSI license and finds BUSL-1.1 in pyproject.toml instead. Before you build on it, try the container build, because it copies MANIFEST.in and setup.cfg that the repository root does not list.

Frequently asked questions

How do I run an SDGym benchmark on a single table?

Call `sdgym.benchmark_single_table` with a list of synthesizer names, such as `GaussianCopulaSynthesizer`, `CTGANSynthesizer`, and the bundled `UniformSynthesizer`. You do not pass datasets, and the harness evaluates performance, memory, and quality across publicly available datasets.

Can SDGym benchmark a synthesizer I wrote myself?

Yes. Supply a training function that takes data and metadata and returns a synthesizer, plus a sampling function that takes the trained synthesizer and num_rows and returns synthetic data. The docs point to a custom synthesizers guide.

Where do I keep custom datasets for SDGym?

On an Amazon S3 bucket, assigned as a folder such as `s3://my-datasets-bucket`, rather than on a local path. The package also installs `google-cloud-compute` and `google-auth` alongside `boto3` and `botocore`.

Which Python versions does SDGym support?

`requires-python` is set to `>=3.9,<3.15` and classifiers name 3.9 through 3.14. numpy, pandas, and cloudpickle each carry separate markers per interpreter band, from numpy 1.22.2 up to 2.3.2.

What license does SDGym use?

`pyproject.toml` declares `license = 'BUSL-1.1'` with `license-files = ['LICENSE']`, while repository metadata reports NOASSERTION. A LICENSE file is present at the repository root.

How do I install SDGym?

Through pip with `pip install sdgym`, or through conda with `conda install -c pytorch -c conda-forge sdgym`. The install section recommends a virtual environment to avoid conflicts with other software.

Official sources

  1. Issues
  2. README
  3. Releases
  4. sdv-dev/SDGym on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/sdv-dev-sdgym.svg)](https://hysenlabs.com/projects/sdv-dev-sdgym)