Library / SDK
sdv-dev/SDV avatar
sdv-dev/SDV

SDV: five numpy floors, and a source-available licence

Synthetic data generation for tabular data

3,568 stars422 forksPythonNOASSERTION

At a glance

What is it?
SDV is a Python library for generating tabular synthetic data, built by a commercial company and declared under a source-available licence rather than an open source one. What the manifest shows in detail is the version matrix: numpy and pandas carry five different floors each across the supported interpreters, and one of the modelling dependencies stops being installed at Python 3.14.
Who is it for?
SDV fits a data team that needs synthetic tables which keep the shape of real data without carrying real values, and that wants a single library rather than assembling one per modality. Two things to settle before you adopt it.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

It is a commercial project, and the licence is not open source

The first thing to establish is who stands behind this and under what terms.

The header says the repository is part of the Synthetic Data Vault project, a project from a company named in the same sentence. The forum, the blog and the project website all sit under that company's domain rather than under a neutral or academic one. The manifest lists the same company as the author, and the documentation links include a shortened address that points at the company's documentation site.

The licence is the part with teeth. The manifest declares a business source licence by its identifier and names the licence file at the root of the repository. The readme says the library is publicly available under the Business Source License. The repository's own licence metadata field, meanwhile, records nothing at all.

So the two documents that exist agree with each other and say source-available, not open source. That is not a criticism of the project, which is unusually explicit about it in the installation section rather than burying it. It is simply the first question to answer, because a source-available licence and an open source licence give you different rights depending on what you intend to do with the library and with what it generates.

The citation points the same way. The paper the project asks you to cite was written by three authors at a university data lab, and it predates the commercial entity.

Three data shapes, and models from a copula to a GAN

The feature list is short and the scope is clear.

The library generates synthetic data using machine learning, with models ranging from classical statistical methods to deep learning methods. The two named examples are a copula-based synthesizer and CTGAN, which is the deep learning one, and the range between them is the design: you can pick a fast classical model or a slower neural one for the same task.

Then the three data shapes. You can generate single tables, multiple connected tables, or sequential tables. Those are three different problems rather than three conveniences: a single table is one relation, connected tables are several with keys between them, and sequential data is a series.

The workflow is described as customizable end to end, which in practice means three separate concerns sit in the same library: preprocessing, anonymization, and constraints. The constraint mechanism is the one worth noting. Business rules are expressed as logical constraints rather than as post-hoc filtering, so a rule such as a relationship between two columns becomes part of what the synthesizer has to satisfy.

Anonymization is offered as a choice of types rather than a single switch, which suggests the library treats it as a tunable rather than a guarantee you switch on and forget.

Metadata is an argument to the synthesizer, not an afterthought

The getting started sequence teaches an API shape that is easy to miss.

The first call downloads a demo dataset, a single table describing guests at a fictional hotel, and it returns two things: the data and a metadata object. That metadata is a description of the dataset, and the readme says it includes the data types of each column and the primary key, which in the demo is a guest email column.

Then the synthesizer is constructed from that metadata alone, and only afterwards is it fitted with the data. The evaluation call takes metadata as a third argument as well.

So metadata is a first-class object that flows through construction, fitting and evaluation. In a pipeline that means you build and version the metadata once per dataset, and the same object describes both the training run and the quality report. That is a reasonable design for reproducible datasets, and it also means metadata handling is something you will be writing, whether you like it or not.

The three-step shape of the demo, download, construct, fit, sample, is short enough that it is worth copying verbatim as a smoke test before you build your own pipeline around it.

The quality score is the average of two families of measure

The evaluation section includes the actual console output from a run, and the output is more informative than the prose around it.

A quality report is generated, and it contains two families of measure. The first is column shapes, which in the example ran over nine columns and scored 89.11 percent. The second is column pair trends, which ran over thirty-six pairs and scored 88.3 percent. The overall score is the average of the two, shown as 88.7 percent, on a scale of zero to one hundred with one hundred being the best.

So the headline number is a mean of a per-column measure and a per-pair measure. That structure tells you what it can and cannot see. It will catch a column whose distribution has drifted and a pair whose correlation has drifted, which is most of what matters for tabular fidelity. It says nothing about whether the synthetic rows are individually plausible, whether rare categories survive, or whether the data is usable for the specific question you have.

There is also a per-column plot helper for looking at one column side by side, and the documented example uses an amenities fee column. So the intended workflow is report first, then plot the columns that scored badly, which is the right order.

What the document promises about the generated rows

After the first sample, the readme lists three properties of the output, and they are worth reading as claims rather than as results.

Sensitive columns are described as fully anonymized: the email, billing address and credit card number columns contain new data, so the real values are not exposed.

Other columns follow statistical patterns. The examples given are the proportion of room types, the distribution of check in dates, and the correlations between room rate and room type.

Keys and other relationships are intact. The primary key is unique for each row, and with multiple tables the connection between a primary and foreign key makes sense.

Three observations. The anonymization claim is about the demo's output, and it is a claim about value generation rather than about a formal guarantee, so it deserves testing on your own schema before you rely on it. The statistical claims are the kind a distribution-based synthesizer is designed to satisfy, which is also the reason the quality report exists: these are the properties that can degrade quietly.

The key claim is the one people forget. Synthetic data that looks right but produces duplicate or dangling keys will break a foreign key constraint the moment you load it, so uniqueness and referential integrity belong in your acceptance criteria rather than in the score.

Five numpy floors, five pandas floors, and a dependency that stops

The dependency list is where this project's compatibility work is visible, and it is more elaborate than a single requirement would suggest.

The manifest allows Python 3.9 up to but not including 3.15, with classifiers for 3.9 through 3.14. Within that range, numpy appears five times with a different floor in each band: a lower bound for interpreters below 3.10, a higher one for 3.10 and 3.11, another for 3.12, another for 3.13, and another for 3.14, stepping across a major version boundary partway through.

Pandas does the same thing with five floors, all capped below the next major.

Two other entries are version-conditional in a more surprising way. The model library that provides the copula is marked as a dependency only below Python 3.14, so on 3.14 that library is not installed as a declared dependency at all. The serialisation helper jumps a major version at the same boundary.

That is what supporting six interpreter versions looks like in practice: a matrix of environment markers rather than one range. It also means an install on the newest interpreter is materially different from an install on the oldest, and the copula case is the one to check if you plan to run 3.14, because a missing declared dependency surfaces at import time rather than at install time.

The cloud SDKs and the graph library are unconditional runtime dependencies, so they are installed whether or not you store anything remotely or draw anything.

Installation is one line from either package manager:

bash
pip install sdv

The Makefile documents itself and still carries a Python 2 shim

The build file has a detail worth reading because it explains how the project documents itself without a separate tool.

The default goal of the make invocation is help, and help is implemented by an embedded Python script that reads the makefile from standard input, matches lines that end in a double-hash comment, and prints the target name padded to a fixed width with that comment as its description. So the makefile is its own documentation: a target is documented by the comment on its own line, and the help output is generated from those comments.

The same block of embedded Python contains a compatibility shim. It tries to import a path-to-URL helper from one module and falls back to importing it from another, which is the pattern for code that has to run on both Python 2 and Python 3. In a project whose manifest requires Python 3.9 or newer, that fallback can never be exercised.

The rest of the file is housekeeping: a set of clean targets for build artefacts, compiled Python files, documentation, coverage and test caches, plus an install target that installs the package into the active environment and a test-install target that adds the test extra.

The makefile also defines a browser-opening helper for the documentation, which is the usual reason such a file exists at all: a project whose documentation builds locally needs a way to open it.

Three requirement files at the root, and two document formats

The root of the repository is where a project's operating habits show up.

There is a requirements file, and it contains a comment and one line: an editable install with the development extra, described as the requirements for development and for the notebook environment. Alongside it is a second file holding the latest pinned requirements, and a third file listing system packages. That trio, plus a file of static analysis tool versions, is a dependency pinning strategy: system packages, Python packages and linter versions each pinned in their own data file so that a build image and a developer's machine resolve the same tools.

Then the documents. A release process file, a history file, an evaluation file and a contributing guide in the reStructuredText format, sitting next to a readme and other documents in Markdown. Two formats for the same purpose in one root is the kind of inconsistency that survives for years because both render correctly.

The build tooling is a Python task runner file rather than something else, and the package itself lives in a directory named after it, with documentation and test directories beside it. There is a coverage service configuration and a coverage badge in the readme, and two separate workflow badges for unit and integration runs.

Editorial conclusion

SDV fits a data team that needs synthetic tables which keep the shape of real data without carrying real values, and that wants a single library rather than assembling one per modality. Two things to settle before you adopt it. The licence is source-available, not open source: the manifest declares a business source licence, so read the terms against how you intend to use the output rather than assuming the usual permissions. And the quality score is an average of two families of measure over columns and column pairs, which is a useful smoke test and not a certification: the documented example lands just under 89, and a score tells you where to look rather than that the data is safe. If the synthetic output has to satisfy an auditor, budget for the evaluation work rather than the score alone.

Frequently asked questions

What does SDV stand for?

It stands for Synthetic Data Vault, a Python library described as a one-stop shop for creating tabular synthetic data by learning patterns from real data and emulating them. The citation the project asks for is a paper of the same name by three authors hosted at a university data lab.

Is SDV free or open source?

The manifest declares a business source licence and names the licence file, and the readme says the library is publicly available under the Business Source License, so it is source-available rather than open source. The project's own licence metadata field records nothing. The forum, blog and website all sit under the commercial company that maintains it.

What kinds of data can SDV model?

Single tables, multiple connected tables and sequential tables, using models that range from classical statistical methods such as a Gaussian copula to deep learning methods such as CTGAN. The workflow can also be customised for preprocessing, anonymization and business rules expressed as logical constraints.

How does SDV evaluate synthetic data?

It compares the synthetic data against the real data and produces a quality report with an overall score from 0 to 100 and detailed breakdowns. The documented example reports a column shapes score of 89.11 percent over nine columns and a column pair trends score of 88.3 percent over 36 pairs, averaging to 88.7, and a per-column plot helper is provided for inspecting individual columns.

Which Python versions does SDV support?

The manifest allows 3.9 up to but not including 3.15, with classifiers for 3.9 through 3.14. numpy and pandas are each pinned with five different floors across that range, and the copula library is a declared dependency only below Python 3.14.

Official sources

  1. Issues
  2. Project website
  3. README
  4. Releases
  5. sdv-dev/SDV on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/sdv-dev-sdv.svg)](https://hysenlabs.com/projects/sdv-dev-sdv)