fg-data-synthetic: synthetic tabular and time-series data with GANs and Gaussian mixtures
Synthetic data generators for tabular and time-series data
At a glance
- What is it?
- fg-data-synthetic is the renamed ydata-synthetic package, a Python library of generative models for tabular and sequential data. It ships a Streamlit UI, a Gaussian-mixture fast path that runs without a GPU, and a TensorFlow 2.15 dependency that shapes most deployment decisions.
- Who is it for?
- Adopt fg-data-synthetic if you need tabular or time-series synthesis in Python and can live with TensorFlow 2.15 and Python 3.9 or 3.10. Skip it if you need a maintained release cadence, a documented rollback path, or a supported Python 3.11+ environment.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 27 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem fg-data-synthetic solves, and who it is actually for
Teams that want to share a dataset, balance a rare class, or train a model without moving raw records across a boundary usually end up writing their own generator. That work is repetitive: pick an architecture, encode categorical columns, sample, then check whether the output resembles the input. fg-data-synthetic packages that loop. The README describes it as "a package to generate synthetic tabular and time-series data leveraging the state of the art generative models", and the repository layout backs that up with separate example directories for regular tabular models and for time series.
The audience is narrow in a useful way. It is for Python engineers and data scientists who already work with pandas DataFrames and want a synthesizer they can call from a notebook or a script. The README lists the intended applications as privacy compliance for data sharing and machine learning development, bias removal, dataset balancing, and augmentation. Those are all cases where you have real data and need more of it, or need a version of it you can hand to someone else.
It is not a data platform. There is no server, no scheduler, no access control layer. The repository contains a Streamlit app for a guided flow, and the README points to YData Fabric for an end-to-end product, which is a clear signal that this package is the library layer rather than the product layer.
How the synthesizers are put together: GANs, a Gaussian-mixture fast path, and TensorFlow 2.15
The repository holds architectures and models for synthetic data, from generative adversarial networks to Gaussian mixtures, and the README states that all the deep learning models are implemented on TensorFlow 2.0. That single decision propagates through everything: requirements.txt pins tensorflow==2.15.*, and the badges in the README advertise Python 3.9 and 3.10 only. If your environment is on a newer interpreter, the install is where you will find out.
The model list is the other half of the architecture story. The tabular synthesizers named in the README are CGAN, WGAN, WGANGP, DRAGAN, CRAMER and CTGAN. CTGAN is called out separately as a conditional architecture for tabular data, which matters because conditional generation is what lets you ask for samples that respect a chosen column value rather than sampling the joint distribution blindly. For sequential data the examples reference TimeGAN and DoppelGANger, each with its own notebook.
The fast path is a Gaussian-mixture model. The README presents it as a way to "quickstart in the world of synthetic data generation without the need for a GPU". That is a real trade-off rather than a marketing line: a mixture model is cheaper to fit and easier to run on a laptop, but it captures the data as a weighted combination of components, so it will not reproduce the long-range structure that a recurrent or adversarial time-series model is designed for. Choose it when you need a first synthetic table in minutes, not when you need realistic sequential dependencies.
Data flow is conventional. You load a DataFrame, you fit a synthesizer, you sample. The repository's examples directory contains the concrete versions of that flow for adult census income, credit card fraud, stock data and the FCC MBA dataset.
Installing fg-data-synthetic and generating a first synthetic table
The package is on PyPI as fg-data-synthetic. The README's migration guide exists because the project was previously published as ydata-synthetic, and the install step for the new name is a plain pip install.
pip install fg-data-syntheticIf you were on the old package, the migration guide is explicit about the order: uninstall first, then install the new distribution. The imports change too, from ydata_synthetic to data_synthetic.
pip uninstall ydata-synthetic
pip install fg-data-syntheticThe README suggests this one-liner to find every file still referencing the old module name before you rename anything.
grep -r "ydata_synthetic" . --include="*.py"The import path after migration keeps the same subpackage structure, so a synthesizer that was imported from ydata_synthetic.synthesizers.regular becomes data_synthetic.synthesizers.regular. The README shows that replacement directly.
# Before
import ydata_synthetic
from ydata_synthetic.synthesizers.regular import RegularSynthesizer
# After
import data_synthetic
from data_synthetic.synthesizers.regular import RegularSynthesizerFor a guided first run, the Streamlit interface has its own extra. The README notes that the app is available from v1.0.0 onwards and supports two flows: training a synthesizer model, and generating and profiling synthetic data samples.
pip install fg-data-synthetic[streamlit]Then launch it from a Python file. The README is explicit that Jupyter Notebooks are not supported for this entry point, and that the same app exists as examples/streamlit_app.py if you prefer to run the file directly.
from data_synthetic import streamlit_app
streamlit_app.run()The README also gives a module-style launch command, python -m streamlit_app. Either way, what you should see is the Streamlit UI with the model picker limited to the architectures the README lists for the app: CGAN, WGAN, WGANGP, DRAGAN, CRAMER and CTGAN.
Where fg-data-synthetic breaks down or is the wrong tool
The dependency pin is the first hard limit. requirements.txt holds tensorflow==2.15.*, and the README badges claim Python 3.9 and 3.10. If your production environment is on a newer Python, you are either building a separate environment for synthesis or you are not using this package. That is not a bug, but it is a constraint that shows up on day one rather than after a prototype.
The second limit is release cadence. The most recent release listed is 2.1.0 on 2026-04-23, following 2.0.1 the same day and 2.0.0 back in 2024. The repository's last push was on 2026-09-03, so work is happening, but the version history shows long gaps between published versions. If you need a fix shipped to PyPI on a predictable schedule, plan for the possibility that you are working from the default branch rather than from a release.
The third is documentation depth. The README covers installation, the migration, the Streamlit app and the example notebooks. It does not document evaluation metrics for the generated data, it does not describe how to persist and reload a trained synthesizer, and it does not describe privacy guarantees. The README's own framing of synthetic data is that it "replicates the statistical components of real data without containing any identifiable information, ensuring individuals' privacy", which is a claim about intent rather than a documented guarantee. If your use case depends on a privacy argument, that argument is not in this repository.
Finally, the packaging metadata disagrees with itself. setup.py carries the classifier 'License :: OSI Approved :: GNU General Public License v3 or later (GPLv3+)', while the repository metadata and the README badge point at MIT. Until that is reconciled, treat the licence as unresolved rather than picking the one you prefer.
How fg-data-synthetic differs from SDV and from writing your own generator
SDV is the closest alternative for tabular synthesis, and the difference is in scope rather than in quality. SDV is built around a DataFrame-first API with its own metadata layer that describes column types and relationships before you fit anything; that metadata step is where multi-table and referential structure gets handled. fg-data-synthetic does not expose a comparable metadata model in the README. What it exposes instead is a set of named architectures, CGAN through CTGAN, plus a Gaussian-mixture fast path and a Streamlit UI that lets a non-programmer drive training and sampling. If your problem is a single flat table and you want to pick an architecture by name, fg-data-synthetic is the more direct route. If your problem is a schema with foreign keys, the metadata layer in SDV is doing work this package does not appear to do.
Against writing your own generator, the difference is the examples. The repository ships runnable notebooks for adult census income with CTGAN, a fast Gaussian-mixture run on the same dataset, TimeGAN on stock data, and DoppelGANger on the FCC MBA dataset. That is a starting point you can modify rather than a blank file. The cost is the TensorFlow 2.15 pin, which you would not inherit if you wrote a small generator against a library you already run.
Maintenance, upgrade cost and the licence question
The repository is not archived, and the last push was on 2026-09-03. That is recent enough that the default branch is moving, but the published release history is uneven: 2.0.0 in September 2024, then 2.0.1 and 2.1.0 both on 2026-04-23. Anyone pinning to a release should assume they may be waiting a while for the next one.
The upgrade cost is dominated by TensorFlow. Because requirements.txt pins tensorflow==2.15.*, moving the package forward means moving that pin, and any environment that shares a TensorFlow installation with other tooling will need coordination. The migration from ydata-synthetic to fg-data-synthetic is the other upgrade you may still owe: it is a rename of the distribution and the import root, and the README provides both the uninstall/install sequence and the grep command to find affected files.
The licence needs a decision from the maintainers, not from you. setup.py declares GPLv3 or later; the repository metadata and the README badge say MIT. Those are very different obligations for anyone embedding the library in a distributed product. This is not legal advice, and the practical step is to open an issue and get the classifier corrected before you rely on either reading.
Editorial conclusion
Adopt fg-data-synthetic if you need tabular or time-series synthesis in Python and can live with TensorFlow 2.15 and Python 3.9 or 3.10. Skip it if you need a maintained release cadence, a documented rollback path, or a supported Python 3.11+ environment. Before committing, run pip install fg-data-synthetic[streamlit] on your target interpreter, confirm the TensorFlow pin resolves against your existing stack, and open one of the example notebooks in examples/regular/ to check that the synthesizer API matches the version you installed. The repository's setup.py classifier says GPLv3+ while the repository metadata says MIT, and that discrepancy is the first thing to resolve with the maintainers.
Frequently asked questions
What does it mean if a dataset is synthetic?
The README defines synthetic data as artificially generated data that is not collected from real world events, replicating the statistical components of real data without containing identifiable information.
Is synthetic data better than real data?
The README does not make a quality comparison between the two. It presents synthetic data as a way to handle privacy compliance for data sharing and machine learning development, and to remove bias, balance datasets and augment datasets.
Can you train AI on synthetic data?
The README lists machine learning development among the applications of synthetic data, alongside privacy compliance for data sharing. It does not document accuracy comparisons between models trained on synthetic versus real data.
Which AI is best for synthetic data generation?
The README does not rank the models. It names CGAN, WGAN, WGANGP, DRAGAN, CRAMER and CTGAN for tabular data, TimeGAN and DoppelGANger in the time-series examples, and a Gaussian-mixture model as the fast option that does not need a GPU.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/data-centric-ai-community-fg-data-synthetic)