# fg-data-profiling: one line of EDA for pandas and Spark DataFrames

> A renamed fork of pandas-profiling that turns a DataFrame into an HTML, JSON or notebook report in a single call. Here is what the report contains, how the rename affects existing code, and where the tool stops being the right choice.

**Data-Centric-AI-Community/fg-data-profiling** — 1 Line of code data quality profiling & exploratory data analysis for Pandas and Spark DataFrames. 

- Repository: https://github.com/Data-Centric-AI-Community/fg-data-profiling
- Website: https://docs.sdk.ydata.ai
- Stars: 13,696 · Forks: 1,796
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/data-centric-ai-community-fg-data-profiling

## What fg-data-profiling produces that df.describe() does not

The README frames the goal narrowly: a one-line Exploratory Data Analysis experience, in the same spirit as pandas df.describe(), but with an extended analysis of a DataFrame that can be exported to HTML and JSON. df.describe() gives you counts, means and quantiles for numeric columns. fg-data-profiling adds type inference (Categorical, Numerical, Date and others), an alerts section that lists potential data quality issues such as high correlation, skewness, uniformity, zeros, missing values and constant values, univariate statistics with distribution histograms, multivariate analysis covering correlations, missing-data patterns and duplicate rows, plus dedicated handling for time-series, text, file and image columns.

The audience is anyone who receives an unfamiliar table and needs a first pass before modelling or before handing data to someone else. The output is a report, not a cleaned dataset. Nothing in the README suggests the package repairs data; it flags problems and leaves the decision to you. That distinction matters when a team expects profiling to be a validation gate. It is a diagnostic, and the alerts section is the part worth reading first.

## The rename from ydata-profiling to fg-data-profiling

This is the single most important operational fact about the project right now. The README states plainly that ydata-profiling is now fg-data-profiling, and that the old package will no longer receive updates or bug fixes. The migration guide is three steps: uninstall the old distribution, install the new one, and replace imports. The import name changes too, from ydata_profiling to data_profiling, which is easy to miss because the distribution name and the module name differ.

If you have notebooks, scripts or internal libraries that import the old module, the README offers a grep one-liner to find affected files. Treat this as a real migration rather than a version bump. A pinned requirements file that still says ydata-profiling will keep resolving to a package the maintainers have said they will not patch. The repository's default branch is develop, and the most recent release listed is 4.19.1, dated 2026-04-22, so the rename has already shipped in a tagged release rather than only on main.

## Installing fg-data-profiling and generating a first report

The README gives two install paths: pip and conda-forge. The pyproject.toml declares Python >=3.10,<3.15, so 3.10 through 3.14 are the supported interpreters. Install with pip:

```bash
pip install fg-data-profiling
```

Or, if your environment is conda-managed, the README lists the conda-forge channel instead:

```bash
conda install -c conda-forge fg-data-profiling
```

With the package installed, the quickstart builds a small random DataFrame and profiles it. Note the import path: the module is data_profiling, not fg_data_profiling.

```python
import numpy as np
import pandas as pd
from data_profiling import ProfileReport

df = pd.DataFrame(np.random.rand(100, 5), columns=["a", "b", "c", "d", "e"])
profile = ProfileReport(df, title="Profiling Report")
```

That call constructs the report object. The README describes three output routes: an HTML report for sharing, JSON for integration into automated systems, and a widget inside a Jupyter Notebook. The examples directory in the repository ships datasets such as titanic, census, bank_marketing_data and meteorites, and the Makefile has an examples target that runs every .py file under ./examples, which is the fastest way to see what a full report looks like on real tables rather than on random floats.

## Spark support and the constraints it carries

Spark is listed as a released integration rather than a roadmap item, and the README links to a dedicated Spark page in the documentation. The Makefile separates Spark work from the default test run: test_spark targets tests/backends/spark_backend/, while the ordinary test target runs tests/unit/, tests/issues/ and the notebook tests. That split is a useful signal about how the project treats the two backends. Spark profiling is exercised by its own suite.

The README also says the project is looking for contributors on Spark and links to a work-in-progress project board, which suggests the integration is not finished. The practical constraint is where the work happens. A profiling report requires collecting statistics from the data, and on a large cluster that collection runs through the driver. The README does not document a distributed execution plan for the report generation itself. If your Spark jobs are sized so that the driver cannot hold the intermediate statistics, this is the point to check before you plan around it. The documentation page linked from the README is the place that would answer it; the README alone does not.

## Where fg-data-profiling is the wrong tool

The HTML report is the product, and that shapes its limits. A profiling report on a wide table with many columns and many rows is a large artifact, and the README does not document a sampling or row-limit setting on the quickstart path. If your table has millions of rows, the report generation cost is a question the README leaves open. Test it on your own data before you put it in a pipeline.

The second limit is scope. The README explicitly points readers who want a scalable solution integrated with database systems toward YData Fabric Data Catalog, naming Oracle, Snowflake, PostGreSQL, GCS and S3 as the systems it connects to. That is a commercial product, and the README's own framing is that the open source package is the local, DataFrame-shaped tool. If your requirement is profiling tables in place inside a warehouse with access control and a catalog, the open source package is not that. If your requirement is a quick, reproducible EDA artifact from a DataFrame you already have in memory, it is.

The third limit is that the report is a snapshot. The README describes a comparison feature for two datasets, but the output formats are HTML, JSON and a notebook widget. There is no documented expectation that a report is diffed against a previous run as a regression gate. Teams that want data quality monitoring over time are looking at a different category of tool.

## How it compares to Great Expectations and to plain pandas

The nearest alternative in the data quality space is Great Expectations. The difference in approach is what you write first. Great Expectations asks you to declare expectations about your data (a column must not be null, values must fall in a range), then validates incoming batches against them and fails when they break. fg-data-profiling asks you to declare nothing. You hand it a DataFrame and it tells you what it noticed. One is a test suite for data, the other is a first look at it.

That makes them complementary rather than substitutes. The alerts section of an fg-data-profiling report is a reasonable place to find the expectations worth writing down, and Great Expectations is where you would encode them. Against plain pandas, the comparison is simpler: df.describe() and a handful of value_counts calls get you part of the way, and if that is genuinely enough for your table, adding a dependency is not worth it. The case for fg-data-profiling is tables with mixed column types, time-series columns or text columns, where the per-type analysis and the alerts list would otherwise be code you write and maintain yourself.

## Licence, maintenance and what an upgrade costs

The project is MIT licensed, and the pyproject.toml carries the matching OSI classifier. MIT is permissive: you can use, modify and redistribute it, including in commercial and closed-source work, provided the copyright notice and licence text are preserved. That is the general shape of the licence, not legal advice; read LICENSE.md in the repository if the terms matter to your organisation.

The dependency ranges in pyproject.toml are the real upgrade cost. pandas is pinned to >1.5,<3.0 and explicitly excludes 1.4.0. numpy is >=1.22,<2.6. scipy is >=1.8,<1.19.0. matplotlib is >=3.5,<3.12. pydantic is >=2,<3, and visions is >=0.7.5,<0.9. Those upper bounds mean a major pandas or numpy release can leave the package unresolvable until the maintainers widen the range. If your environment pins a newer numpy than 2.6, expect a conflict.

On maintenance: the most recent release listed is 4.19.1 on 2026-04-22, and the last push to the repository was on 2026-04-22. That is roughly five months before today, so the project is not being pushed to continuously. The repository is not archived. The README's own warning about the old package name is the thing to act on, and the migration is a find-and-replace plus a dependency change, which is cheap now and more expensive the longer it waits.

## Conclusion

Adopt fg-data-profiling if you want a shareable EDA artifact from a pandas DataFrame without writing plotting code, and budget the rename work first: run grep -r "ydata_profiling" . --include="*.py", uninstall ydata-profiling, then install fg-data-profiling and switch imports to data_profiling. Skip it if you need a governed catalog across Snowflake, Postgres or S3 (the README points at YData Fabric for that), or if your Spark jobs already run on a cluster where pulling a profiling library onto the driver is not acceptable. Verify two things before committing: that your pinned pandas, numpy and scipy versions fall inside the ranges in pyproject.toml, and that the report you get on your own wide table is one you are willing to hand to a reviewer.

## FAQ

### What does data profiling do?

It produces a structured summary of a dataset before you analyse or model it. In fg-data-profiling that summary covers type inference, univariate statistics with histograms, multivariate correlations, missing-data patterns, duplicate rows, and an alerts section listing potential problems such as skewness, uniformity, zeros and constant values.

### What is Pandas Profiling and how does it work?

Pandas Profiling is the earlier name of this lineage of tool; the README states that ydata-profiling is now fg-data-profiling, and that the old package will no longer receive updates or bug fixes. It works by taking a DataFrame and generating a report covering type inference, univariate and multivariate analysis, time-series and text handling, exportable as HTML, JSON or a Jupyter widget.

### What is data profiling in ETL?

The README does not describe an ETL integration or scheduling model for fg-data-profiling. The package's stated outputs are an HTML report, JSON for integration into automated systems, and a notebook widget, so any ETL use would be built around the JSON export rather than provided by the tool.

### What are the best data profiling tools?

The README does not compare fg-data-profiling to other tools by name, so it offers no ranking. It does distinguish the package from YData Fabric Data Catalog, which it describes as the scalable option for connecting to databases and storage systems such as Oracle, Snowflake, PostGreSQL, GCS and S3.

## Sources

- [Data-Centric-AI-Community/fg-data-profiling on GitHub](https://github.com/Data-Centric-AI-Community/fg-data-profiling)
- [License: MIT](https://github.com/Data-Centric-AI-Community/fg-data-profiling/blob/develop/LICENSE)
- [Project website](https://docs.sdk.ydata.ai)
- [README](https://github.com/Data-Centric-AI-Community/fg-data-profiling/blob/develop/README.md)
- [Releases](https://github.com/Data-Centric-AI-Community/fg-data-profiling/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/data-centric-ai-community-fg-data-profiling
