Library / SDK
Data-Centric-AI-Community/fg-data-quality avatar
Data-Centric-AI-Community/fg-data-quality

fg-data-quality: a DataQuality object that scores a pandas DataFrame in one call

Data Quality assessment with one line of code

456 stars56 forksJupyter NotebookMIT

At a glance

What is it?
The fg-data-quality library wraps duplicates, drift, collinearity, fairness and missing-value checks behind a single DataQuality class. It is a fast screening pass for tabular data, not a pipeline gate, and the README leaves several limits unsaid.
Who is it for?
Adopt fg-data-quality if you have a pandas DataFrame and want a first-pass quality report before modelling, and you accept the pinned dependency set. Do not adopt it expecting an automated pipeline gate or a maintained release cadence: the last push was on 2026-04-23, the classifiers mark it Alpha, and the README documents no threshold configuration, no CI integration and no rollback path.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 160 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What fg-data-quality checks, and who the warning output is for

The library targets a narrow but common moment: you have loaded a table into pandas, you are about to train a model or hand the data to someone else, and you want a structured list of things that look wrong before you spend time on them. The README positions it as an assessment tool for "the multiple stages of a data pipeline development", which in practice means the notebook or script where a DataFrame first appears. The audience is data scientists and analytics engineers working in Python, which the setup.py classifiers confirm with Intended Audience entries for Developers, Science/Research and several industry verticals. It is not a schema validator and it is not a data contract system. It reads a DataFrame you already have in memory and reports on it. The unit of output is a warning with a priority level, not a pass or fail. That distinction matters: the README's own example prints five warnings and then says Priority 2 warnings are "usage allowed, limited human intelligibility", which is a judgement call handed back to you rather than a decision the library makes.

How the DataQuality engine turns one DataFrame into prioritised warnings

The mechanism is visible in the quickstart. You construct a DataQuality object with the DataFrame as the df argument, then call evaluate(). The class is described as holding "all quality modules", so the object is a container and evaluate() is the run loop across them. The README's sample output shows the modules that fired on a census dataset: duplicate columns, high collinearity among numerical variables measured with Variance Inflation Factor above 5.0, high collinearity among categorical variables at a p-value below 0.05, predefined erroneous data, and exact duplicate rows. Those names map onto the tutorial notebooks listed in the README, which cover bias and fairness, data expectations, data relations, drift, duplicates, labelling for categoricals and numericals, missings, and erroneous data. So the architecture is a set of independent check modules plus a reporting layer that buckets findings into Priority 1 ("heavy impact expected") and Priority 2. The second call, dq.get_warnings(), returns the same findings as a list rather than a printed summary, which is the hook you would use if you wanted to act on them programmatically. The README does not document how to disable individual modules, change the VIF or p-value thresholds, or reorder priorities.

Installing fg-data-quality and running a first evaluation

Installation is a single pip command from PyPI. The README gives it directly, and setup.py confirms the distribution name is fg-data-quality while the importable package is data_quality.

bash
pip install fg-data-quality

The README badges list Python 3.6, 3.7 and 3.8, while setup.py declares python_requires=">=3.7". Those two statements do not agree, and the pinned requirements.txt is the more practical constraint: dython==0.6.7, pandas==1.2.*, numpy==1.20.3, scikit-learn==0.24.2, statsmodels==0.12.2, pydantic==1.8.2 and matplotlib==3.4.2. A modern environment will likely need its own resolution rather than these exact pins.

The first real use is three lines, taken from the quickstart. Load a CSV into a DataFrame, wrap it, evaluate it.

python
from data_quality import DataQuality
import pandas as pd

df = pd.read_csv('./datasets/transformed/census_10k.csv')
dq = DataQuality(df=df)
results = dq.evaluate()

What you should see is a printed block headed Warnings, with a total count, a breakdown by priority, and one line per finding. The README's example shows "TOTAL: 5 warning(s)" split into one Priority 1 and four Priority 2 entries, each line naming the module in brackets, for example [DUPLICATES - DUPLICATE COLUMNS] or [DATA RELATIONS - HIGH COLLINEARITY - NUMERICAL]. To get the same content as data instead of console text, call dq.get_warnings().

python
warnings = dq.get_warnings()

If you are migrating from the older ydata-quality package, the README provides a migration guide: uninstall the old distribution, install fg-data-quality, and replace imports. The README suggests a grep to find affected files.

bash
grep -r "ydata_quality" . --include="*.py"

The pinned dependency set is the first thing that will break

The most concrete limitation is not in the warning logic, it is in requirements.txt. Exact pins on numpy 1.20.3, scikit-learn 0.24.2 and statsmodels 0.12.2 will conflict with almost any environment built in the last few years, and pip will either refuse or silently upgrade around them depending on how the library declares its dependencies. The README does not discuss this, and setup.py's truncated dependency section means the published wheel's actual constraints are not visible here. A second limitation is the output model itself. Priorities are assigned by the library, and the README offers no documented way to change thresholds such as VIF>5.0 or p-value < 0.05, so a team with different tolerances cannot encode them without touching the source. Third, the checks are statistical heuristics over a DataFrame in memory. High collinearity and duplicate detection are useful signals, but the README's own Priority 2 wording concedes that some findings are advisory. If your goal is a hard gate that blocks a deployment, this library does not provide one; it provides a report you still have to interpret. Finally, the package classifiers mark it Development Status :: 3 - Alpha, which is a signal about API stability that the README does not repeat.

fg-data-quality versus Great Expectations and pandas-profiling

The closest comparison is Great Expectations, and the difference is in what each one treats as the source of truth. Great Expectations is built around expectations you declare in advance: you write assertions about columns, ranges and nullability, and the tool checks data against them, which makes it suitable for running inside a pipeline where a failure should stop things. fg-data-quality works the other way around. You declare nothing. The DataQuality object inspects the DataFrame and surfaces what it finds, which is why the README can promise a comprehensive check in a few lines. That makes it better for exploration and worse for enforcement. The other common comparison is pandas-profiling, now distributed under a different name. That tool produces a descriptive report per column: distributions, quantiles, missingness, correlations. fg-data-quality overlaps on missing values and correlations but adds model-oriented checks that a profiling report does not frame as warnings, notably variance inflation factor and categorical collinearity, plus bias and fairness and drift modules. If you need a column-by-column statistical profile, profiling is the more direct tool. If you need a short list of things that will hurt a model, the warning summary here is closer to the point.

Maintenance status, licence and the cost of upgrading

The repository is not archived, but the last push was on 2026-04-23, and the three releases listed (0.1.1, 0.1.2 and 0.2.0) all landed on that same day. That is a burst of release activity rather than a steady cadence, and nothing shown here indicates activity after that date. Treat the project as stable-but-quiet and plan to read the source when something breaks, because the README documents no deprecation policy and no upgrade path beyond the ydata-quality migration guide. The licence situation is worth reading carefully. The README's License section links to GNU General Public License v3.0, and setup.py carries the classifier GNU General Public License v3 or later (GPLv3+), while the repository metadata listed alongside it says MIT. Those disagree, and the LICENSE file in the repository root is the one that governs. GPLv3 is copyleft: if you distribute software that incorporates this library, the licence terms attach to that distribution in ways an MIT or Apache-2.0 dependency would not. This is not legal advice; if you are shipping a product, have someone check the LICENSE file against your distribution model. On upgrade cost, the pinned requirements are the practical tax. Installing fg-data-quality into an existing project means either accepting old numpy and scikit-learn or resolving conflicts yourself, and the README gives no guidance on which pins are load-bearing.

Editorial conclusion

Adopt fg-data-quality if you have a pandas DataFrame and want a first-pass quality report before modelling, and you accept the pinned dependency set. Do not adopt it expecting an automated pipeline gate or a maintained release cadence: the last push was on 2026-04-23, the classifiers mark it Alpha, and the README documents no threshold configuration, no CI integration and no rollback path. Verify first that your Python and pandas versions can coexist with requirements.txt, that the warning list from dq.get_warnings() is something your team will actually triage, and that the GPLv3 licence fits how you distribute your own code.

Frequently asked questions

What does fg-data-quality actually do?

It is a Python library that assesses a pandas DataFrame across several quality dimensions and prints a prioritised warning summary. You create a DataQuality object with your DataFrame and call evaluate().

How do I install fg-data-quality?

The README gives one command, pip install fg-data-quality, which installs the package from PyPI. The importable module is data_quality, not the distribution name.

How do I get the warnings as data instead of printed text?

After calling evaluate(), call dq.get_warnings(), which the README describes as retrieving a list of detected data quality warnings for detailed inspection.

Can I move from ydata-quality to fg-data-quality?

Yes, and the README has a migration guide: uninstall ydata-quality, install fg-data-quality, then replace the old imports with data_quality. It suggests grepping for ydata_quality across your .py files to find the affected code.

What Python versions does fg-data-quality support?

The README badges list Python 3.6, 3.7 and 3.8, while setup.py declares python_requires=">=3.7". The pinned requirements.txt is the tighter constraint in practice.

Official sources

  1. Data-Centric-AI-Community/fg-data-quality on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/data-centric-ai-community-fg-data-quality.svg)](https://hysenlabs.com/projects/data-centric-ai-community-fg-data-quality)