Great Expectations (GX Core): data quality tests for pandas and SQL pipelines
Always know what to expect from your data.
At a glance
- What is it?
- GX Core turns data quality checks into Python objects you can version, run and document. It installs with pip, supports Python 3.10 through 3.13, and is licensed Apache-2.0. The hard part is not the library, it is deciding which checks are worth keeping.
- Who is it for?
- Adopt GX Core if you already write Python against pandas or SQL and want data checks that live in the same repository as the pipeline code. Skip it if you need a hosted dashboard with no code, or if your team will not maintain the Expectation suites after the first sprint.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What GX Core actually solves
The README describes Expectations as "expressive and extensible unit tests for your data". That framing is the whole product. A pipeline that loads a CSV, joins two tables or writes a partition can succeed at the transport level and still be wrong: a column that used to hold integers now holds strings, a date range silently shrank, a category list grew a typo. GX Core gives those assumptions a name, a place in the repository and a pass or fail result.
The intended audience is data engineers and analytics engineers who already work in Python. The topics list on the repository confirms the orientation: data-quality, data-unit-tests, pipeline-testing, mlops. There is no claim of a no-code product here. You write Python, you get validation results, and the library can generate documentation for each set of results so the rest of the organisation can read what was checked.
The second problem it addresses is institutional memory. When the person who wrote the ingestion job leaves, the implicit knowledge about acceptable null rates and valid ranges leaves with them. An Expectation suite is that knowledge written down in a form a machine re-runs.
How Expectations, Data Contexts and validation results fit together
The README's quickstart shows two objects. `gx.get_context()` returns a Data Context, which is the entry point that holds configuration and connections. Expectations are the individual assertions. Validation runs an Expectation suite against a batch of data and produces results, and the README states that documentation can be generated automatically for each set of validation results.
The dependency list in requirements.txt tells you what the core install drags in: pandas, numpy, scipy, pydantic, marshmallow, jsonschema, jinja2, altair, ruamel.yaml. That is a data-frame-and-reporting stack, not a thin assertion library. If you only want to check that a column is non-null, you are paying for a lot of machinery. If you want generated HTML documentation and a schema for the suite itself, the dependencies are the reason those features exist.
The repository layout reinforces the split. There is a `great_expectations/` package, a `contrib/` directory that pyproject.toml explicitly excludes from mypy checking, and a `docs/` tree with a Docusaurus site. The `contrib` exclusion is worth noting: integrations that live there are not held to the same type-checking bar as the core package.
Installing GX Core and running a first check
The README recommends deploying GX Core inside a virtual environment and running the install from an empty base directory. The command is a plain pip install with no extras:
pip install great_expectationsAfter that, the README's second step is to import the module and create a Data Context. The call takes no arguments in the example, so it resolves a context from the working directory or a default:
import great_expectations as gx
context = gx.get_context()If that runs without an exception, the library is importable and a context object exists. The README stops there. It does not show connecting a data source, defining an Expectation or running a validation in the quickstart; it points to the Introduction to GX Core page in the documentation for the next steps. That is a real gap for anyone evaluating the tool from the README alone.
Python version support is explicit. The README states GX Core supports Python 3.10 through 3.13, and setup.py encodes that as `>=3.10,<3.14`. The README also notes that experimental support for Python 3.14 and later can be enabled by setting a `GX_PYTHON_EXPERIMENTAL` environment variable when installing `great_expectations`. setup.py shows that this variable removes the upper bound entirely, returning `>=3.10` rather than the pinned range. Treat that as a way to try a newer interpreter, not as a supported configuration.
Where GX Core is the wrong tool
The library is Python-only. There is no CLI-only or language-agnostic path described in the README, so a team whose pipelines are written in Scala, Go or pure SQL with no Python orchestration layer gets nothing from it without adding a Python service that reaches into those systems.
The integration policy is another boundary. The README says GX supports a set of data sources and integrations and defers to a compatibility reference in the docs for the details. It does not enumerate them in the README. That means the question "does it work with my warehouse" cannot be answered from the repository's front page, and the answer changes between releases.
There is also a maintenance cost that the README does not discuss. An Expectation suite is code. It needs owners, review and updates when the schema legitimately changes. A suite written during a migration and never revisited will fail on the first intentional column rename and then get disabled. The library makes the checks cheap to write, not cheap to keep.
The dependency surface is the last consideration. Pulling pandas, scipy and altair into an environment that already pins conflicting versions of numpy is a common source of resolution pain. requirements.txt carries separate numpy pins for Python 3.10, 3.12 and 3.13, which is a sign the project tracks those boundaries carefully, but your own constraints file still has to agree with them.
GX Core compared with writing plain assertions
The obvious alternative is a handful of `assert` statements or a pytest file that loads a sample and checks it. That approach has real advantages: no new dependency, no configuration format, and the checks run in the same test runner as everything else.
The difference in approach is what happens after the check fails. A bare assertion tells you a condition was false. GX Core's model separates the Expectation (what should be true), the suite (the collection), the batch (what was actually tested) and the validation result (the outcome), and the README states that documentation can be generated for each set of validation results. That structure is what makes a failure readable by someone who did not write the check, and it is what lets the same suite run against a sample in CI and a full partition in production.
A second alternative is a metrics-and-alerting tool that monitors tables from the outside. Those tools observe data that already landed. GX Core is designed to be called from inside the pipeline, so a failed validation can stop a load before bad rows reach consumers. The trade-off is that you must wire that call in yourself; the library does not sit between your pipeline and your warehouse by default.
Maintenance, releases and the Apache-2.0 licence
The repository is not archived. The last push was on 2026-09-18, the same day release 1.23.1 was published, following 1.23.0 on 2026-09-10 and 1.22.0 on 2026-08-31. That cadence, three minor releases in under three weeks, means the surface you pin today will move. Pin a version in your requirements file rather than tracking the latest, and read the release notes before upgrading, because the README does not document a rollback procedure and nothing in the repository layout suggests a migration tool.
Licensing is Apache-2.0, stated in the repository and confirmed by the LICENSE file at the top level. Apache-2.0 is a permissive licence that includes an explicit patent grant, which matters if you are embedding the library in a commercial product. It also means there is no copyleft obligation on your own code. This is a description of the licence text, not legal advice; your legal team should read LICENSE and the CLA.md and CONTRIBUTING.md files if you plan to contribute code back.
The presence of AGENTS.md, CONTRIBUTING_CODE.md and CONTRIBUTING_WORKFLOWS.md at the top level indicates a project with documented contributor processes. For an adopter that is mildly reassuring about governance, but it says nothing about whether the specific data source you need is supported.
Editorial conclusion
Adopt GX Core if you already write Python against pandas or SQL and want data checks that live in the same repository as the pipeline code. Skip it if you need a hosted dashboard with no code, or if your team will not maintain the Expectation suites after the first sprint. Before committing, verify the compatibility reference for your specific data source, confirm Python 3.10 through 3.13 matches your runtime, and read the Apache-2.0 LICENSE file in the repository for how the project itself is licensed.
Frequently asked questions
How do I install Great Expectations in Python?
Run `pip install great_expectations` inside a virtual environment, from an empty base directory, as the README instructs. The README recommends a virtual environment and gives that single pip command as the install step. Python 3.10 through 3.13 is the supported range.
How do I use Great Expectations in Python?
Import the module as `great_expectations as gx` and call `gx.get_context()` to create a Data Context, which is the entry point the README's quickstart shows. From there the README points to the Introduction to GX Core documentation for defining Expectations and running validations.
What is Great Expectations as a data quality tool?
It is a Python library whose README describes Expectations as expressive and extensible unit tests for your data. Validation runs a suite of Expectations against a batch of data and produces results, and the README states documentation can be generated for each set of validation results.
What is Great Expectations in data engineering?
It is the GX Core library, aimed at data engineers who write Python and want checks on pandas or SQL data inside their pipelines. The repository's topics include data-quality, pipeline-testing and data-unit-tests, which matches that positioning.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/fivetran-great-expectations)