deequ: writing assertions about data instead of hoping the pipeline held
Deequ is a library built on top of Apache Spark for defining "unit tests for data", which measure data quality in large datasets.
At a glance
- What is it?
- An Apache Spark library that treats data quality constraints as unit tests, translating a list of assertions into Spark jobs and returning a per-constraint verdict.
- Who is it for?
- deequ makes sense the moment a pipeline feeds something you cannot easily inspect, because it turns implicit assumptions about columns into a result object you can assert on in CI. The important thing to get right before adopting it is the artifact suffix: a Deequ build is compiled against one Spark version, and mixing it with another is the most likely reason a first attempt fails.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 21 days ago.
- What is it written in?
- Mainly Scala, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Assertions on data, expressed the same way as assertions on code
The framing in the README is deliberate. Deequ is a library on top of Apache Spark for defining unit tests for data, and the pitch is that most data applications carry implicit assumptions about their input: attributes have certain types, columns do not contain NULL, an ID column is unique. When those assumptions break, the application crashes or silently produces wrong output, usually somewhere downstream of the actual cause.
The API follows that metaphor closely. `VerificationSuite` is the entry point, you attach `Check` objects to it, and each check is a set of constraints on attributes of the data. A check carries a name and a `CheckLevel`, and the README's own example uses `CheckLevel.Error` with the name `unit testing my data`.
What makes this readable is that the constraints look like assertions. `.hasSize(_ == 5)`, `.isComplete("id")`, `.isUnique("id")`, `.isContainedIn("priority", Array("high", "low"))`, `.isNonNegative("numViews")`. A developer who has written Scala tests already knows how to read that, and the underscore syntax keeps the Scala function literal rather than a string-based rule language.
Picking the artifact that matches your Spark version
This is the practical hurdle and the README handles it head on. Deequ releases are built for specific Apache Spark versions, and the artifact suffix has to match the Spark version your application uses:
| Spark 3.1.x | Deequ 2.x | `-spark-3.1` |
| Spark 3.5.x | Deequ 2.x | `-spark-3.5` |The dependency declaration itself is unremarkable:
<dependency>
<groupId>com.amazon.deequ</groupId>
<artifactId>deequ</artifactId>
<version>2.0.21-spark-3.5</version>
</dependency>For sbt the same choice is a single line: `libraryDependencies += "com.amazon.deequ" % "deequ" % "2.0.21-spark-3.5"`. Below that, the Java floor matters too. Deequ 2.1.0 and later require Java 11, earlier 2.0.x releases require Java 8, and Spark 3.0 support now lives on a `legacy-spark-3.0` branch. Releases for Spark 2.2.x and 2.3.x were built against Scala 2.11 while later supported Spark releases use Scala 2.12.
The repository is a Maven project with a `pom.xml`, a `Makefile` whose `build` target is `mvn clean install`, and a `deequ-scalastyle.xml` for style enforcement. That last file is a small signal about how seriously the maintainers treat consistency, which matters in a library where the constraint names are part of the contract.
The three data frames in the README example
The walkthrough in the README is worth reading end to end because it shows the level of abstraction the library expects. It starts with a case class describing an item:
+case class Item(
id: Long,
productName: String,
description: String,
priority: String,
numViews: Long
)Then it builds a Spark DataFrame from a parallelized sequence of records, deliberately including nulls and a stray URL, because the toy dataset has to violate something for the test to be interesting. The third frame is the verification itself:
val verificationResult = VerificationSuite()
.onData(data)
.addCheck(
Check(CheckLevel.Error, "unit testing my data")
.hasSize(_ == 5)
.isComplete("id")
.isUnique("id")
.hasApproxQuantile("numViews", 0.5, _ <= 10))
.run()The last constraint is the interesting one. `hasApproxQuantile` is an approximate algorithm rather than an exact median, which is the price of computing it across a distributed dataset. Deequ is designed for datasets of billions of rows in a warehouse or distributed filesystem, so approximation is the default posture rather than an exception.
What run() actually does to your Spark cluster
The mechanism is the part that separates this from a validation script. When you call `run`, Deequ translates the test into a series of Spark jobs and executes them to compute metrics on the data. Only after the metrics come back does it invoke your assertion functions, such as `_ == 5` for the size check, to decide whether the constraints hold.
So the shape is compute first, assert second. Your lambda is not evaluated per row; it is evaluated once against an aggregated metric. That is what makes a uniqueness check or a quantile check affordable on a large dataset, and it is also why an assertion like `_ >= 0.5` in `containsURL("description", _ >= 0.5)` reads as a threshold on a computed ratio rather than a per-record predicate.
The return value is a `VerificationResult`, and the README shows how to walk it. You check `verificationResult.status` against `CheckStatus.Success`, then flatten `verificationResult.checkResults` into `constraintResults`, filter on `ConstraintStatus.Success` and print what failed. You get per-constraint output rather than a single boolean, which is what makes the result usable in a CI job that wants to report six failures as six lines.
Two of the sources are worth knowing about by path: `VerificationSuite.scala` and `VerificationResult.scala` under src/main/scala/com/amazon/deequ, and `Check.scala` under the checks package. The `docs/` directory in the tree is where the fuller material lives.
PyDeequ and the limits of the analogy
Python users are pointed at PyDeequ, described in the README as a Python interface for Deequ, with its own GitHub repository, readthedocs pages and a PyPI package. That matters because the constraint vocabulary you learn is portable even when the entry points are not.
The analogy to unit testing has real limits worth stating. A unit test asserts on a value you control and it either passes or fails deterministically. A Deequ check runs against data that can change shape between runs, so a suite that is green this morning can be red tomorrow because a producer changed, not because your code broke. That is the feature, not the bug: the library exists to catch that upstream drift before a consumer does. But it means you are writing assertions about the world, and you need to think about who gets paged when one fails.
Another limit is that Deequ computes quality metrics on Spark. If your data is small enough to fit in a notebook or a local process, the overhead of shipping checks to a cluster may exceed what a handful of assertions in your own code would cost. The library's own framing, very large datasets on distributed storage, is a reasonable guide to where it belongs.
Maintenance signals in the recent release notes
The repository is not archived and the last push was on 2026-09-16, the same day as release 2.1.0. That release is worth reading past the version bump. It replaces deprecated stateful UDAFs with typed aggregators, fixes the KLL rank map and CDF under-counting repeated data points, requires Java 11, updates the Spark compatibility and installation guidance, and backtick-escapes column names in places where Spark's V2 DataSource rejected them.
Release 2.0.21 before it fixed profiling of columns whose names require escaping, such as a column with a leading space. That is a narrow bug with a wide blast radius: column-name escaping is the kind of thing that only breaks when someone finally has a column with a space in it.
The Apache-2.0 licence and the `NOTICE` file in the tree are unremarkable for an AWS Labs project. The thing that should shape your decision is the dependency matrix rather than the licence: a library that must be rebuilt for each Spark version carries a real upgrade cost whenever Spark moves, and that cost lands on whoever owns the pipeline.
Editorial conclusion
deequ makes sense the moment a pipeline feeds something you cannot easily inspect, because it turns implicit assumptions about columns into a result object you can assert on in CI. The important thing to get right before adopting it is the artifact suffix: a Deequ build is compiled against one Spark version, and mixing it with another is the most likely reason a first attempt fails. Version 2.1.0 raised the floor to Java 11 and the release notes for that version include a fix to the KLL rank map and CDF under-counting repeated data points, which is the kind of correctness bug you want to know landed. Start with the VerificationSuite example in the README, then read the checks source under src/main/scala/com/amazon/deequ/checks to see the full constraint vocabulary.
Frequently asked questions
How to use Deequ?
Add the Maven or sbt dependency whose suffix matches your Spark version, build a Spark DataFrame, then pass it to a VerificationSuite with one or more Check objects. Call run, and inspect the returned VerificationResult: its status says whether the suite passed, and its checkResults carry a per-constraint status and message describing what failed.
Which Deequ version should I use with Apache Spark?
Match the artifact suffix to your Spark major and minor version, for example a -spark-3.5 suffix for Spark 3.5.x. Deequ 2.x covers Spark 3.1 through 3.5, Spark 3.0 support lives on the legacy-spark-3.0 branch with Deequ 1.x, and Deequ 2.1.0 or later requires Java 11 while earlier 2.0.x releases require Java 8.
Does Deequ have a Python interface?
Yes. The README points Python users to PyDeequ, a Python interface for Deequ, available on GitHub, on readthedocs and on PyPI. The underlying library is Scala and runs on Apache Spark, so the Python package is a way to reach the same constraints rather than a separate implementation.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/awslabs-deequ)