CLI tool
google/differential-privacy avatar
google/differential-privacy

google/differential-privacy: Google's DP libraries for Go, C++, Java and the JVM

Google's differential privacy libraries.

3,365 stars436 forksGoApache-2.0

At a glance

What is it?
The repository holds end-to-end frameworks (Privacy on Beam, PipelineDP4j) plus low-level building blocks, an accounting library, a stochastic tester and a ZetaSQL CLI. The building blocks do not bound user contributions, so the end-to-end tools are the safer entry point.
Who is it for?
Adopt Privacy on Beam or PipelineDP4j if you already run Apache Beam or Apache Spark and need end-to-end DP aggregation, and use the C++, Go or Java building blocks only when you can enforce a per-user contribution bound upstream yourself.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What google/differential-privacy actually ships

This is not one library. It is a monorepo that packages several distinct pieces under the Apache-2.0 licence: Privacy on Beam, an end-to-end differential privacy framework for Go built on Apache Beam; PipelineDP4j, the JVM counterpart for Java, Kotlin and Scala, which supports Apache Beam and Apache Spark; and three "DP building block" libraries in C++, Go and Java that implement noise addition primitives and differentially private aggregations. Privacy on Beam and PipelineDP4j sit on top of those building blocks. Alongside them the repository carries a stochastic tester under cc/testing, an accounting library under python/dp_accounting for tracking privacy budget, a command line interface under examples/zetasql for running differentially private SQL queries with ZetaSQL, and DP Auditorium under python/dp_auditorium for auditing differential privacy guarantees.

The intended audience splits cleanly. The README says the building blocks, Privacy on Beam, PipelineDP4j and the external PipelineDP are suitable for research, experimental or production use cases, while the other tools are currently experimental and subject to change. That sentence is the most useful line in the document: it tells you which half of the repository you can put behind a real workload and which half you should treat as movable. If you are a Go shop doing aggregation over user data, Privacy on Beam is the target. If you are on the JVM, PipelineDP4j is. If you are writing your own aggregation engine, you want the building blocks, and you take on the responsibilities described below.

The contribution bound the building blocks refuse to enforce

Differential privacy needs a bound on how many rows a single user can contribute to one aggregation. The README is explicit that the DP building block libraries do not perform this bounding. Their implementation assumes each user contributes only a fixed number of rows to each partition, that number is configurable by the user, and the library neither verifies nor enforces it. Pre-processing to enforce it is the caller's job.

The reasoning given is architectural rather than lazy. Enforcing the bound requires a global operation over the data: group by user, then aggregate or subsample that user's contributions before they reach the DP aggregators. Under scalability constraints that step belongs in a distributed data processing framework, which is why Privacy on Beam relies on Apache Beam and PipelineDP4j relies on Apache Beam or Apache Spark. The README therefore recommends the end-to-end tooling where possible.

This is the single most consequential design decision in the repository, and it is easy to get wrong. If you call a building block aggregator directly on raw per-event rows without collapsing them per user first, you have not implemented differential privacy, whatever the API suggests. The library will produce a number and add noise; the guarantee will not hold. The support matrix in the README lists Laplace mechanism, Gaussian mechanism, Count, Sum, Mean, Variance, Quantiles, truncated geometric thresholding, Laplace thresholding, Gaussian thresholding and pre-thresholding as Supported across C++, Go and Java, with automatic bounds approximation marked Supported in C++ and Java but Planned for Go. That gap matters if you are choosing a language on the assumption the three are interchangeable.

Installing and building: Bazelisk, then the per-directory README

There is no package manager step here. The README says the build process assumes you have cloned the Git repository, and that most tools and libraries use Bazel. To get Bazel, you install Bazelisk, which manages Bazel versions and installs the correct one.

bash
git clone https://github.com/google/differential-privacy.git
cd differential-privacy

After cloning, follow the instructions in the directory for the tool you want. The README points out that all tools except the DP building block libraries keep their documentation in their own directories, so Privacy on Beam has its README in privacy-on-beam, and PipelineDP4j has one in pipelinedp4j. The shared documentation for the building blocks is at the top level, with language-specific notes in cc, go and java.

bash
ls privacy-on-beam/README.md pipelinedp4j/README.md
ls cc/ go/ java/

What you should see is a README in each of those directories, plus the top-level MODULE.bazel, WORKSPACE and .bazelversion files that the Bazel build reads. The examples directory holds runnable examples split by language: examples/cc, examples/go, examples/java, examples/pipelinedp4j and examples/zetasql. The README states that the documentation in the tooling directories refers to these examples, so the fastest route to a first real use is to open the example for your language and adapt it rather than to start from the API surface. I would not expect a pip install, a go get or a Maven coordinate to be the documented path; the README's build section is Bazel-first throughout.

Privacy on Beam versus PipelineDP4j: picking the right front end

Both are described as end-to-end differential privacy frameworks intended to be easy to use, even by non-experts. The difference is the runtime you already have. Privacy on Beam is Go and built on Apache Beam. PipelineDP4j targets JVM languages (Java, Kotlin, Scala) and supports different data processing frameworks, specifically Apache Beam and Apache Spark.

That Spark support is the practical dividing line. If your aggregation jobs already run as Spark pipelines and you cannot justify standing up a Beam pipeline purely for the privacy layer, PipelineDP4j is the only one of the two that meets you where you are. If your stack is Go, Privacy on Beam is the natural fit and the README's own recommendation for that case. Choosing the other one means either a language migration or a second processing framework in your infrastructure, and neither is a small decision.

There is also a Python option, but it is not in this repository. The README notes that PipelineDP is an end-to-end framework for Python, is the Python version of PipelineDP4j, is a collaboration between Google and OpenMined, and lives in the OpenMined repository. PyDP, also from OpenMined, is a Python wrapper of the C++ building block library. If your team writes Python, the honest answer is that you will be depending on OpenMined's projects, and the Google repository is upstream of them rather than the thing you install.

Where this is the wrong tool

The repository is not a privacy layer you can drop in front of an arbitrary query engine. The building blocks cover a fixed set of aggregations: Laplace and Gaussian mechanisms, count, sum, mean, variance, quantiles, and several thresholding variants. Anything outside that list means composing the mechanisms yourself, and the README points out that the mechanism implementations use secure noise generation, with a linked PDF in common_docs. It also warns that if you plan to use the low-level libraries you may need a more in-depth understanding of differential privacy than the friendly introduction the README recommends for newcomers.

The second boundary is the experimental set. The ZetaSQL CLI, the stochastic tester and DP Auditorium are not described the way the production-suitable components are. The README says they are currently experimental and subject to change. Building a reporting product on the ZetaSQL CLI means accepting that its interface can move. The stochastic tester is a development aid for catching regressions that would break the differential privacy property, not a runtime guard.

The third boundary is organisational rather than technical. Nothing here decides what your privacy budget should be or how you account for it across queries. The accounting library under python/dp_accounting tracks budget, but the policy is yours. If nobody on the team can state what epsilon your workload is allowed to spend, adopting this repository will produce noisy numbers without producing a defensible privacy position.

Alternatives and the real difference in approach

The most direct alternative for Python teams is PipelineDP, which the README describes as the Python version of PipelineDP4j and a collaboration between Google and OpenMined, hosted in the OpenMined repository. The difference is not just language. PipelineDP4j is maintained inside this monorepo alongside the building blocks it depends on, while PipelineDP is a separate project that builds on the same ideas. If your data processing is Python and your team does not want to run Beam or Spark jobs written in Go or a JVM language, that split is the deciding factor.

PyDP is a different kind of alternative: a Python wrapper of the C++ DP building block library. Choosing it means choosing the building block level of abstraction, with everything that entails. You get the primitives in Python, and you take on the contribution-bounding problem yourself, exactly as you would with the C++ library directly. It is not a substitute for an end-to-end framework, and the README does not present it as one.

Within the repository, the choice is between the front ends and the building blocks, and the README's own recommendation is to prefer the end-to-end tooling. That recommendation is the strongest signal available about how the maintainers expect the pieces to be used.

Maintenance, licence and what upgrades cost

The repository is not archived, and the last push was on 2026-09-21, three days before this writing. Releases are infrequent and deliberate: v4.1.0 on 2026-02-06, v4.0.0 on 2025-05-08, v3.0.0 on 2024-03-12. The gaps are roughly nine months and fourteen months. A major version bump every year or so is a reasonable cadence for a library whose correctness argument is mathematical, but it means you should plan for migration work rather than expect continuous small updates.

The licence is Apache-2.0, which is permissive and includes an explicit patent grant. That is a statement about the licence text, not legal advice; if you are embedding these libraries in a distributed product, have your own counsel review the notice and attribution requirements.

Upgrade cost is dominated by the Bazel toolchain rather than the API. The repository pins a Bazel version through .bazelversion and uses Bazelisk to select it, so a version bump can ripple into your build environment. The README's caveat section also means that any change to how you bound per-user contributions is a change to your privacy argument, not just to your code, and that review is the expensive part of an upgrade here.

Editorial conclusion

Adopt Privacy on Beam or PipelineDP4j if you already run Apache Beam or Apache Spark and need end-to-end DP aggregation, and use the C++, Go or Java building blocks only when you can enforce a per-user contribution bound upstream yourself. Do not reach for the building blocks as a first step, because the README states they neither verify nor enforce that limit, and do not treat the ZetaSQL CLI, the stochastic tester or DP Auditorium as production components, since the README calls those experimental. Before committing, check the per-language README for the actual build target and confirm which algorithms your language supports, because the support matrix marks automatic bounds approximation as Planned for Go.

Frequently asked questions

What is differential privacy in the context of google/differential-privacy?

The repository describes its purpose as generating epsilon- and (epsilon, delta)-differentially private statistics over datasets. It provides end-to-end frameworks in Privacy on Beam and PipelineDP4j, plus lower-level building blocks in C++, Go and Java that implement noise addition primitives and differentially private aggregations.

What are the limitations of the DP building blocks in google/differential-privacy?

The README states that the building block libraries do not bound the maximum number of contributions each user can make to a single aggregation. They assume a fixed, user-configurable number of rows per user per partition and neither verify nor enforce it, leaving pre-processing to the caller.

Which differential privacy algorithms does google/differential-privacy support?

The README lists Laplace mechanism, Gaussian mechanism, Count, Sum, Mean, Variance, Quantiles, truncated geometric thresholding, Laplace thresholding, Gaussian thresholding and pre-thresholding as supported across C++, Go and Java. Automatic bounds approximation is listed as Supported for C++ and Java but Planned for Go.

Official sources

  1. google/differential-privacy on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/google-differential-privacy.svg)](https://hysenlabs.com/projects/google-differential-privacy)