TensorFlow Privacy: DP-SGD for Keras without microbatches
Library for training machine learning models with privacy for training data
At a glance
- What is it?
- TensorFlow Privacy wraps standard Keras optimizers in differentially private versions and ships the accounting tools to measure the guarantee. The catch is a pinned TensorFlow ceiling and a Python range that stops before 3.12.
- Who is it for?
- Adopt it if you train Dense and Embedding networks in Keras and need a documented epsilon rather than a promise. Skip it if your architecture needs per-example clipping on convolutions, or if you are already on Python 3.12 or TensorFlow 2.16, since setup.py caps both.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 34 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
The problem TensorFlow Privacy addresses
A model trained on user records can memorize them. TensorFlow Privacy exists so that the training step itself carries a formal guarantee, rather than relying on post-hoc claims about what the network did or did not retain. The library provides TensorFlow optimizer implementations for training machine learning models with differential privacy, plus analysis tools for computing the privacy guarantees those runs provide. That pairing is the point: an optimizer that adds noise is not useful on its own, because you need to know how much privacy the noise bought you.
The audience is narrow but real. You are already writing TensorFlow and Keras code. You have a dataset of records that belong to individuals. You need to state a number, epsilon, and defend it. If any of those three is false, this library is not aimed at you. It is not a general anonymization toolkit, it does not scrub identifiers from CSV files, and it does not produce synthetic data.
How DP-SGD is implemented in the optimizers
The mechanism is differentially private stochastic gradient descent. Instead of averaging gradients over a batch, the optimizer clips each example's gradient to a fixed norm and then adds calibrated noise before applying the update. The wrapping is done at the optimizer level, so existing code keeps its model definition and swaps the optimizer class.
Version 0.9.0 split the repository into two PyPI packages. The first keeps the name tensorflow-privacy and contains the training side. The second, tensorflow-empirical-privacy, contains the parts related to testing for empirical privacy. If you only want to train, you install the first and never touch the second.
The interesting engineering is in the Keras path. A 2023-02-21 note in the README describes a new implementation of efficient per-example gradient clipping for DP Keras models consisting only of Dense and Embedding layers, using fast gradient calculation results from arXiv 1510.01799. The README states the implementation should allow DP training without meaningful memory or runtime overhead, and that it removes the need to tune the number of microbatches because it clips with respect to each example. That last part matters more than the performance claim. Microbatch tuning is the step where most DP training setups go wrong, and eliminating it removes a whole class of silent mistakes. The constraint is right there in the sentence: Dense and Embedding layers only.
Installing tensorflow-privacy and a first DP training run
The README gives two paths. If you only want the library, install it from PyPI. TensorFlow itself is a prerequisite, version 1.14 or later according to the README, and GPU support is recommended for better performance.
pip install tensorflow-privacyIf you intend to read or modify the source, clone the repository and install the local package in editable mode so it lands on your PYTHONPATH. The README recommends forking first if you plan to contribute.
git clone https://github.com/tensorflow/privacy
cd privacy
pip install -e .For TensorFlow 2, the README points at the Keras-based estimators in tensorflow_privacy/privacy/optimizers/dp_optimizer_keras.py, and notes that using them with tf.keras.Model and tf.estimator.Estimator requires TensorFlow 2.4 or later. The tutorials directory holds a walkthrough that teaches wrapping existing optimizers such as SGD or Adam into their differentially private counterparts, tuning the parameters that DP optimization introduces, and measuring the resulting guarantee with the analysis tools. The README is explicit that the tutorials are not part of the API and can change without warning, so do not import them from production code.
Version pins that decide your environment before you start
setup.py is the most honest document in the repository. It declares python_requires '>=3.9.0,<3.12', which rules out Python 3.12 and later. It pins tensorflow to '>=2.4.0,<=2.15.0'. The ceiling is not a suggestion. If your platform has moved to a newer TensorFlow, you are choosing between the privacy library and that upgrade.
Other pins constrain the surrounding environment: numpy~=1.21, scipy~=1.9, scikit-learn>=1.0,==1.*, tensorflow-probability~=0.22.0, dm-tree==0.1.8, and dp-accounting==0.4.4, the last carrying a TODO comment in the source. The requirements.txt file adds tooling that the install does not pull in, including pandas, matplotlib, statsmodels==0.14.0, tensorflow-datasets, tensorflow-estimator, immutabledict and tf-models-official. Those are for the tutorials and research code, not the library.
The practical consequence is that TensorFlow Privacy wants its own virtual environment. Dropping it into a shared environment that already has a newer numpy or scipy will produce resolver conflicts, and the pinned dp-accounting version is what the privacy calculations were tested against.
Where the Dense-and-Embedding restriction bites
The fast per-example clipping path covers models made of Dense and Embedding layers. A convolutional image classifier is not that. Neither is a transformer with attention blocks built from custom layers. For those architectures you fall back to the general DP optimizer path, which means microbatches, and microbatches mean another parameter to tune and another way to get a worse-than-expected result.
This is the honest limitation of the library and it is stated in the README rather than buried. The efficient implementation is scoped to a layer set, and the scope is narrow.
A second limitation is structural. The research directory is described in the README as not maintained as carefully as the tutorials directory and intended as a convenient archive. If you find a promising technique there, treat it as a paper reproduction you may have to repair, not as a supported feature.
A third is that the library gives you a mechanism and an accountant, not a decision. Nothing in the repository tells you what epsilon is acceptable for your data. That judgement depends on your threat model and your obligations, and the library will happily report a guarantee for a value that satisfies nobody.
TensorFlow Privacy compared with Opacus and TensorFlow Federated
The closest alternative in the PyTorch world is Opacus, which implements the same DP-SGD idea through a PrivacyEngine that attaches to an existing optimizer and a GradSampleModule that wraps the model to capture per-sample gradients. The difference in approach is where the wrapping happens. Opacus wraps the model and the optimizer from the outside, so it can support a broader range of layer types through its grad sampler, at the cost of an extra module layer in your code. TensorFlow Privacy puts the privacy inside the optimizer class, which fits Keras idioms more naturally and is why the fast path is limited to layer types the optimizer knows how to handle.
TensorFlow Federated is a different comparison, and the requirements.txt comments make the relationship concrete: they instruct maintainers to keep dependency versions in setup.py in sync with what TensorFlow Federated requires. If your privacy work happens in a federated setting, the dependency alignment is already part of the maintenance burden here. If it does not, that coupling is just one more thing that can force a version bump.
Maintenance, licensing and the upgrade bill
The repository is not archived and the last push was on 2026-08-26, so the codebase is receiving changes. The most recent release listed is v0.9.0 from 2024-02-14, which introduced the two-package split, followed by v0.8.12 and v0.8.11 in October 2023. The gap between the last push and the last release is worth noting if you depend on tagged versions rather than master.
The upgrade cost is dominated by the TensorFlow ceiling. Every time TensorFlow moves, this library's setup.py has to move with it, and until it does you are pinned. The Python range has the same shape: '<3.12' means an interpreter upgrade on your side is blocked by a dependency you did not choose for its interpreter support. Budget for that. The tutorials are explicitly outside the API, so any code you copy from them is yours to maintain.
Licensing is Apache-2.0, stated in both the LICENSE file and setup.py, with copyright 2019 Google LLC. Contributions require signing the Google CLA, and the project asks for PEP8 with two spaces, checked with autopep8 and pylint against TensorFlow's configuration file. The README also states that pull requests adding git submodules are not accepted. None of this is legal advice; if you are embedding the library in a product, read the licence text and your own obligations.
Editorial conclusion
Adopt it if you train Dense and Embedding networks in Keras and need a documented epsilon rather than a promise. Skip it if your architecture needs per-example clipping on convolutions, or if you are already on Python 3.12 or TensorFlow 2.16, since setup.py caps both. Verify your Python version against the >=3.9.0,<3.12 floor and ceiling before you plan the work.
Frequently asked questions
What Python and TensorFlow versions does TensorFlow Privacy require?
setup.py declares python_requires '>=3.9.0,<3.12' and pins tensorflow to '>=2.4.0,<=2.15.0'. The README separately notes that using the Keras estimators with tf.keras.Model and tf.estimator.Estimator needs TensorFlow 2.4 or later.
How do I install TensorFlow Privacy?
If you only want the library, run pip install tensorflow-privacy. To work from source, clone the repository, change into the privacy directory and run pip install -e . to install it in editable mode.
What changed in TensorFlow Privacy 0.9.0?
According to the README's 2024-02-14 note, version 0.9.0 split the repository into two PyPI packages: tensorflow-privacy for DP model training, and tensorflow-empirical-privacy for testing empirical privacy.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/tensorflow-privacy)
Community notes