# Opacus: DP-SGD Training for PyTorch Models

> Opacus wraps an existing PyTorch training loop in a PrivacyEngine and reports the privacy budget as you train. It is a good fit for practitioners who already have a working model and need per-sample gradient clipping, and a poor fit for anyone expecting DP to be free.

**meta-pytorch/opacus** — Training PyTorch models with differential privacy

- Repository: https://github.com/meta-pytorch/opacus
- Website: https://opacus.ai
- Stars: 1,962 · Forks: 398
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/meta-pytorch-opacus

## The problem Opacus solves for PyTorch teams

Standard SGD updates a model from the average gradient over a batch. A single training example can, in principle, shift that average in a way that is detectable in the released weights, which is the mechanism behind membership inference attacks. Differential privacy replaces that with a guarantee about the distribution of outputs: the README describes Opacus as a library that "enables training PyTorch models with differential privacy", and the target audience section names two groups, ML practitioners who want a gentle introduction with minimal code changes, and differential privacy researchers who want to experiment.

The second group matters for how you read the project. Opacus is not a product that hides DP behind a switch. The privacy guarantee comes from DP-SGD, which clips each per-sample gradient to a fixed norm and adds calibrated Gaussian noise, and the parameters that control both are exposed directly to you. The PrivacyEngine tracks the budget spent so far, so you can stop training when epsilon crosses a threshold you chose in advance. If you cannot state that threshold, the library will still run, but you have no basis for calling the result private.

## How the PrivacyEngine rewrites your training step

The mechanism is a wrapper, not a fork of the optimizer. You construct a model, an optimizer and a DataLoader as usual, then hand all three to PrivacyEngine.make_private(), which returns private counterparts. According to the README, the engine needs a module, an optimizer, a data_loader, a noise_multiplier and a max_grad_norm. After that call, the training loop is unchanged: you still call optimizer.zero_grad(), loss.backward() and optimizer.step().

What changed underneath is the gradient computation. DP-SGD needs a gradient per example rather than one averaged gradient per batch, because clipping is applied before averaging. Opacus provides that through grad samplers, and the project ships a guide to grad samplers as one of its tutorials, which tells you the per-layer machinery is configurable rather than fixed. The README also documents two optimizations added in August 2024, Fast Gradient Clipping and Ghost Clipping, described as significantly reducing the memory requirements of DP-SGD. Those matter because the naive per-sample path multiplies activation memory by batch size. The README points to a PyTorch blog post on clipping in Opacus for detail rather than explaining the algorithms inline, so the repository is the place to look if you need to know which sampler your architecture will select. The tutorials folder also includes a guide to the module validator, which is the component that inspects your model and reports layers it cannot handle. That validator is the first thing that will fail on an unusual architecture.

## Installing Opacus and running a first private step

The README gives two package manager routes and a source route. The pip form is the shortest. Run it and you should end up with the opacus package importable alongside your existing torch install; requirements.txt pins torch>=2.6.0, numpy>=1.15, scipy>=1.2 and opt-einsum>=3.3.0, so pip will pull those if they are missing.

```bash
pip install opacus
```

The conda route is the alternative the README lists, using the conda-forge channel. Use it when your environment is already conda-managed and you do not want pip to resolve torch again.

```bash
conda install -c conda-forge opacus
```

For unreleased features the README suggests a source install, with the caveat that you get quirks and potentially occasional bugs. Note that setup.py requires Python 3.7.5 or newer and exits with an error message otherwise.

```bash
git clone https://github.com/pytorch/opacus.git
cd opacus
pip install -e .
```

The first real use is the make_private() call itself. The README's example defines a model, an SGD optimizer and a DataLoader with batch_size=1024, then wraps them. The two privacy parameters are noise_multiplier=1.1 and max_grad_norm=1.0. Higher noise buys more privacy and costs accuracy; max_grad_norm sets the clipping threshold.

```python
privacy_engine = PrivacyEngine()
model, optimizer, data_loader = privacy_engine.make_private(
    module=model,
    optimizer=optimizer,
    data_loader=data_loader,
    noise_multiplier=1.1,
    max_grad_norm=1.0,
)
```

After this, the README says it is business as usual. For an end-to-end run, examples/mnist.py is the reference the README names, and the examples folder holds more, including cifar10.py, imdb.py, dcgan.py and a Lightning variant. The tutorials folder has a text classifier notebook built on BERT that the release notes say was updated in December 2024 to show LoRA and the peft library used together with DP-SGD, which is the realistic path if you want to fine-tune a large model without the per-sample gradient cost of full fine-tuning.

## Where DP-SGD breaks your accuracy budget

The honest limitation is that differential privacy is not free, and Opacus does not pretend otherwise in its interface. You choose noise_multiplier, and the noise is added to the clipped gradient sum every step. At a fixed epsilon, more training steps means more noise, so a longer schedule is not automatically a better model. The README does not publish an accuracy table, so the only way to know your trade-off is to run it: pick a target epsilon, train, and compare against the non-private baseline on the same data.

The second limitation is architectural. Per-sample gradients require Opacus to know how to compute them for each layer, which is why the module validator exists and why the project ships a guide to fixing modules it rejects. Custom layers, some attention variants, and anything with a backward pass that does not decompose cleanly per example will need either a registered grad sampler or a rewrite. The third is memory. Fast Gradient Clipping and Ghost Clipping reduce the footprint, but the README describes them as reducing memory requirements rather than eliminating the overhead, and the naive path scales with batch size.

There is also a category error worth naming. Opacus protects the training data. It says nothing about what a trained model leaks at inference time, and it does not add noise to predictions. If your threat model is a user querying a deployed API to reconstruct training examples, DP-SGD during training is the relevant control, but a deployment-time defence is a different piece of work.

## Opacus compared with TensorFlow Privacy and manual DP-SGD

The closest alternative in the deep learning ecosystem is TensorFlow Privacy, which implements the same DP-SGD idea inside the TensorFlow/Keras training stack. The difference is not the algorithm, it is the integration surface. TensorFlow Privacy expects a Keras optimizer and its own privacy accounting, so choosing between them is mostly a question of which framework your model already lives in. If you have a PyTorch training loop, porting to Keras to get DP is a large change; if you have a Keras model, Opacus does not help you.

The other alternative is implementing DP-SGD yourself: per-sample gradient computation, clipping, noise injection, and an accountant for the privacy budget. The appeal is control over the accountant and the sampler. The cost is that the accountant is where mistakes are quiet and expensive, and the tutorials list shows Opacus has invested specifically in the pieces that are tedious to get right, including grad samplers and non-wrapping mode. The README also points to a technical report on arXiv describing the design principles, mathematical foundations and benchmarks, which is the document to read if you want to judge whether the library's accounting matches what you need. For researchers who want to modify the mechanism, the README explicitly names them as a target audience, so forking or vendoring is an expected use rather than an abuse.

## Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-07-13. The most recent release listed is v1.6.0 on 2026-05-05, following v1.5.4 on 2025-05-27 and v1.5.3 on 2025-02-18. The gap between 1.5.4 and 1.6.0 is roughly a year, so plan for a slow release cadence rather than a monthly one. The repository does include a Migration_Guide.md at the top level, which is the file to read before jumping a minor version, and a CHANGELOG.md for the itemized list.

The licence is Apache-2.0, and the README states it plainly. Apache-2.0 is permissive and includes an explicit patent grant, which matters for a library that may end up inside a commercial training pipeline. It also means you can vendor the code if you need to patch a grad sampler yourself. This is not legal advice; if your organisation has a policy on patent clauses or attribution, read the LICENSE file with whoever handles that.

The practical upgrade cost is tied to torch. requirements.txt requires torch>=2.6.0, so an Opacus upgrade can force a torch upgrade, and torch upgrades are where training loops break. Budget for running the examples folder against your own model after any bump, not just your test suite.

## Conclusion

Adopt Opacus if you already have a working PyTorch training loop and a concrete reason to bound what a single training example can reveal, and you can afford the accuracy and throughput cost of DP-SGD. Do not adopt it if you only need to protect a deployed model from extraction queries, or if your model relies on layers the module validator cannot account for and you are not prepared to rewrite them. Before committing, verify three things on your own data: that make_private() accepts your model without unsupported-module errors, what epsilon you actually reach at your target accuracy, and whether the per-sample gradient path fits in your GPU memory at your batch size.

## FAQ

### What is Opacus?

Opacus is a library from Meta that enables training PyTorch models with differential privacy. It works by wrapping your model, optimizer and data loader in a PrivacyEngine, and it lets you track the privacy budget spent during training.

### What does Opacus mean?

The README does not explain the origin of the name. It only describes what the library does, so the naming rationale is not documented there.

### What is the meaning of Opacus?

The README does not define the word itself. It describes Opacus as a library for training PyTorch models with differential privacy, aimed at ML practitioners and differential privacy researchers.

### What are Opacus clouds?

The README and the repository do not describe any cloud offering under that name. Opacus is distributed as a Python package through pip and conda, and the source is hosted on GitHub.

### What does Opacus do?

It enables training PyTorch models with differential privacy by wrapping your model, optimizer and data loader in a PrivacyEngine and tracking the privacy budget expended as training proceeds.

## Sources

- [License: Apache-2.0](https://github.com/meta-pytorch/opacus/blob/main/LICENSE)
- [meta-pytorch/opacus on GitHub](https://github.com/meta-pytorch/opacus)
- [Project website](https://opacus.ai)
- [README](https://github.com/meta-pytorch/opacus/blob/main/README.md)
- [Releases](https://github.com/meta-pytorch/opacus/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/meta-pytorch-opacus
