# AIJack: a C++ and Python simulator for machine learning attacks and defences

> AIJack packages adversarial attacks, differential privacy, homomorphic encryption and federated learning into one Python library with a C++ backend. It is a research simulator, not a production hardening tool, and the documentation is thinner than the feature list.

**Koukyosyumei/AIJack** — Security and Privacy Risk Simulator for Machine Learning (arXiv:2312.17667)

- Repository: https://github.com/Koukyosyumei/AIJack
- Website: https://koukyosyumei.github.io/AIJack/
- Stars: 429 · Forks: 66
- Language: C++
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/koukyosyumei-aijack

## What AIJack is for, and who it is not for

AIJack targets a specific kind of reader: someone who wants to run a poisoning attack, a model inversion attack, a membership inference attack or a backdoor against a model, and who does not want to write that attack from scratch. The README frames the project as an "easy-to-use open-source simulation tool for testing the security of your AI system against hijackers" and claims support for more than 30 methods. The paper reference is arXiv:2312.17667, so the intended audience is at least partly academic.

The same library also carries the defensive side: differential privacy, homomorphic encryption, k-anonymity and federated learning. That combination is the actual pitch. You can build a federated learning setup, attach a server-side gradient inversion manager to it, and observe what the server can recover from the gradients. That is a simulation, not a deployment.

The word simulator matters. Nothing in the README describes hardening a live service, monitoring traffic, or integrating with an existing ML platform. If your goal is a model that survives contact with real adversaries in production, this is not the tool. If your goal is to understand how much a given attack recovers under controlled conditions, it is aimed at you.

## The Client, Server, API and Manager model behind federated scenarios

For distributed learning, AIJack exposes four abstractions: `Client`, `Server`, `API` and `Manager`. Clients and servers hold the participants in a scheme. The `API` object wires them together and runs training when you call `run()`. The `Manager` is the interesting piece: it wraps an existing class through an `attach` method and returns an extended class with extra behaviour.

That indirection is what lets the library express an attack without rewriting the training loop. The README shows a federated averaging setup where a server-side manager turns a normal `FedAVGServer` into an attacking server. The manager is constructed with an input shape, attaches to the server class, and the resulting class is used in place of the original when the API is built. The training call does not change.

The design has a cost. Because the attack lives in a manager that mutates a class, the failure mode when something goes wrong is a stack trace inside generated code rather than inside a readable subclass. It also means the API surface is not obvious from the class you instantiate. You have to know which manager exists for which attack, and the README points to the API reference for that rather than listing them inline.

## Installing AIJack and running a first attack

The README gives a pip path and states two prerequisites: Boost and pybind11. The build configuration in pyproject.toml requires setuptools, wheel, ninja and cmake, and setup.py drives a CMake build of a native extension, which is consistent with the C++ backend the topics list mentions.

Install the system dependency first, then pybind11, then the package itself:

```bash
apt install -y libboost-all-dev
pip install -U pip
pip install "pybind11[global]"
pip install aijack
```

If you want the current state of the main branch rather than the release on PyPI, the README offers a direct install from GitHub:

```bash
pip install git+https://github.com/Koukyosyumei/AIJack
```

A Dockerfile is also provided in the repository root. It is short: a slim Python base image, the same Boost install, the same pybind11 install, and a pip install from GitHub. Note that the Dockerfile pins `python:3.13.0a2-slim`, an alpha Python image, so treat that file as a starting point rather than a maintained artefact.

For a first real use, the README gives a poisoning attack against an sklearn SVM. The attacker is constructed with the classifier and the training data, and `attack` returns both the poisoned data and a log:

```python
from aijack . attack import Poison_attack_sklearn

attacker = Poison_attack_sklearn(clf , X_train , y_train)
malicious_data , log = attacker.attack(initial_data , 1, X_valid , y_valid)
```

After this call you have `malicious_data`, which you can feed to a fresh fit, and `log`, which records the optimisation. The README does not document the meaning of the positional arguments beyond this example, so read the source before changing them.

## AIValut and the SQL debugging layer

AIValut is a separate component inside the repository, described as a simple DBMS for SQL-based algorithms. It has its own storage engine and query parser, and the README states that it currently supports Rain, a SQL-based debugging system for ML models, with k-anonymity, homomorphic encryption and differential privacy listed as future work.

The worked example trains logistic regression with a `Logreg` statement and then repairs it with a `Complaint` statement. In the example, the model is trained on a small bankruptcy table, produces an AUC of 0.520000, and then a `Complaint` query removes one record so that samples with `debt` greater than or equal to 100 are classified positive, after which the reported AUC is 1.000000. The fixed parameters are printed alongside the original ones, and predictions are stored in tables named after the model and the complaint.

That AUC jump is the point of the example, not a performance claim about the library. It shows a single removed row flipping the decision boundary, which is exactly the kind of fragility the tool exists to expose. The README carries a blunt warning above this section: use AIValut only for research purposes. The top-level README says nothing about persistence guarantees, concurrency or transaction semantics, so do not treat the storage engine as a database you can put behind a service.

## Where AIJack gets in the way

The install path is the first real obstacle. A pure `pip install aijack` is not self-contained: it needs Boost headers on the system and pybind11 installed beforehand, and the build goes through CMake and ninja. On a machine without a compiler toolchain, or inside a minimal container that lacks Boost, the install fails before any AIJack code runs. The Dockerfile exists precisely because of this, which tells you the maintainers know the plain pip path is not always enough.

The second obstacle is documentation depth. The README shows one concrete attack, one federated scenario and one SQL session. It states that more than 30 methods are supported and points to the API reference for the rest, but it does not enumerate them in the text available here. When you move past the examples, you are reading source.

The third is scope. AIJack is a simulator. Its federated learning backends and its privacy primitives are there so you can measure an attack or a defence under controlled conditions. Nothing in the README describes deployment, key management for the Paillier-based encryption, or integration with an existing training stack. If you need any of those, this is the wrong layer.

## How AIJack differs from a privacy library like Opacus

Opacus is the obvious comparison for the differential privacy half of AIJack, and the difference is one of scope rather than quality. Opacus exists to make a PyTorch training loop differentially private: you wrap the optimizer and the data loader, and the library handles per-sample gradient computation and clipping. It does one thing and it is built to sit inside a real training run.

AIJack covers differential privacy as one entry in a much wider catalogue that also includes poisoning, model inversion, backdoor and free-rider attacks, homomorphic encryption and k-anonymity. The `Manager` and `attach` pattern is designed to let you bolt an attack onto a training scheme, which is a different job from making a training run private by default.

If your question is "how do I train this model with a privacy budget", Opacus is the narrower and more direct answer. If your question is "how much does this attack recover, and does this defence reduce it", AIJack is built for that comparison. The two are not substitutes, and the README does not position AIJack as a replacement for a privacy training library.

## Maintenance, licensing and what an upgrade costs

The repository is not archived, and the last push was on 2026-05-30. The most recent release listed is v0.0.1-beta.2 from 2024-01-01, preceded by v0.0.1-beta.1 and v0.0.1-alpha.2. The version string in pyproject.toml matches v0.0.1-beta.2. A beta version with no 1.0 release is the relevant fact here: the public API has not been declared stable, so pinning a version and reading the diff before upgrading is the practical stance.

The licence is Apache-2.0, declared in the LICENSE file and referenced from pyproject.toml. Apache-2.0 is permissive and includes a patent grant, which matters for a library that implements cryptographic primitives, but this is not legal advice and the patent and notice clauses deserve a read if you redistribute.

One dependency detail is worth flagging for upgrade planning. The declared dependencies include `torch >= 1.11.0`, `mpi4py`, `pybind11`, `opencv-python` and `statsmodels`. `mpi4py` in particular implies an MPI runtime for the federated learning backend, which is a system-level dependency that pip will not install for you. Upgrading AIJack in an environment without MPI will not give you the federated paths.

## Conclusion

AIJack suits researchers and graduate students who need to reproduce or compare attack and defence methods in Python without reimplementing them, and who can read the source when the documentation stops. It is the wrong choice if you need a supported, versioned product for a production pipeline, or if your team cannot build a C++ extension from source. Before adopting it, check the pip install path on your target platform, confirm that the specific attack or defence you need appears in the supported algorithm table, and read the aivalut/README.md if you intend to use the SQL layer, since the top-level README marks AIValut as research-only.

## FAQ

### How do I install AIJack?

Install Boost and pybind11 first, then run pip install aijack. The README gives the sequence as apt install -y libboost-all-dev, pip install -U pip, pip install "pybind11[global]", then pip install aijack. A Dockerfile is also provided in the repository.

### Does AIJack work with PyTorch and scikit-learn models?

The README describes the design as PyTorch-friendly and compatible with scikit-learn, and states that AIJack mainly supports PyTorch or sklearn models. The poisoning example uses an sklearn classifier, and the federated examples use PyTorch-style model objects.

### Can I use AIJack in production?

The README presents AIJack as a simulation tool for testing an AI system against attacks, and the AIValut section carries an explicit instruction to use it only for research purposes. Nothing in the README describes deployment, monitoring or integration with a production training pipeline.

## Sources

- [Koukyosyumei/AIJack on GitHub](https://github.com/Koukyosyumei/AIJack)
- [License: Apache-2.0](https://github.com/Koukyosyumei/AIJack/blob/main/LICENSE)
- [Project website](https://koukyosyumei.github.io/AIJack/)
- [README](https://github.com/Koukyosyumei/AIJack/blob/main/README.md)
- [Releases](https://github.com/Koukyosyumei/AIJack/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/koukyosyumei-aijack
