# AIF360: IBM's Fairness Toolkit for Python and R Models

> AIF360 bundles group fairness metrics, metric explainers and bias mitigation algorithms behind one Python and R interface. It is a research library with a research library's install costs, and much of its value depends on which extra you install.

**Trusted-AI/AIF360** — A comprehensive set of fairness metrics for datasets and machine learning models, explanations for these metrics, and algorithms to mitigate bias in datasets and models.

- Repository: https://github.com/Trusted-AI/AIF360
- Website: https://aif360.res.ibm.com/
- Stars: 2,872 · Forks: 913
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/trusted-ai-aif360

## What AIF360 actually solves, and who ends up using it

The problem AIF360 addresses is not building a fair model. It is answering the question that comes after a model already exists: is this model treating groups differently, by how much, and what happens if I change it? The README frames the package as three things in one: metrics for datasets and models, explanations for those metrics, and algorithms that mitigate bias in datasets and models. That bundling is the point. A team can measure disparate impact, then run a preprocessing or postprocessing algorithm, then measure again on the same objects, without rewriting the evaluation harness.

The intended audience is stated plainly: the project says it is designed to translate algorithmic research from the lab into practice in domains such as finance, human capital management, healthcare and education. In practice that means data scientists and ML engineers, not legal or compliance staff. The API is Python-first, with an R package published on CRAN as a second front end. The repository also carries an mlops/ directory and a demo_FACTS.ipynb example, which suggests some attention to pipeline integration, though the README does not describe a serving story.

One honest detail sits in the README itself: because the toolkit is a comprehensive set of capabilities, the authors acknowledge it may be confusing to figure out which metrics and algorithms fit a given use case, and they point to separate guidance material. That is a real admission. A library that ships more than a dozen mitigation algorithms has not solved the selection problem for you.

## Metrics, explainers and mitigation algorithms in one object model

The mechanism is built around a small number of shared abstractions. Datasets and models are wrapped so that a protected attribute, a privileged group and a label are declared once, and every metric or algorithm reads from that declaration. That is why the same code path can compute a metric before and after a mitigation step: both sides speak the same vocabulary.

The metric side is grouped in the README into several families. There are group fairness metrics derived from selection rates and error rates, including rich subgroup fairness; sample distortion metrics; the Generalized Entropy Index; Differential Fairness and Bias Amplification; and Bias Scan with Multi-Dimensional Subset Scan. The split matters. Selection-rate metrics such as disparate impact are cheap and interpretable but blind to error-rate differences. Error-rate metrics such as equalized odds capture those differences but require labels you may not have at prediction time. Bias Scan is a different kind of tool: it searches for subsets of the data where bias is concentrated, rather than reporting one number for the whole population.

The mitigation side is organized by where the intervention happens. Preprocessing algorithms transform the training data: Optimized Preprocessing, Disparate Impact Remover, Reweighing, Learning Fair Representations, Fair Data Adaptation. Inprocessing algorithms change the training procedure itself: Prejudice Remover Regularizer, Adversarial Debiasing, the Meta-Algorithm for Fair Classification, Exponentiated Gradient Reduction, Grid Search Reduction, Sensitive Set Invariance. Postprocessing algorithms adjust predictions after the model is trained: Equalized Odds Postprocessing, Calibrated Equalized Odds Postprocessing, Reject Option Classification. The README lists each with its source paper, which is the most useful part of that table: it tells you what guarantees the method was designed to provide, and therefore what it does not promise.

## Installing AIF360 with pip and running a first reweighing check

The README recommends a virtual environment before anything else, on the grounds that AIF360 requires specific versions of many Python packages that may conflict with other projects. Conda is the recommended manager, though the README says Virtualenv is generally interchangeable. The supported Python versions are 3.10 through 3.13 on macOS, Ubuntu and Windows.

Create the environment and activate it:

```bash
conda create --name aif360 python=3.11
conda activate aif360
```

The shell prompt should change to show (aif360). Then install the base package from PyPI:

```bash
pip install aif360
```

At this point every metric in the README works. The algorithms do not. Each algorithm family with non-trivial dependencies is gated behind an extra, and the README gives this example:

```bash
pip install 'aif360[LFR,OptimPreproc]'
```

The full list of extras is OptimPreproc, LFR, AdversarialDebiasing, DisparateImpactRemover, LIME, ART, Reductions, FairAdapt, inFairness, LawSchoolGPA, notebooks, tests, docs and all. Installing 'aif360[all]' pulls in everything, which is convenient for exploration and heavy for a production image. For an editable install from a clone, the README gives:

```bash
git clone https://github.com/Trusted-AI/AIF360
pip install --editable '.[all]'
```

The example notebooks live in examples/ and are meant to be run after a manual install, with datasets placed according to aif360/data/README.md. The R route is a single call from CRAN:

```r
install.packages("aif360")
```

For a first real use, the Reweighing notebook (examples/demo_reweighing_preproc.ipynb) is the shortest path: it is a preprocessing method, so it needs no TensorFlow or torch, and it produces sample weights you can inspect directly. The tutorial_credit_scoring.ipynb notebook is the more realistic end-to-end example if you want to see metrics and mitigation composed on a lending dataset.

## Where AIF360 gets in the way: extras, versions and missing attributes

The dependency surface is the first practical limitation. The extras map to genuinely heavy packages: AdversarialDebiasing pulls tensorflow, LFR pulls torch, Reductions pins fairlearn to a 0.7 series, FairAdapt needs rpy2, and inFairness needs skorch and inFairness. requirements.txt pins several of these tightly, including rpy2==3.4.5, skorch==0.11.0, pot==0.9, lightgbm==3.1.1 and igraph[plotting]==0.9.8. If your environment already carries a different TensorFlow or fairlearn version, the extras are where the conflict will surface. The README's own advice to use a virtual environment is the mitigation, and it is worth taking literally rather than treating as boilerplate.

A second limitation is conceptual rather than technical. Every metric in the group fairness families requires a protected attribute with a defined privileged group. If the attribute is absent, incomplete, or must be inferred, the toolkit has nothing to compute on. The README does not describe an imputation path for protected attributes.

A third is the selection problem the authors flag themselves. Fifteen mitigation algorithms with different fairness definitions do not compose into a single answer, and the definitions conflict: satisfying equalized odds and satisfying calibration simultaneously is not generally possible, which is precisely why Calibrated Equalized Odds Postprocessing exists as a separate method rather than a default. Choosing an algorithm is a policy decision that the library does not make for you.

Finally, the README is silent on deployment. There is no documented rollback procedure, no guidance on serving a postprocessed model, and no description of how the mlops/ directory is meant to be used in production. Treat AIF360 as an analysis and experimentation toolkit until you have verified otherwise.

## AIF360 compared with fairlearn

The closest alternative in the same Python ecosystem is fairlearn, and the difference is one of scope and dependency weight rather than of quality. Fairlearn is narrower by design: it centers on a small set of reduction-based mitigation approaches and a metrics interface tied closely to scikit-learn estimators. AIF360 is broader: it spans preprocessing, inprocessing and postprocessing, and it includes metric families fairlearn does not carry, such as Bias Scan and the Generalized Entropy Index.

The dependency story differs in a way that matters for adoption. AIF360 depends on fairlearn for its Reductions extra, pinned to the 0.7 series in requirements.txt. That means the two are not fully independent choices: installing AIF360's reductions support brings fairlearn along. If your team already standardizes on fairlearn, adding AIF360 is additive and mostly about the extra metrics and the preprocessing and postprocessing algorithms.

There is also an interface difference. Fairlearn leans on scikit-learn conventions, so a model that is already a scikit-learn estimator fits naturally. AIF360 wraps datasets and models in its own objects, which is what enables the before-and-after comparison across preprocessing and postprocessing, but it also means more adaptation work if your pipeline is not already structured that way. The examples/sklearn/ directory in the repository suggests the project is aware of this and provides some bridging.

## Maintenance, licence and what upgrading costs

The repository is not archived, and the last push was on 2026-06-15. That is recent enough that the codebase is being touched, but the release cadence tells a different story: v0.6.1 shipped on 2024-04-08, v0.6.0 on 2024-02-23, and v0.5.0 before that on 2022-09-03. Between v0.5.0 and v0.6.0 there were roughly eighteen months, and the gap between v0.6.1 and today is longer still. Commits land on main, but tagged releases are infrequent, so pinning to a release means pinning to something that may be well behind the branch.

Upgrade cost is dominated by the extras, not by the core package. The tight pins in requirements.txt (rpy2, skorch, pot, lightgbm, igraph) are the places where a version bump elsewhere in your stack will collide. The setup.py extras are declared with looser bounds than requirements.txt, for example tensorflow>=1.13.1 and fairlearn~=0.7, so installing via an extra and installing via requirements.txt can resolve differently. If reproducibility matters, pin explicitly and test the extras you actually use rather than installing [all].

The licence is Apache-2.0, as stated in the repository's LICENSE file and in the project metadata. Apache-2.0 is a permissive licence that includes an express patent grant, which is one reason it is common for research-derived libraries. That is a statement about the licence text, not legal advice; if you are embedding this in a commercial product, have your own counsel review the notice and attribution requirements.

One more cost is documentation drift. The README covers setup and lists algorithms and metrics, and the project points to Read the Docs and a separate resources page for guidance. The README does not document rollback, serving, or the MLops directory, so those are areas where you will be reading source.

## Conclusion

AIF360 suits teams that already have a trained model and labelled protected attributes, and want to compare several mitigation strategies on the same dataset rather than commit to one method. It is the wrong tool if you need a single deployable fix, if your protected attributes are missing or inferred, or if your runtime is already pinned to a TensorFlow or fairlearn version that conflicts with the extras. Before adopting, verify three things on your own data: that the metric you plan to report is implemented for your label type, that the extra for your chosen algorithm installs cleanly in your environment, and that the postprocessing methods behave acceptably on your holdout split, since the README does not document rollback or production serving behaviour.

## FAQ

### What is AI Fairness 360?

AIF360 is an open source toolkit from IBM Research that packages fairness metrics for datasets and models, explanations for those metrics, and algorithms that mitigate bias in datasets and models. It ships as both a Python package and an R package on CRAN.

### How do I install AIF360?

The README recommends creating a virtual environment first, for example with conda, then running pip install aif360 for the base package. Algorithms with heavier dependencies are gated behind extras such as pip install 'aif360[LFR,OptimPreproc]', or 'aif360[all]' for everything.

### How do I use AIF360?

The README points to the interactive experience for concepts, the tutorials and notebooks in examples/ for a data scientist-oriented introduction, and the full API for reference. The example notebooks require a manual install with datasets placed as described in aif360/data/README.md.

### Is AIF360 free?

Yes. The repository is licensed under Apache-2.0, and the package is published on PyPI for Python and on CRAN for R. Apache-2.0 is a permissive licence that also includes an express patent grant.

### How does fairlearn compare with AIF360?

Fairlearn is narrower and leans on scikit-learn conventions, while AIF360 spans preprocessing, inprocessing and postprocessing algorithms plus metric families such as Bias Scan and the Generalized Entropy Index. AIF360's Reductions extra depends on fairlearn, pinned to the 0.7 series in requirements.txt.

## Sources

- [License: Apache-2.0](https://github.com/Trusted-AI/AIF360/blob/main/LICENSE)
- [Project website](https://aif360.res.ibm.com/)
- [README](https://github.com/Trusted-AI/AIF360/blob/main/README.md)
- [Releases](https://github.com/Trusted-AI/AIF360/releases)
- [Trusted-AI/AIF360 on GitHub](https://github.com/Trusted-AI/AIF360)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/trusted-ai-aif360
