Model or dataset
unitaryai/detoxify avatar
unitaryai/detoxify

detoxify ships three Jigsaw toxicity models, and the most current thing in its README is a news entry from 2021

Trained models & code to predict toxic comments on all 3 Jigsaw Toxic Comment Challenges. Built using ⚡ Pytorch Lightning and 🤗 Transformers. For access to our API, please email us at [email protected].

1,305 stars146 forksPythonApache-2.0

At a glance

What is it?
detoxify is a Python library of pre-trained classifiers for the three Jigsaw toxic comment challenges, built on Transformers with weights served through torch hub. Its leaderboard table compares itself against ensemble submissions rather than single models, its language breakdown lists six of the seven languages the model claims, and its runtime dependencies are three open-ended lower bounds while PyTorch Lightning is called an inference requirement and installed only as a dev extra.
Who is it for?
detoxify is worth using if you want a strong single-model baseline for toxicity triage with an honest account of its failure modes, because the limitations section names the exact failure, profanity read as toxicity regardless of tone, instead of burying it.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 8, 2026, and from our analysis. They are not legal advice.

Editorial analysis

PyTorch Lightning is called an inference dependency and is not installed

The dependency list has two parts. For inference it names Hugging Face Transformers and PyTorch Lightning. For training it adds the Kaggle API, because that is how the data arrives. The installed package disagrees with the first list. The runtime dependency array in pyproject.toml contains exactly three entries: `sentencepiece >= 0.1.94`, `torch >=2`, and `transformers >= 3`. PyTorch Lightning appears only in the development extra, as `pytorch-lightning>2`. So a headline that says the library is built with PyTorch Lightning and a dependency list that says you need PyTorch Lightning to run it are describing different things, most likely a training-time dependency that survived in the prose. All three of the bounds you do get are floors with no ceiling, which is the opposite discipline from the pin-everything projects this one is often compared with.

The news section ends in 2021 while the releases continued to 2026

The news and updates block has four entries and the newest is dated 22-10-2021. Before that: a new unbiased model on 03-09-2021 with a test score of 93.74 against 93.64 before, a Scientific American opinion piece on 15-02-2021, and the lightweight Albert models on 14-01-2021 where the small original model reached a mean AUC of 98.28 against 98.64 before. The 2021 entry is also where a breaking change hides: all models were made to return consistent class names, so `identity_attack` replaced `identity_hate` in the `original` model's output to match the `unbiased` classes. That renames a key in your results dictionary. The release history tells a different timeline: v0.5.1 on 2022-12-19, v0.5.2 on 2024-02-01 after a two-year gap, and v0.5.3 on 2026-03-26, with the last commit on 2026-07-06. Five years of releases, five years of silence in the changelog, and a default branch still called master.

The leaderboard comparison is against ensembles

The three rows pair a challenge year and its data source against a leaderboard score and a score from this library. The 2018 toxic comment challenge on Wikipedia comments gives `original` 98.64 against a top leaderboard figure of 98.86. The 2019 unintended bias challenge on Civil Comments gives `unbiased` 93.74 against 94.73. The 2020 multilingual challenge on Wikipedia plus Civil comments gives `multilingual` 92.11 against 95.36. Then the sentence that should be read first: the top leaderboard scores were achieved using model ensembles, and the stated purpose of the library was to be user-friendly and straightforward rather than to win a competition. That is a fair framing and the document is upfront about it. It does mean the 3.25-point gap on the multilingual row is not a ranking, and that anyone quoting these numbers as a leaderboard position is quoting a comparison the author has already disclaimed.

The prediction API is the small part that is easy to get wrong. Each model takes either a string or a list of strings:

python
from detoxify import Detoxify

# each model takes in either a string or a list of strings

results = Detoxify('original').predict('example text')

results = Detoxify('unbiased').predict(['example text 1','example text 2'])

The device is chosen with an argument that defaults to cpu and accepts any torch.device input, so the same code runs on a laptop and on a GPU without a separate path.

Seven languages are claimed and the breakdown lists six

The multilingual model is described as trained on seven languages, with the instruction that it should only be tested on English, French, Spanish, Italian, Portuguese, Turkish, or Russian, and the prediction example passes exactly seven strings, one per language. The language breakdown table then lists six. Italian scores 89.18 on 8,494 examples, French 89.61 on 10,920, Russian 89.81 on 10,948, Portuguese 91.00 on 11,012, Spanish 92.74 on 8,438, and Turkish 97.19 on 14,000. The pattern across those six is that the largest subgroup has the best score and the two smallest have the two worst, which is the opposite of what a per-language breakdown usually shows and worth checking against your own traffic. The missing seventh row is English, which is the one language the example code passes first and the one the model was not built for.

The limitations section is the reason to read this library

Most model cards bury this part. Here it comes early and it is specific. If a comment contains words associated with swearing, insults, or profanity, it is likely to be classified as toxic regardless of the tone or the intent of the author, and the example given is humour and self-deprecation. The stated consequence is that the model can present biases toward already vulnerable minority groups. The intended uses are then narrowed to three: research, fine-tuning on carefully constructed datasets that reflect real-world demographics, and helping content moderators flag harmful content faster. Three references are given for the risk research rather than the marketing, on racial bias in hate speech detection, on automated detection and the problem of offensive language, and on bias in hate speech and abusive language datasets. The labelling schema behind all of this is up to ten annotators collapsing into four buckets, very toxic, toxic, hard to say, and not toxic.

The test suite runs your docstrings and the lint rules are three letters wide

The pytest configuration does two unusual things. `--doctest-modules` means every docstring example in the package is executed as a test, so the snippets in documentation are load-bearing and a stale one fails the suite. `--strict-markers` means an unregistered marker is an error rather than a warning. Alongside them sit `--durations=0` and `--color=yes`, and three `filterwarnings` entries that silence deprecation aliases coming from tensorboard's `tensorflow_stub` and `tensor_util` and from pyarrow's pandas compatibility layer, which tells you the dependency graph has some shims in it. The ruff configuration sets a line length of 120, targets py39, selects the E, F, and W rule families, and ignores four specific codes, with double quotes and space indentation enforced by the formatter. `requires-python` is `>=3.9,<3.13`, which is the one bounded range in the file.

Training needs Kaggle credentials and the weights load through torch hub

Two things about how the models move. A `hubconf.py` at the repository root is the PyTorch Hub hook, so the classifiers are reachable through the hub mechanism rather than only through the Python package. A `convert_weights.py` at the same level explains the Lightning-to-Transformers path, which is consistent with a training stack built on Lightning and a runtime stack built on Transformers with only three installed dependencies. Training also needs the Kaggle API to fetch the data, so neither the weights nor the datasets are fully reproducible from this repository alone; you get the converted weights and the code. `train.py`, `run_prediction.py`, a `model_eval/` directory, and a `configs/` directory cover the rest, with CITATION.cff and CONTRIBUTING.md at the root. Two contact addresses appear: the repository description points API access at one address on the company's domain, and the package author field uses a different one.

Editorial conclusion

detoxify is worth using if you want a strong single-model baseline for toxicity triage with an honest account of its failure modes, because the limitations section names the exact failure, profanity read as toxicity regardless of tone, instead of burying it. It is a poor fit if you need multilingual coverage outside its six documented subgroups, if you need numbers comparable to a leaderboard winner, or if you need reproducibility, since training data sits behind the Kaggle API. Before you build on it, read the limitations section and decide whether your use tolerates the profanity false positives, treat the AUC figures as single-model results rather than rankings, and note that PyTorch Lightning is named as an inference dependency but is not installed with the package.

Frequently asked questions

What is detoxify and which models does it ship?

It is a Python library of trained classifiers for the three Jigsaw toxic comment challenges. `original` covers the 2018 toxic comment classification challenge on Wikipedia comments, `unbiased` covers the 2019 unintended bias challenge on Civil Comments, and `multilingual` covers the 2020 multilingual challenge. Smaller Albert variants exist as `original-small` and `unbiased-small`.

Which languages does the detoxify multilingual model support?

The documentation says seven: English, French, Spanish, Italian, Portuguese, Turkish, and Russian, and the prediction example passes one string per language. The published language breakdown lists only six of them, omitting English, with Turkish scoring highest at 97.19 on 14,000 examples and Italian lowest at 89.18 on 8,494.

How do I install detoxify and run a prediction?

With `pip install detoxify`, then `from detoxify import Detoxify` and a call such as `Detoxify('original').predict('example text')`. Each model accepts a string or a list of strings, and the device is chosen with an argument that defaults to cpu and accepts any torch.device input, for example `Detoxify('original', device='cuda')`.

What are the known limitations of detoxify?

Comments containing words associated with swearing, insults, or profanity are likely to be classified as toxic regardless of tone or intent, with humour and self-deprecation given as examples, which can present biases toward already vulnerable minority groups. The stated intended uses are research, fine-tuning on carefully constructed datasets reflecting real-world demographics, and helping content moderators flag harmful content faster.

Can I train a detoxify model from this repository?

The training path also needs the Kaggle API to download the data, so the datasets are not in the repository. What is included is the training code, a `configs/` directory, a `model_eval/` directory, `train.py`, and `convert_weights.py` for converting the checkpoints, plus `hubconf.py` so the released weights can be loaded through PyTorch Hub.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. unitaryai/detoxify on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/unitaryai-detoxify.svg)](https://hysenlabs.com/projects/unitaryai-detoxify)