# category_encoders: sklearn-compatible encoders for categorical variables

> category_encoders collects unsupervised and supervised transformers for turning categorical columns into numbers, and the supervised ones are where the design decisions get interesting. It is for pandas and scikit-learn users who have outgrown OneHotEncoder on high-cardinality columns.

**scikit-learn-contrib/category_encoders** — A library of sklearn compatible categorical variable encoders

- Repository: https://github.com/scikit-learn-contrib/category_encoders
- Website: http://contrib.scikit-learn.org/category_encoders/
- Stars: 2,508 · Forks: 415
- Language: Python
- License: BSD-3-Clause
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/scikit-learn-contrib-category-encoders

## What category_encoders solves, and who it is for

The problem is narrow and specific: a pandas DataFrame with object or category columns that a scikit-learn estimator will not accept. OneHotEncoder handles the low-cardinality case, but a column with thousands of levels produces a sparse matrix that is wide, slow, and often worse than a single informative number. category_encoders supplies that number through a set of transformers that follow the scikit-learn convention, so they drop into a Pipeline next to any estimator. The README describes the project as "a set of scikit-learn-style transformers for encoding categorical variables into numeric by means of different techniques."

The intended user is a practitioner already working in pandas and scikit-learn who wants to try several encodings without leaving that stack. The README states that supported input formats include numpy arrays and pandas dataframes, and that if the cols parameter is not passed, every column with object or pandas categorical dtype is encoded. That default matters in practice: a DataFrame carrying an id column of strings will be encoded unless you name columns explicitly. The library is not a feature store, not a serving layer, and not a replacement for domain-specific binning. It is a preprocessing step.

## Unsupervised and supervised encoders in one namespace

The README splits the encoders into two groups. Unsupervised methods include Backward Difference Contrast, BaseN, Binary, Gray, Count, Hashing, Helmert Contrast, Ordinal, One-Hot, Multi-Hot, Rank Hot, Polynomial Contrast, and Sum Contrast. These depend only on the feature column. Count replaces a level with its frequency; Binary and BaseN represent the level index in a compact numeral system; Hashing maps levels into a fixed number of buckets; the contrast family (Helmert, Sum, Polynomial, Backward Difference) produces columns whose coefficients are interpretable against a reference, which is why they appear in statistics curricula.

Supervised methods use the target and include CatBoost, Count Target Encoding, Generalized Linear Mixed Model, James-Stein Estimator, LeaveOneOut, M-estimator, Target Encoding, Weight of Evidence, Quantile Encoder, and Summary Encoder. The split is not cosmetic. A supervised encoder that is fitted on the whole training set and then used to transform that same training set leaks the target into the features, and the downstream model learns a shortcut that does not exist at prediction time. The library's answer is documented explicitly: for supervised methods on training data you should call fit_transform() rather than fit().transform(), because the two do not have to produce the same result. The README uses LeaveOneOut as the example: fit_transform() runs a nested cross-validation for the training data to reduce overfitting of the downstream model, while transform() uses all the training data to get as accurate an estimate as possible on new rows. Two wrappers extend this further. PolynomialWrapper extends supervised encoders to polynomial targets, and NestedCVWrapper is described as helping to prevent overfitting.

## Installing category_encoders and running a first encoding

The README lists three installation routes: python setup.py install, pip install category_encoders, and conda install -c conda-forge category_encoders. For a development build it gives pip install --upgrade git+https://github.com/scikit-learn-contrib/category_encoders. The package requires numpy, scipy, and statsmodels according to the README, but pyproject.toml is more precise: statsmodels is an optional dependency declared under the glmm extra, and it is only needed for the Generalized Linear Mixed Model encoder. If you install with pip and later reach for that encoder, you will need the extra.

```bash
pip install category_encoders
```

After that, an unsupervised encoder needs only the feature frame. The README's BinaryEncoder example builds a small DataFrame with gender, country, and age, fits the encoder on the two categorical columns, and transforms the frame.

```python
from category_encoders import *
import pandas as pd

X = pd.DataFrame({
    'gender': ['male', 'female', 'female', 'male', 'female'],
    'country': ['US', 'UK', 'US', 'CA', 'UK'],
    'age': [25, 32, 47, 51, 38],
})

enc = BinaryEncoder(cols=['gender', 'country']).fit(X)
numeric_dataset = enc.transform(X)
```

What you should see is a frame where gender and country have been replaced by integer columns and age is untouched. The supervised path takes a target. The README's TargetEncoder example fits on X_train and y_train and transforms a separate X_test.

```python
from category_encoders import *
import pandas as pd

X_train = pd.DataFrame({
    'gender': ['male', 'female', 'female', 'male'],
    'country': ['US', 'UK', 'US', 'CA'],
})
y_train = pd.Series([1, 0, 1, 0])
X_test = pd.DataFrame({
    'gender': ['female', 'male'],
    'country': ['UK', 'US'],
})

enc = TargetEncoder(cols=['gender', 'country'])
training_numeric_dataset = enc.fit_transform(X_train, y_train)
testing_numeric_dataset = enc.transform(X_test)
```

Note the asymmetry: fit_transform() on training data, transform() on test data. Swapping those two calls is the single most common way to get a misleading validation score with this library.

## The fit_transform versus transform split is the real API

Most preprocessing libraries treat fit().transform() and fit_transform() as interchangeable. Here they are not, and the README says so directly. The reason is that LeaveOneOut, and to a lesser degree the smoothing-based estimators, need two different behaviours: an out-of-fold estimate when the rows are part of training, and a full-data estimate when the rows are new. If you write a scikit-learn Pipeline and call pipeline.fit(X_train, y_train), the internal fit_transform path is used and you get the safer behaviour. If you instead fit the encoder once on the full training set and transform the same set, you get the leakier behaviour. That distinction is easy to lose when the encoder is buried inside a ColumnTransformer, which is why the repository ships examples/column_transformer_example.py and examples/grid_search_example.py alongside the general encoding_examples.py.

The trade-off is that fit_transform() is more expensive. LeaveOneOut performs a nested cross-validation on the training data, so the cost scales with the number of folds rather than with a single pass. On a large frame with a high-cardinality column this can dominate the runtime of the whole pipeline, and the README does not document a way to cap the fold count for that encoder. If training time is the binding constraint, the cheaper route is an unsupervised encoder such as Count or Hashing, accepting that you lose the target signal.

## Where category_encoders is the wrong tool

The pandas dependency is structural, not incidental. pyproject.toml requires pandas >=1.0.5 and scikit-learn >=1.6.0, and the README's examples are all DataFrame-based. If your pipeline is pure numpy, or if you are serving in an environment where pandas is not available, adding this library pulls in a dependency you may not want. scikit-learn's own OneHotEncoder and OrdinalEncoder cover the common cases without it.

Target leakage is the second boundary. The supervised encoders are designed to reduce it, not eliminate it. The README is explicit that fit_transform() and transform() can differ, and that NestedCVWrapper exists to help prevent overfitting, which is an admission that the plain encoders can overfit when used carelessly. If you cannot restructure your workflow so that the encoder is fitted inside a cross-validation fold, a supervised encoder will inflate your validation metrics. A target encoder fitted on the full training set before the split is the wrong tool no matter how good the implementation is.

The third boundary is the Generalized Linear Mixed Model encoder. Because statsmodels is declared as an optional extra in pyproject.toml rather than a hard dependency, a plain pip install category_encoders may leave that encoder unusable. The README's dependency sentence lists statsmodels alongside numpy and scipy without flagging it as optional, so the two sources disagree on this point. Treat pyproject.toml as the accurate one and plan to install the extra.

## How it compares with scikit-learn's built-in encoders

The closest alternative is scikit-learn's own preprocessing module, specifically OneHotEncoder and OrdinalEncoder. The difference in approach is breadth versus integration. scikit-learn's encoders are maintained inside the library that consumes them, so version compatibility is guaranteed and the dependency footprint stays at one package. They cover one-hot and ordinal, and OneHotEncoder has options for handling unknown categories and for producing sparse output. What they do not provide is the supervised family: there is no TargetEncoder, no Weight of Evidence, no James-Stein estimator, no LeaveOneOut, and no contrast coding beyond what you assemble yourself.

That is the actual decision. If your categorical columns are low-cardinality and one-hot is acceptable, scikit-learn's encoder is the smaller dependency and the README offers no argument against it. category_encoders becomes worth the extra package when cardinality is high enough that one-hot is impractical, or when you want the target-based estimators and are prepared to handle them inside cross-validation. The contrast encoders are a third case: they exist mainly for interpretability, and scikit-learn does not ship them at all.

## Maintenance, licence, and what an upgrade costs

The repository is not archived, and the last push was on 2026-09-08, the same day as the v2.11.1 release; v2.11.0 landed on 2026-09-07 and 2.10.0 on 2026-07-26. That is a steady release cadence, and the README states that the project is "under active development" with a CONTRIBUTING.md for new contributors. The repository also carries a CHANGELOG.md, a CITATION.cff, a .zenodo.json, and a JOSS submission directory, which suggests a project that expects academic citation as well as production use.

The version floor is the main upgrade cost. pyproject.toml requires Python >=3.11 and scikit-learn >=1.6.0, so upgrading category_encoders on an older environment can force a scikit-learn upgrade, and that in turn can break other estimators in the same environment. The declared dependencies are numpy >=1.14.0, scipy >=1.0.0, pandas >=1.0.5, and the optional statsmodels >=0.9.0. Pinning category_encoders without pinning scikit-learn is how a routine bump turns into a pipeline rebuild.

The licence is BSD-3-Clause, as stated in the repository metadata and as "BSD-3" in pyproject.toml. That is a permissive licence, which in practice means you can use the library in commercial and closed-source work provided you retain the copyright notice and disclaimer. This is a description of the licence text, not legal advice; if the licence terms matter to your organisation, read LICENSE.md and consult your own counsel.

## Conclusion

Adopt category_encoders if your pipeline is already pandas plus scikit-learn and you need target, count, hashing or contrast encodings behind a familiar fit/transform API, especially if you want LeaveOneOut or NestedCVWrapper to keep target leakage out of training folds. Do not adopt it if you cannot accept the pandas dependency, if you need the Generalized Linear Mixed Model encoder and will not install the optional statsmodels extra, or if a single OneHotEncoder already covers your cardinality. Before wiring it into a pipeline, verify three things on your own data: that you use fit_transform() rather than fit().transform() for supervised encoders, that your scikit-learn version satisfies the >=1.6.0 floor declared in pyproject.toml, and that you have a held-out split to confirm the target encoder is not leaking.

## FAQ

### How do I install category_encoders?

The README gives three routes: pip install category_encoders, conda install -c conda-forge category_encoders, or python setup.py install from a checkout. For the development version it gives pip install --upgrade git+https://github.com/scikit-learn-contrib/category_encoders.

### What is category_encoders?

It is a set of scikit-learn-style transformers that encode categorical variables into numeric values using different techniques. The README groups them into unsupervised methods such as Binary, Count, Hashing and the contrast encoders, and supervised methods such as Target Encoding, LeaveOneOut and Weight of Evidence.

### What are the different types of encoders in category_encoders?

The README lists thirteen unsupervised encoders, including Backward Difference Contrast, BaseN, Binary, Gray, Count, Hashing, Helmert Contrast, Ordinal, One-Hot, Multi-Hot, Rank Hot, Polynomial Contrast and Sum Contrast, and ten supervised ones, including CatBoost, Count Target Encoding, Generalized Linear Mixed Model, James-Stein Estimator, LeaveOneOut, M-estimator, Target Encoding, Weight of Evidence, Quantile Encoder and Summary Encoder.

### Is category_encoders compatible with sklearn?

Yes. The README states that all of the encoders are fully compatible sklearn transformers, so they can be used in pipelines or existing scripts. pyproject.toml declares scikit-learn >=1.6.0 as a dependency.

### What is categorical encoding?

It is the conversion of categorical variables into numeric values so that a machine learning estimator can consume them. category_encoders offers this conversion through several techniques, split in the README into unsupervised methods that use only the feature column and supervised methods that also use the target.

## Sources

- [License: BSD-3-Clause](https://github.com/scikit-learn-contrib/category_encoders/blob/master/LICENSE)
- [Project website](http://contrib.scikit-learn.org/category_encoders/)
- [README](https://github.com/scikit-learn-contrib/category_encoders/blob/master/README.md)
- [Releases](https://github.com/scikit-learn-contrib/category_encoders/releases)
- [scikit-learn-contrib/category_encoders on GitHub](https://github.com/scikit-learn-contrib/category_encoders)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/scikit-learn-contrib-category-encoders
