Library / SDK
feature-engine/feature_engine avatar
feature-engine/feature_engine

Feature-engine: sklearn-Compatible Python Library for Feature Engineering and Selection

Feature engineering and selection open-source Python library compatible with sklearn.

2,283 stars378 forksPythonBSD-3-Clause

At a glance

What is it?
Feature-engine is a Python library that provides over 60 transformers for feature engineering and feature selection tasks, all following the scikit-learn fit/transform API convention. It covers missing data imputation, categorical encoding, discretisation, outlier handling, variable creation, and multiple feature selection methods including MRMR.
Who is it for?
Feature-engine suits data scientists and ML engineers who want a dedicated, well-documented library of feature engineering and selection transformers that fit into scikit-learn pipelines without modification. It is a poor fit for projects that are locked to older pandas (below 2.2.0) or scikit-learn (below 1.4.0) versions, since the current release requires both.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 11 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Feature-engine provides and who uses it

Feature-engine addresses a gap in scikit-learn's transformer coverage. Scikit-learn provides a core set of preprocessing tools, but many feature engineering tasks, such as encoding rare categories, applying Weight-of-Evidence encoding, discretizing with decision trees, or selecting features by MRMR, require either custom code or external libraries.

Feature-engine fills that space with transformers that follow the same fit/transform pattern scikit-learn transformers use. A Feature-engine transformer learns its parameters from training data in the fit() call, then applies the learned transformation in transform(), making it compatible with scikit-learn's Pipeline, GridSearchCV, and cross-validation utilities without modification.

The library is maintained by Soledad Galli, a machine learning educator who authored the Python Feature Engineering Cookbook published by Packt and runs the trainindata.com online course platform. The courses explicitly use Feature-engine for their examples, which means the library's API is tuned to match how the underlying concepts are taught.

Installing Feature-engine and its dependencies

Feature-engine is available on PyPI and conda-forge:

code
pip install feature_engine

Or from Anaconda:

code
conda install -c conda-forge feature_engine

The runtime dependencies specified in pyproject.toml are:

- numpy >= 1.18.2 - pandas >= 2.2.0 - scikit-learn >= 1.4.0 - scipy >= 1.4.1

The pandas >= 2.2.0 and scikit-learn >= 1.4.0 lower bounds are significant. Projects running pandas 1.x or scikit-learn 1.3.x cannot install the current release without upgrading those dependencies. The package supports Python 3.9 through 3.14.

Transformer categories: what Feature-engine covers

The README lists thirteen categories of functionality:

Imputation covers seven transformers: MeanMedianImputer, ArbitraryImputer, RandomSampleImputer, EndTailImputer, CategoricalImputer, MissingIndicator, and DropMissingData. The EndTailImputer fills missing values with the distribution's tails rather than the mean, which avoids shifting the distribution center for skewed variables.

Categorical Encoding has eight transformers including WoEEncoder (Weight of Evidence), RareLabelEncoder, DecisionTreeEncoder, and StringSimilarityEncoder. The RareLabelEncoder groups infrequent categories into a single Rare label, which is common in production pipelines where unseen categories during deployment break downstream models.

Discretisation offers five transformers including EqualFrequencyDiscretiser, EqualWidthDiscretiser, and DecisionTreeDiscretiser, which bins numerical variables using decision tree splits to optimize predictive power.

Outlier Handling has three transformers: Winsoriser (caps values at a specified percentile), ArbitraryOutlierCapper, and OutlierTrimmer (removes rows with outliers).

Feature Selection includes fourteen transformers. Notable ones: SmartCorrelationSelection (removes correlated features while retaining the best predictor), DropHighPSIFeatures (drops features whose population stability index exceeds a threshold, useful for model monitoring), and MRMR (minimum redundancy maximum relevance).

Variable Creation covers MathFeatures (arithmetic combinations), RelativeFeatures (ratios), CyclicalFeatures (sin/cos encoding for periodic variables), DecisionTreeFeatures, and GeoDistanceFeatures.

Time Series features include LagFeatures, WindowFeatures, and ExpandingWindowFeatures for temporal forecasting pipelines.

How Feature-engine works in a scikit-learn Pipeline

The README's example demonstrates the fit_transform pattern directly. The RareLabelEncoder groups categories with fewer than 10% of observations:

python
import pandas as pd
from feature_engine.encoding import RareLabelEncoder

data = {'var_A': ['A'] * 10 + ['B'] * 10 + ['C'] * 2 + ['D'] * 1}
data = pd.DataFrame(data)
rare_encoder = RareLabelEncoder(tol=0.10, n_categories=3)
data_encoded = rare_encoder.fit_transform(data)

After transformation, categories C and D, which together account for fewer than 10% of observations, are grouped into a single Rare label. The fit() call on training data learns which categories qualify as rare. The transform() call on new data applies the same grouping, including mapping any new, unseen categories to Rare by default.

Because every Feature-engine transformer implements the scikit-learn estimator interface, they compose directly into scikit-learn Pipeline objects using the standard make_pipeline or Pipeline constructors. Feature-engine also ships its own Pipeline and make_pipeline exports that behave identically to scikit-learn's but allow Feature-engine-specific transformers to be included with no additional adapters.

This compatibility means cross-validation, hyperparameter search with GridSearchCV or RandomizedSearchCV, and model evaluation utilities from scikit-learn all work normally with Feature-engine transformers as steps. A pipeline that starts with MeanMedianImputer, moves through RareLabelEncoder, and ends in a gradient boosting classifier can be tuned with a single GridSearchCV call without writing any custom cross-validation logic.

Where Feature-engine differs from scikit-learn's built-in preprocessors

Scikit-learn includes OrdinalEncoder, OneHotEncoder, SimpleImputer, PowerTransformer, and a selection of feature selection utilities. These cover many common cases but have specific limits.

Scikit-learn's SimpleImputer fills missing values with mean, median, most frequent value, or a constant. It does not fill with end-tail values (distribution mean plus three standard deviations, for example), and it does not flag missing values as a binary indicator column while also imputing, which Feature-engine's MissingIndicator and separate imputers support in a single pass.

Scikit-learn's OrdinalEncoder and OneHotEncoder handle categories seen in fit() but have different default behaviors for unseen categories during transform(). Feature-engine's RareLabelEncoder provides explicit control over which categories survive encoding and how infrequent ones are treated.

For feature selection, scikit-learn includes SelectKBest and RFE. Feature-engine adds MRMR, SmartCorrelationSelection, DropHighPSIFeatures (which is production monitoring logic, not a training-time selector), and ProbeFeatureSelection, none of which have direct equivalents in scikit-learn's selection module.

Limitations and version history

The requirements pin pandas >= 2.2.0 and scikit-learn >= 1.4.0. These are recent lower bounds. An existing project that has not yet upgraded from pandas 1.x will encounter a version conflict when adding Feature-engine as a dependency. The pyproject.toml also pins scipy >= 1.4.1 and numpy >= 1.18.2, which are less restrictive but still worth verifying in environments that pin transitive dependencies for reproducibility.

The release history visible in the repository shows v1.2.0 in January 2022 and v1.9.4 in July 2026. That gap spans over four years of development with no intermediate tags visible in the recent releases list, meaning many minor and patch releases happened in between. Users upgrading from 1.2.0 to 1.9.4 should expect breaking changes and read the changelog before updating in a production pipeline.

Feature-engine does not cover deep learning-based feature representations, NLP vectorization beyond basic text statistics, or image feature extraction. It focuses exclusively on tabular data transformation tasks.

Maintenance and license

The repository is not archived. The last push was on 2026-09-19. The most recent release is v1.9.4 from 2026-07-21. The project has a published JOSS (Journal of Open Source Software) paper, cited via DOI 10.21105/joss.03642, and a Zenodo DOI for citation.

The license is BSD 3-Clause. The dependency chain (NumPy, pandas, scikit-learn, scipy) also carries permissive licenses, making Feature-engine suitable for use in proprietary commercial applications without license constraints.

The project is developed and maintained by Soledad Galli at trainindata.com, with JetBrains listed as an open-source sponsor.

Editorial conclusion

Feature-engine suits data scientists and ML engineers who want a dedicated, well-documented library of feature engineering and selection transformers that fit into scikit-learn pipelines without modification. It is a poor fit for projects that are locked to older pandas (below 2.2.0) or scikit-learn (below 1.4.0) versions, since the current release requires both. Before upgrading from an older Feature-engine version, check the changelog in the repository's CONTRIBUTING.md or documentation site since the jump from 1.2.0 (January 2022) to 1.9.4 (July 2026) spans multiple breaking changes that are not summarized in the README.

Frequently asked questions

How do I install Feature-engine?

Install from PyPI with pip install feature_engine, or from Anaconda with conda install -c conda-forge feature_engine. The library requires Python 3.9 through 3.14, pandas 2.2.0 or later, and scikit-learn 1.4.0 or later.

How does Feature-engine compare to scikit-learn for feature engineering?

Feature-engine follows the same fit/transform API as scikit-learn and works inside scikit-learn Pipelines. It adds transformers not available in scikit-learn, such as RareLabelEncoder, WoEEncoder, EndTailImputer, MRMR selection, DropHighPSIFeatures, and decision-tree-based discretisation and encoding.

Does Feature-engine work with time series data?

Yes, partially. Feature-engine includes LagFeatures, WindowFeatures, and ExpandingWindowFeatures for creating lag-based and window-based features from time series. It does not cover forecasting models or time series cross-validation, which fall outside its scope.

Official sources

  1. feature-engine/feature_engine on GitHub
  2. License: BSD-3-Clause
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/feature-engine-feature-engine.svg)](https://hysenlabs.com/projects/feature-engine-feature-engine)