imbalanced-learn: resampling techniques for skewed datasets in scikit-learn pipelines
A Python Package to Tackle the Curse of Imbalanced Datasets in Machine Learning
At a glance
- What is it?
- imbalanced-learn provides a suite of resampling algorithms (SMOTE, random undersampling, combinations) for datasets where one class vastly outnumbers others. It plugs into scikit-learn pipelines and works with pandas DataFrames, TensorFlow, and Keras.
- Who is it for?
- Adopt imbalanced-learn if you have classification datasets where one class is 10 or 100 times rarer than the majority and your classifier struggles to detect the minority. Skip it if your data is already balanced, your metric is not recall or F1, or cost-sensitive learning (class_weight in your classifier) is sufficient.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 93 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Resampling algorithms for the minority class problem
imbalanced-learn is a Python package that implements resampling algorithms to handle datasets with severe class imbalance. In a balanced binary classification problem, both classes have roughly equal samples. In an imbalanced problem, one class is rare and easily ignored by the loss function during training. The standard machine learning solution is to alter the training data to balance the classes via over-sampling the minority (creating more copies or synthetic examples), under-sampling the majority (removing examples), or a combination of both. imbalanced-learn bundles these resampling techniques as transformers that follow scikit-learn's API and fit into scikit-learn pipelines. The package is compatible with scikit-learn and part of the scikit-learn-contrib ecosystem. The README states that the development of this project aligns with the scikit-learn community standards. The last push was on 2026-06-29, about three months before today. The latest release, 0.14.2, came out in June 2026. The project has a published academic paper from 2017 in the Journal of Machine Learning Research documenting the design and algorithms.
Installation with pip or conda
Install imbalanced-learn via pip:
pip install -U imbalanced-learnOr via conda-forge:
conda install -c conda-forge imbalanced-learnThe package requires Python 3.10 or later and scikit-learn 1.4.2 or later. Core dependencies include NumPy 1.25.2 or later, SciPy 1.11.4 or later, joblib for parallelization, and threadpoolctl for controlling thread pools. For working with pandas DataFrames while preserving column names, install pandas 2.0.3 or later. If you use TensorFlow or Keras models, install tensorflow 2.16.1 or later and keras 3.3.3 or later as optional dependencies. The pyproject.toml file lists complete version constraints for all dependencies. Development installation with pip install --no-build-isolation --editable . allows contributing to the package. The examples in the documentation require matplotlib 3.7.3 or later and seaborn 0.12.2 or later for visualization.
SMOTE, undersampling, and combination methods
imbalanced-learn implements multiple resampling approaches for different data scenarios. SMOTE (Synthetic Minority Over-sampling Technique) creates synthetic samples of the minority class by interpolating between existing minority examples in feature space. The generated samples lie on the line segments connecting k nearest neighbors of each minority example. Other over-sampling methods include random duplication of existing minority samples and variants like ADASYN that focus generation on harder-to-learn regions. Under-sampling removes majority-class samples to reduce the class imbalance ratio. Near Miss is one under-sampling strategy that selects majority samples closest to the decision boundary. Combination methods apply both strategies sequentially: for example, SMOTE followed by edited nearest neighbors to clean up overlapping or misclassified samples. The README directs users to the documentation at https://imbalanced-learn.org/stable/user_guide.html for a complete list of implemented algorithms and their behavior. Each technique has trade-offs: SMOTE can produce unrealistic synthetic samples if the feature space is sparse or high-dimensional, and synthetic samples may not represent true minority patterns; under-sampling discards potentially useful data from the majority class and may lose important patterns that could inform the decision boundary; combination methods are more complex to tune and require careful validation on held-out test data to avoid overfitting to the resampled distribution.
Building end-to-end pipelines with scikit-learn
imbalanced-learn transformers implement scikit-learn's Pipeline API (fit, transform, and fit_transform methods), allowing you to chain resampling with preprocessing and classification seamlessly. The repository examples directory shows how to build end-to-end pipelines that apply resampling during training only, preventing test data leakage. A typical pipeline starts with feature preprocessing (scaling, encoding), then applies resampling on each cross-validation training fold, then trains the final classifier. imbalanced-learn works seamlessly with pandas DataFrames, preserving column names in resampled output so you can track which features remain after under-sampling and verify that the resampling did not corrupt your data. The optional support for TensorFlow and Keras models broadens compatibility for deep learning workflows where you want to resample before training a neural network. The pyproject.toml shows that imbalanced-learn depends on joblib, which enables parallelization of some resampling operations across multiple cores, speeding up algorithms significantly on multi-core systems.
Cost-sensitive learning and threshold tuning as alternatives
Resampling addresses class imbalance by changing the training distribution, but it is one lever among several. Cost-sensitive learning assigns higher misclassification cost to the minority class; many scikit-learn classifiers support class_weight='balanced' to apply this without data manipulation. Random forest, decision trees, and logistic regression all support this flag. Threshold moving after training changes the decision boundary based on the cost matrix: instead of classifying samples with probability > 0.5 as positive, you lower the threshold to 0.3 or 0.2 to increase recall at the expense of precision. If your classifier was trained correctly in the first place, resampling may not help; verify that the raw classifier performance (using metrics like precision, recall, and F1 on the minority class) is actually poor before applying resampling. The fundamental challenge is that no resampling method will extract signal that is not in the data. If the minority examples are outliers or the signal is too weak to distinguish them, both resampling and cost-sensitive approaches will struggle.
MIT license, production stability, and ongoing maintenance
imbalanced-learn is MIT licensed with no restrictions on commercial use. The development status is marked as Production/Stable in pyproject.toml, indicating it is ready for production workloads. The README includes a citation for the academic paper published in the Journal of Machine Learning Research in 2017 that describes the project's design and algorithms. The project endorses the Scientific Python Specification (SPEC) for minimum supported dependencies, committing to a defined window of support for older versions of dependencies. This eases transitions in large Python environments where upgrading all packages at once is not feasible. The project receives active maintenance; version 0.14.0 came out in August 2025, 0.14.1 in December 2025, and 0.14.2 in June 2026. The repository uses GitHub Actions for continuous integration, CircleCI for testing, and Codecov for coverage reporting, showing robust testing infrastructure.
Editorial conclusion
Adopt imbalanced-learn if you have classification datasets where one class is 10 or 100 times rarer than the majority and your classifier struggles to detect the minority. Skip it if your data is already balanced, your metric is not recall or F1, or cost-sensitive learning (class_weight in your classifier) is sufficient. Verify first that you understand the difference between resampling in training (correct) and in test data (leakage). Run the examples from the documentation on your own data to understand which algorithms help your specific problem; there is no universal best resampling strategy.
Frequently asked questions
How do I install imbalanced-learn using pip?
Run pip install -U imbalanced-learn. You need Python 3.10 or later and scikit-learn 1.4.2 or later. You can also use conda install -c conda-forge imbalanced-learn.
What is imbalanced-learn?
imbalanced-learn is a Python package providing resampling algorithms (SMOTE, undersampling, combinations) for datasets where one class is much rarer than others. It works as a transformer in scikit-learn pipelines.
How can I balance an imbalanced dataset?
imbalanced-learn provides multiple strategies: SMOTE creates synthetic minority samples, undersampling removes majority samples, and combination methods use both. Apply resampling to training data only, never to test data.
How can I fix imbalanced classes in my classifier?
Use resampling (SMOTE, undersampling) via imbalanced-learn, or enable cost-sensitive learning with class_weight='balanced' in scikit-learn classifiers. Verify that your evaluation metric (recall, F1, precision-recall curve) reflects what matters for your problem.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/scikit-learn-contrib-imbalanced-learn)