# scikit-learn: the Python machine learning library, and when to reach for something else

> scikit-learn is a BSD-licensed Python module for classical machine learning, built on NumPy and SciPy. It fits tabular data and small to medium models, and it is the wrong tool for deep learning and GPU-scale work.

**scikit-learn/scikit-learn** — scikit-learn: machine learning in Python

- Repository: https://github.com/scikit-learn/scikit-learn
- Website: https://scikit-learn.org
- Stars: 67,429 · Forks: 27,463
- Language: Python
- License: BSD-3-Clause
- Published: 2026-08-04 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/scikit-learn-scikit-learn

## What scikit-learn actually solves

The library covers the part of machine learning that happens before neural networks: fitting a model to a table of features, scoring it, and choosing between candidates. Its estimators share a common interface, so a linear regression, a random forest and a support vector machine are all constructed, fitted and predicted through the same method names. That uniformity is the product. You learn fit and predict once and reuse the knowledge across dozens of algorithms.

The audience is engineers and researchers who have numeric data already loaded as NumPy arrays or pandas DataFrames and want a model today, not a training pipeline next month. The project describes itself as "a Python module for machine learning built on top of SciPy", and the pyproject.toml lists its runtime dependencies as numpy, scipy, joblib, narwhals and threadpoolctl. Nothing in that list is a deep learning framework. The scope is deliberate: preprocessing, feature selection, model selection, and the classical estimators themselves.

It is a poor fit when your problem is images, audio or text at scale, or when training must run on GPUs across many machines. Those workloads need a different stack, and the library does not pretend otherwise.

## The estimator, pipeline and cross-validation mechanism

Everything in scikit-learn is organised around three objects. An estimator holds hyperparameters and learns from data. A transformer is an estimator that also converts data, which is how scaling, imputation and encoding fit into the same interface. A pipeline chains transformers and a final estimator so that the whole sequence is fitted and applied as one unit.

The data flow is array in, array out. You pass a two-dimensional array of shape (n_samples, n_features) plus a target vector, call fit, then call predict or transform on new data. Because the pipeline owns the transformers, preprocessing statistics such as the mean and standard deviation used by a scaler are computed on the training fold only. That is the mechanism that keeps cross-validation honest: model_selection splits the data, and each split refits the entire pipeline rather than reusing statistics leaked from the test portion.

Hyperparameter search follows the same pattern. A search object wraps an estimator, a parameter grid or distribution, and a cross-validation splitter, then exposes the best estimator after fitting. The repository layout reflects this structure: the sklearn/ directory holds the implementation, examples/ is organised by task (classification, cluster, decomposition, model_selection and others), and asv_benchmarks/ holds the airspeed velocity benchmarks the README links to. Nothing in that layout suggests a graph execution engine or a distributed scheduler; the design assumes one process working on in-memory arrays, with joblib handling parallelism across cores.

## Installing scikit-learn with pip and running a first model

The README gives two installation paths. If NumPy and SciPy are already present, pip is the shortest route. The -U flag upgrades an existing installation rather than failing on a version conflict.

```bash
pip install -U scikit-learn
```

Conda users get the same package from the conda-forge channel, which is the path the README documents for environments where binary dependencies are managed by conda rather than wheels.

```bash
conda install -c conda-forge scikit-learn
```

Both commands install the current release line; the project's most recent releases listed are 1.9.0, 1.8.0 and 1.7.2. For a first real use, the README points at the examples directory, which is organised by task, and at the documentation at https://scikit-learn.org. The linear_model examples under examples/linear_model/ are the natural entry point for a first regression, and the module that holds the estimators is sklearn.linear_model. A minimal script imports the estimator, fits it, and inspects the fitted attributes; the README does not print a worked snippet, so the exact call sequence belongs to the example files rather than to this page.

What you should expect after a successful fit is a NumPy array of coefficients, one per feature, stored on the estimator. From there the natural next step is wrapping the estimator in a pipeline with a scaler and evaluating it with cross-validation, which is where the interface starts paying for itself. The README also notes that plotting helpers (functions starting with plot_ and classes ending with Display) need Matplotlib, so those calls will fail on a minimal install. Running the test suite from outside the source directory requires pytest 7.1.2 or newer, and the README documents the SKLEARN_SEED environment variable for controlling random number generation during testing.

## Where scikit-learn stops being the right tool

The most common mismatch is scale. Estimators expect the full dataset in memory as a dense or sparse array. A dataset that does not fit in RAM is a problem the library does not solve at its core, and the documentation does not present an out-of-core training path for arbitrary estimators. If your data lives in a distributed store and cannot be materialised locally, this is the wrong starting point.

The second mismatch is model family. There is no deep neural network implementation here, and the dependency list confirms it: numpy, scipy, joblib, narwhals and threadpoolctl. Anyone arriving from a TensorFlow or PyTorch background expecting GPU-accelerated training will not find it. The two ecosystems answer different questions. scikit-learn gives you a fixed set of well-understood algorithms with a uniform API; a deep learning framework gives you differentiable building blocks and hardware acceleration. Asking which is better is usually a category error.

A quieter limitation is that the uniform interface hides cost. Fitting a pipeline inside a large hyperparameter search multiplies the work by the number of candidates and folds, and the library will happily attempt it on a single machine. The repository ships benchmarks precisely because estimator performance varies enough to matter, but the README does not offer guidance on when a search becomes impractical.

## How it compares with a deep learning framework

The honest alternative for many users is PyTorch or TensorFlow, and the difference is architectural rather than a matter of quality. scikit-learn provides finished estimators: you choose an algorithm, set hyperparameters, and call fit. The library owns the training loop. A deep learning framework provides tensors with automatic differentiation and lets you write the training loop yourself, which is what you need when the model architecture is the research contribution.

That distinction decides most cases. Tabular regression, classification, clustering and dimensionality reduction with a few thousand to a few million rows are squarely scikit-learn territory, and the pipeline and cross-validation machinery has no real equivalent in a deep learning framework without extra libraries. Sequence models, image models, and anything requiring gradient descent over a custom architecture belong on the other side. The dependency footprint also differs: scikit-learn builds on NumPy and SciPy, while a deep learning framework pulls in its own tensor runtime and, for GPU work, vendor-specific acceleration libraries.

There is overlap. A deep learning framework can train a logistic regression, and scikit-learn can be used for the preprocessing and evaluation around a neural network. Treating them as mutually exclusive is unnecessary; treating them as interchangeable is a mistake.

## Maintenance, releases and the BSD-3-Clause licence

The repository is not archived, and the last push was on 2026-06-02, the same date as the 1.9.0 release. Before that, 1.8.0 landed on 2025-12-10 and 1.7.2 on 2025-09-09. The cadence is roughly two feature releases a year with patch releases between them, which gives you a predictable window for planning upgrades. The README states the project is maintained by a community of contributors with support from several organisations, and the changelog at https://scikit-learn.org/dev/whats_new.html is the authoritative list of notable changes.

Upgrade cost is mostly API deprecation. The changelog exists because behaviour changes between minor versions, and the project follows a deprecation cycle rather than breaking interfaces without warning. Pin your version in production and read the changelog before moving a minor version, particularly if you rely on estimators that have been reworked.

On licensing, the package is distributed under the 3-Clause BSD license, and pyproject.toml declares license = "BSD-3-Clause" with license-files = ["COPYING"]. That is a permissive licence, which generally means you can use, modify and redistribute the library, including in commercial products, provided the copyright notice and licence text are retained. This is a description of what the repository declares, not legal advice; if the licence terms matter to your organisation, have counsel read COPYING.

## Conclusion

Adopt scikit-learn when your data fits in memory as arrays or DataFrames and your models are regression, classification, clustering or dimensionality reduction. Skip it for deep neural networks, GPU training or out-of-core datasets, where PyTorch or TensorFlow is the better fit. Before committing, verify the Python 3.11 floor and the NumPy, SciPy, joblib, narwhals and threadpoolctl minimums in pyproject.toml against your environment, and read the changelog at https://scikit-learn.org/dev/whats_new.html for API changes between the 1.7, 1.8 and 1.9 lines.

## FAQ

### What is scikit-learn used for?

It is a Python module for classical machine learning: regression, classification, clustering, dimensionality reduction, preprocessing and model selection, all built on NumPy and SciPy. It is aimed at data that fits in memory as arrays or DataFrames rather than deep learning workloads.

### Which is better, TensorFlow or scikit-learn?

They answer different questions rather than competing directly. scikit-learn provides finished estimators with a uniform fit and predict interface and no deep learning implementation, while TensorFlow is a deep learning framework. The dependency list in pyproject.toml shows scikit-learn relies on numpy, scipy, joblib, narwhals and threadpoolctl, with no tensor runtime.

### Is scikit-learn still relevant in 2026?

The repository is not archived, and the last push was on 2026-06-02, matching the 1.9.0 release. Releases 1.8.0 and 1.7.2 arrived on 2025-12-10 and 2025-09-09 respectively, so the project continues to ship.

### Is sklearn obsolete?

The project is not archived and continues to publish releases, with 1.9.0 dated 2026-06-02. The name sklearn is the import name used in code, while the package is installed as scikit-learn.

### How to install scikit-learn?

The README gives two commands: pip install -U scikit-learn when NumPy and SciPy are already present, or conda install -c conda-forge scikit-learn. The package requires Python 3.11 or newer along with the NumPy, SciPy, joblib, narwhals and threadpoolctl minimums listed in pyproject.toml.

### How to use scikit-learn linear regression?

Construct LinearRegression, call fit with a feature matrix and target vector, then read the coefficients from the coef_ attribute. The examples/ directory includes a linear_model section with worked cases, and sklearn.linear_model is the module that holds the estimator.

## Sources

- [Official documentation](https://scikit-learn.org)
- [Official README](https://github.com/scikit-learn/scikit-learn#readme)
- [Project repository](https://github.com/scikit-learn/scikit-learn)
- [Release notes](https://github.com/scikit-learn/scikit-learn/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/scikit-learn-scikit-learn
