skrub: machine learning with dataframes, without hand-writing the cleaning pipeline
Machine learning with dataframes
At a glance
- What is it?
- skrub is a Python library that wraps pandas and scikit-learn so that messy dataframe columns can go straight into a model. The documentation is generous, the dependency list is light, and the hard part is knowing when its automatic encoders are the right call.
- Who is it for?
- skrub fits teams already using pandas and scikit-learn who spend their time writing per-column cleaning code before a model ever sees the data. It does not fit projects that need a dataframe engine other than pandas, or teams that want a fully declarative schema language.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem skrub targets: cleaning code that never gets reused
Most tabular machine learning work in Python starts the same way. You load a dataframe, look at the dtypes, and start writing per-column fixes: parse this date string, group those rare categories, impute the missing numeric values, one-hot the low-cardinality columns. That code lives in a notebook. When the next table arrives, you write it again.
skrub is aimed at that gap. The README describes it as "a Python library that facilitates machine learning with dataframes", and the repository topics list data-cleaning, data-preparation and data-wrangling alongside dataframes and machine-learning. The intended user is someone who already works in pandas and scikit-learn and wants the cleaning step to be a reusable object rather than a block of ad hoc code.
The package metadata classifies it as Production/Stable and targets Python 3.10 through 3.14. It is not a new dataframe engine. It sits on top of pandas and scikit-learn, which is a deliberate constraint: anything you already know about those two libraries still applies.
How skrub sits between pandas and scikit-learn
The dependency list in pyproject.toml tells most of the architecture story. skrub requires numpy, pandas, scikit-learn, scipy, jinja2, matplotlib, requests, pydot, cloudpickle and joblib. pandas and scikit-learn are the two load-bearing dependencies. jinja2 and pydot point at HTML and graph rendering, which is how the library reports what it did. joblib and cloudpickle point at parallel fitting and serialisation of fitted objects.
The data flow implied by that stack is: a pandas dataframe goes in, skrub inspects each column's dtype and contents, applies an encoder chosen for that column, and emits a numeric matrix that scikit-learn estimators accept. Because the output is a scikit-learn compatible transformer, it can live inside a Pipeline and be cross-validated like any other step.
What makes this more than a convenience wrapper is that the column-level decisions are made per column rather than globally. A date column and a free-text column do not need the same treatment, and skrub's job is to notice the difference from the data rather than from a hand-written list of column names.
The documentation ships inside the package. After installing, the README gives this snippet to locate it:
Installing skrub and running a first TableVectorizer
The README states that skrub installs via pip or conda and points at the installation page for details. The pip route is the one shown in the package metadata, which declares the distribution name as skrub.
pip install skrubAfter installation, the bundled documentation can be found from Python. The README gives this exact example, and it is a reasonable first check that the package landed correctly.
import skrub
print(skrub.__docs_dir__)The printed path is a directory on disk containing the docs and examples that shipped with your installed version. That matters because it means the examples match the version you have, not the version on the website.
For a first real use, the pattern is to hand a dataframe to skrub's vectorizer and let it decide per-column encodings, then pass the result to a scikit-learn estimator. The README does not include a full worked example, so the safest next step is to open the examples directory pointed at by skrub.__docs_dir__ and follow one of the notebooks there rather than guessing at constructor arguments.
One environment constraint to check before you start: skrub requires Python 3.10 or newer, and scikit-learn 1.4.2 or newer. On an older environment pip will either refuse or resolve to an older skrub release.
Where skrub stops being the right tool
The pandas dependency is the clearest boundary. If your data lives in Polars, DuckDB, Spark or a warehouse and never becomes a pandas dataframe, skrub's encoders have nothing to operate on. The pyproject dependencies contain no dataframe library other than pandas, so this is not a configuration question.
The second boundary is scale. pandas holds data in memory, and skrub's per-column inspection adds work on top of that. For tables that fit comfortably in memory this is fine. For tables that do not, the cleaning step is not the place to solve it.
The third boundary is control. Automatic per-column encoding is a trade-off: you get less code, and you also get a decision made from heuristics rather than from your knowledge of the domain. A column that looks numeric but encodes a category, or a column whose missing values carry meaning, may not be treated the way you intend. The library's rendering of what it did is what lets you check, and if you never look at that output you are trusting a heuristic you have not read.
Finally, the README does not document a rollback or migration path between minor versions. The CHANGES.rst file in the repository root is where release notes live, and pinning a version before upgrading is the conservative move.
skrub against writing your own ColumnTransformer
The obvious alternative is scikit-learn's own ColumnTransformer combined with SimpleImputer, OneHotEncoder and StandardScaler, wired together by hand. That approach is fully explicit: you name every column, choose every encoder, and nothing happens that you did not write.
The difference in approach is where the knowledge lives. With ColumnTransformer, the mapping from column to encoder is in your code, and it breaks when the schema changes. With skrub, the mapping is inferred from the data at fit time, so a new column with a recognisable type gets handled without an edit. You trade explicitness for adaptability.
There is a middle ground worth naming: keep ColumnTransformer for the columns where you have strong domain knowledge, and use skrub's vectorizer for the long tail of columns you would otherwise handle with a generic fallback. The two are scikit-learn transformers, so they compose in the same Pipeline.
The other alternative is to move off pandas entirely, for example to a columnar engine with its own expression API. That solves a different problem (scale and lazy execution) and does not give you the scikit-learn integration, so it is not a like-for-like swap.
Maintenance, releases and the BSD-3-Clause licence
The repository is not archived, and the last push was on 2026-09-10. Releases are frequent: 0.9.0 on 2026-05-06, 0.10.0 on 2026-07-06, and 0.10.1 on 2026-09-01. That cadence means minor versions arrive roughly every two months, so an upgrade is a recurring task rather than a one-off.
Upgrade cost is mostly the usual scikit-learn and pandas compatibility surface. skrub pins minimum versions of both (scikit-learn >= 1.4.2, pandas >= 1.5.3) rather than upper bounds, so a skrub upgrade can pull in newer versions of its dependencies. The dev extra in pyproject.toml is large and documentation-oriented, including sphinx and jupyterlite packages, and is not needed to use the library.
The licence is BSD-3-Clause, declared in pyproject.toml with license-files pointing at LICENSE.txt. That is a permissive licence, which in practice means you can use skrub in commercial and closed-source work provided you keep the copyright notice and licence text with redistributed copies. This is a description of the licence identifier, not legal advice; if your organisation has a policy on permissive licences, route it through the people who own that policy.
The README also gives a Zenodo DOI for citation if you use skrub in a scientific publication.
Editorial conclusion
skrub fits teams already using pandas and scikit-learn who spend their time writing per-column cleaning code before a model ever sees the data. It does not fit projects that need a dataframe engine other than pandas, or teams that want a fully declarative schema language. Before adopting it, check that your Python is 3.10 or newer and that scikit-learn is at least 1.4.2, then run TableVectorizer on one of your own tables and read the transformed column names to see what the automatic encoders decided.
Frequently asked questions
What does skrub mean as a Python package?
The repository describes skrub as a Python library that facilitates machine learning with dataframes, built on pandas and scikit-learn. The README does not give an expansion or etymology for the name.
How do I install skrub?
The README states that skrub can be installed via pip or conda and links to the installation instructions on the project website. The package name in pyproject.toml is skrub, and it requires Python 3.10 or newer.
What Python and scikit-learn versions does skrub require?
pyproject.toml declares requires-python >=3.10 and classifiers for Python 3.10 through 3.14, with scikit-learn >=1.4.2, pandas >=1.5.3 and numpy >=1.23.5 among the dependencies.
Where can I find skrub examples and documentation after installing?
Documentation and examples are bundled with the package in skrub/_docs. The README gives the snippet import skrub; print(skrub.__docs_dir__) to print the directory on your machine.
What licence is skrub released under?
skrub is released under BSD-3-Clause, declared in pyproject.toml with license-files pointing at LICENSE.txt. The repository also provides a Zenodo DOI for citation in scientific publications.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/skrub-data-skrub)