Library / SDK
henrikbostrom/crepes avatar
henrikbostrom/crepes

crepes: conformal classifiers, regressors and predictive systems for scikit-learn models

Python package for conformal prediction

584 stars46 forksPythonBSD-3-Clause

At a glance

What is it?
crepes is a Python package that wraps any scikit-learn style classifier or regressor and turns its predictions into calibrated p-values, prediction sets and intervals. The wrapper API is small, but the exchangeability assumption behind the guarantees is the part that decides whether it fits your problem.
Who is it for?
Use crepes when you already have a fitted scikit-learn model and need prediction sets or intervals with a coverage guarantee, and you can hold out a calibration set that is exchangeable with the data you will score in production. Do not reach for it if you need a model, a training framework or a serving stack: it does none of those, and it assumes the underlying learner is already good enough that calibration is the remaining problem.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 84 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What crepes adds on top of a fitted scikit-learn model

A random forest or gradient boosted model returns a class label or a number. It does not tell you how often that answer is right, and its predicted probabilities are frequently miscalibrated. crepes addresses exactly that gap. The package implements conformal classifiers, regressors and predictive systems that wrap any standard classifier or regressor, converting the original predictions into p-values, cumulative distribution functions, prediction sets or prediction intervals with coverage guarantees. That phrasing comes from the README, and it is a fair summary of the scope.

The audience is narrower than "everyone doing machine learning". You need a model that already produces a score or a probability, a labelled calibration set drawn from the same distribution as the data you will score later, and a downstream decision that benefits from a set of plausible labels or a numeric range rather than a single point. If your problem is "which class is most likely", crepes is overhead. If your problem is "which classes cannot be ruled out at 99 percent confidence", it is the right shape of tool.

The wrap, calibrate, predict sequence

The mechanism is a two-stage split. The README's quickstart splits a dataset into a training and a test set, then splits the training set again into a proper training set and a calibration set. The underlying learner is fitted on the proper training set only. The calibration set is then used to compute non-conformity scores, which is what the calibrate method does.

Wrapping is done through WrapClassifier and WrapRegressor, which take an unfitted estimator as their argument. After fit and calibrate, the wrapper exposes predict_p for p-values (one column per class), predict_set for prediction sets at a given confidence, and evaluate for metrics when true labels are available. The README shows the output of predict_set as a list of arrays of class labels, or, with labels=False, as binary vectors indicating presence or absence of each class. That second form is the one to use if you are feeding the result into another numeric pipeline.

Mondrian conformal classifiers are the mechanism for conditional coverage. You pass a MondrianCategorizer, or any function, as the mc argument to calibrate, and the same categorisation is applied under the hood when prediction sets are generated for test objects. The README illustrates this with the model's own predicted labels as categories, and notes that setting class_cond=True produces the class-conditional variant, where categories come from the true labels. The cost is visible in the README's own evaluation output: the class-conditional classifier reports a lower error and a higher average set size than the standard one on the same test data. Conditional coverage is bought with larger sets, which is the trade-off to expect, not a defect.

The package also ships crepes.extras for standard difficulty estimates, non-conformity scores and Mondrian categories, and crepes.martingales for testing the exchangeability assumption. The README states that you can supply your own functions for the first three, so the built-in options are a starting point rather than a constraint.

Installing crepes and producing a first prediction set

The README gives two install paths, PyPI and conda-forge. The setup.py declares numpy, pandas and scipy as install requirements and python_requires of 3.10 or later, so the interpreter version is a real constraint rather than a suggestion.

bash
pip install crepes

or, if your environment is conda-based:

bash
conda install conda-forge::crepes

The README's quickstart fetches a dataset from OpenML and makes the two splits. Note the second split: the calibration set is carved out of the training data, not the test data.

python
from sklearn.datasets import fetch_openml
from sklearn.model_selection import train_test_split

dataset = fetch_openml(name="qsar-biodeg", parser="auto")

X = dataset.data.values.astype(float)
y = dataset.target.values

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.5)

X_prop_train, X_cal, y_prop_train, y_cal = train_test_split(X_train, y_train,
                                                            test_size=0.25)

Wrapping and calibrating is two calls after the fit. The wrapper takes the estimator, not a fitted model, and the README fits it to the proper training set before calibrating.

python
from crepes import WrapClassifier
from sklearn.ensemble import RandomForestClassifier

rf = WrapClassifier(RandomForestClassifier(n_jobs=-1))

rf.fit(X_prop_train, y_prop_train)

rf.calibrate(X_cal, y_cal)

After that, predict_p returns an array with as many columns as there are classes, and predict_set returns sets at a chosen confidence. The README shows confidence=0.99 producing a mix of single-label and two-label sets on the test data, which is the expected behaviour: objects near a decision boundary get larger sets. To get a numeric matrix instead of label arrays, pass labels=False.

python
rf.predict_set(X_test, labels=False, confidence=0.99)

If you have the true labels, evaluate returns a dictionary. The README's example, at 99 percent confidence, contains an error rate, average set size under avg_c, the fraction of singleton sets under one_c, the fraction of empty sets, a Kolmogorov-Smirnov statistic, and fit and evaluation timings. Read avg_c and one_c together: they tell you whether the sets are actually useful or whether the guarantee is being met by returning everything.

The exchangeability assumption is the load-bearing wall

Conformal guarantees rest on exchangeability between the calibration data and the data being scored. crepes does not remove that requirement, and it cannot detect a violation for you unless you look. The package provides crepes.martingales for testing the assumption, which is a useful signal that the authors treat this as a first-class concern rather than a footnote, but a test is not a fix. If your production traffic drifts away from the calibration distribution, the coverage number you reported stops meaning what you think it means, and nothing in the wrapper API will raise an error.

There is a second limitation that the README states plainly rather than hiding. For an inductive conformal predictor, the predicted p-values and the errors made on a test set are not independent, because the calibration set is fixed. The documented remedy is semi-online operation: pass online=True to the predict methods and supply the true labels, and the calibration set is updated immediately after each prediction. That only works if labels arrive promptly. In a setting where ground truth shows up days later, semi-online calibration is not available, and you are back to a fixed calibration set with dependent errors across the batch.

A third consideration is set size. The README's own numbers show that the class-conditional variant increases average set size relative to the standard classifier. If a downstream consumer can only act on a single label, larger sets are not a better answer, they are a different kind of unusable. crepes is the wrong tool when the decision layer needs one answer and cannot express uncertainty.

How crepes differs from MAPIE and from a probabilistic classifier

MAPIE is the closest well-known alternative in the Python ecosystem, and the difference is in scope and in the shape of the API. MAPIE is built around scikit-learn's own estimator and cross-validation conventions, exposing conformal prediction through wrappers that plug into pipelines and model selection. crepes takes a different route: it introduces its own WrapClassifier and WrapRegressor classes with an explicit fit-then-calibrate sequence, and it adds predictive systems that return cumulative distribution functions rather than only intervals or sets. If you want conformal prediction to feel like another scikit-learn estimator, MAPIE's design will be more familiar. If you want the calibration step to be an explicit, separate object you control, and you want CDF output for regression, crepes is the more direct fit.

The second comparison is with simply reading the probabilities off a calibrated classifier. Platt scaling or isotonic regression adjusts the numbers a model outputs; it does not give you a finite-sample coverage guarantee on a set-valued prediction. That distinction matters when the output feeds a decision with a stated confidence level. It does not matter when you only need better-ranked probabilities for a ranking task, in which case a calibration layer is less machinery for the same practical result.

Maintenance, licence and the cost of upgrading

The repository is not archived. The last push was on 2026-07-08, and the most recent release listed is v0.9.1 from 2026-06-12, following v0.9.0 in October 2025 and v0.8.0 in March 2025. That is a steady release rhythm rather than a burst, and the gap between minor versions is measured in months, so an upgrade is a planned activity rather than something that happens to you.

Upgrade cost is low in the common case because the surface is small: a wrapper class, a calibrate call, and a handful of predict and evaluate methods. The risk sits in the transition between minor versions, where the CHANGELOG.md at the repository root is the place to look before bumping. The package tracks the output of evaluate as a dictionary of named metrics, so any change to those keys will break code that indexes them by name rather than reading them defensively.

The licence is BSD-3-Clause, declared in setup.py's classifiers as an OSI-approved BSD licence and shown in the README badge. That is a permissive licence, which in practice means you can use the package in commercial and closed-source work provided you keep the copyright notice and licence text with redistributed copies, and you do not use the author's name to endorse derived products. This is a description of the licence, not legal advice; read the LICENSE file at the repository root and, if the stakes are high, have counsel read it too. Note that the dependencies are separate works under their own licences, and BSD-3-Clause on crepes says nothing about them.

Editorial conclusion

Use crepes when you already have a fitted scikit-learn model and need prediction sets or intervals with a coverage guarantee, and you can hold out a calibration set that is exchangeable with the data you will score in production. Do not reach for it if you need a model, a training framework or a serving stack: it does none of those, and it assumes the underlying learner is already good enough that calibration is the remaining problem. Before adopting it, check three things in the repository: that the latest release on PyPI matches the version declared in setup.py, that the calibration split is large enough for the confidence level you intend to report, and that the exchangeability assumption survives your data collection process, since the package offers crepes.martingales to test it but cannot repair it.

Frequently asked questions

What is crepes and what does it do?

crepes is a Python package for conformal prediction that wraps any standard classifier or regressor and turns its predictions into calibrated p-values, cumulative distribution functions, prediction sets or intervals with coverage guarantees. It implements standard and Mondrian conformal classifiers, standard, normalized and Mondrian conformal regressors and predictive systems.

How do I install crepes?

The README gives two options: pip install crepes from PyPI, or conda install conda-forge::crepes from conda-forge. The package requires Python 3.10 or later and pulls in numpy, pandas and scipy.

Does crepes train the underlying model for me?

No. You pass an estimator such as RandomForestClassifier to WrapClassifier, then call fit on the proper training set and calibrate on a separate calibration set. crepes handles the conformal layer; the learner and its hyperparameters remain your responsibility.

How does crepes handle class-conditional coverage?

By setting class_cond=True in the call to calibrate, which forms a Mondrian conformal classifier whose categories are the true labels. More generally, you can pass any function or a MondrianCategorizer as the mc argument to calibrate to define your own categories.

Can crepes update its calibration set as new labels arrive?

Yes, for semi-online conformal predictors. Setting online=True when calling the prediction methods, while also providing the true labels, updates the calibration set immediately after each prediction, which makes the p-values and errors for a test set independent rather than dependent as they are for an inductive conformal predictor.

Official sources

  1. henrikbostrom/crepes on GitHub
  2. Issues
  3. License: BSD-3-Clause
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/henrikbostrom-crepes.svg)](https://hysenlabs.com/projects/henrikbostrom-crepes)