Framework
zama-ai/concrete-ml avatar
zama-ai/concrete-ml

Concrete ML: FHE Inference for scikit-learn and PyTorch Models

Concrete ML: Privacy Preserving ML framework using Fully Homomorphic Encryption (FHE), built on top of Concrete, with bindings to traditional ML frameworks.

1,452 stars201 forksPythonNOASSERTION

At a glance

What is it?
Concrete ML turns quantized scikit-learn and PyTorch models into fully homomorphic equivalents you can run on encrypted inputs. The API is familiar; the compile step and the licence are where adoption decisions get made.
Who is it for?
Adopt Concrete ML when your inference inputs are the sensitive part and you can accept a quantization and compile step before deployment; it is a poor fit for training on encrypted data or for Python 3.13 environments. Before committing, check that your model family appears in the built-in model list, that your OS is covered by the installation table, and that the BSD-3-Clause-Clear terms are acceptable to your legal team.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 57 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Concrete ML Solves, and for Whom

Fully homomorphic encryption lets a server compute on ciphertext without holding the key. The catch is that FHE circuits are built from additions and multiplications over quantized integers, so a plain floating-point scikit-learn model does not translate directly. Concrete ML exists to close that gap. The README describes it as a set of tools built on top of Concrete that lets data scientists "automatically turn machine learning models into their homomorphic equivalents, and use them without knowledge of cryptography."

The audience is therefore narrow and specific. It is not a library for cryptography engineers who want to hand-write FHE circuits; those users would work with Concrete directly. It is for a data scientist who already has a logistic regression, a tree ensemble or a small neural network and needs inference to happen on inputs the server never sees. Concrete ML's built-in models expose an interface equivalent to their scikit-learn and XGBoost counterparts, so the migration path is mostly a change of import and an extra compile call. If your problem is training on encrypted data rather than inference, the README does not present that as the primary path, and you should not assume it.

The Mechanism: Quantize, Fit, Compile, Predict in FHE

The data flow has four stages, and the compile stage is the one that surprises newcomers. First, the model is constructed with a quantization parameter, for example n_bits=8, which fixes the integer precision used for weights and activations. Second, fit runs in the clear on plaintext training data, exactly as in scikit-learn. Third, compile takes a representative data set and traces the model into an FHE-compatible circuit, which is where the expensive work happens. Fourth, predict with fhe="execute" runs that circuit on encrypted inputs and returns decrypted predictions.

Two consequences follow from this design. The compile step needs representative data, so you must be able to hand the compiler a sample that looks like production traffic; the README's example simply passes X_train. And predictions in FHE are not guaranteed to match the clear-text predictions exactly, because quantization is lossy. The README's own example prints a similarity score between the two, which is the honest way to present it: you measure agreement rather than assume it. For custom models, the route is different: the README states that models using quantization-aware training are developed by the user in PyTorch or keras/tensorflow and imported into Concrete ML through ONNX.

Installing Concrete ML and Running a First Encrypted Prediction

The installation table in the README is the first thing to check, because pip support is not universal. Linux, Windows Subsystem for Linux, macOS 11+ on Intel and macOS 11+ on Apple Silicon all support pip. Native Windows does not; the README lists Docker as the route there. Concrete ML supports Python 3.8 through 3.12 only.

On a supported platform, the pip path is two commands. The first upgrades the build tooling, the second installs the package:

bash
pip install -U pip wheel setuptools
pip install concrete-ml

The Docker alternative pulls the published image, which is also the way to get a working environment on native Windows:

bash
docker pull zamafhe/concrete-ml:latest

With the package installed, the README gives this example, which is close to scikit-learn except for the quantization parameter and the compile call:

python
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from concrete.ml.sklearn import LogisticRegression

x, y = make_classification(n_samples=100, class_sep=2, n_features=30, random_state=42)
X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.2, random_state=42)

model = LogisticRegression(n_bits=8)
model.fit(X_train, y_train)
y_pred_clear = model.predict(X_test)
model.compile(X_train)
y_pred_fhe = model.predict(X_test, fhe="execute")

What you should see is printed output for the clear predictions, the FHE predictions, and a similarity percentage between them. In the README's run the two prediction arrays match and the similarity prints as 100%. Treat that number as something to measure on your own data, not a property of the library.

Where Concrete ML Is the Wrong Tool

The compile step is the main practical constraint. It runs on a representative sample and it is not free, which means a model that is retrained frequently will need a fresh compile each time, and a model whose input distribution drifts away from the representative set is being compiled against the wrong circuit. If your deployment pattern is continuous retraining, that loop is awkward.

Quantization is the second constraint, and it is a modelling constraint rather than a tooling one. Setting n_bits is a precision decision, and the README's similarity print exists precisely because FHE output can diverge from clear output. A task that needs exact floating-point behaviour, or where a single misclassified case is unacceptable, is a bad candidate. A third boundary is training: the documented workflow fits in the clear. If your requirement is that the training data itself never leaves the data owner's premises unencrypted, Concrete ML's documented path does not give you that.

Finally, the platform matrix rules some users out before any modelling starts. Native Windows without WSL is pip-unsupported, and Python 3.13 is outside the declared range in pyproject.toml, which pins python to ">=3.8.1,<3.13". A team standardised on 3.13 must change interpreter or containerise.

Concrete ML Versus Working With Concrete Directly

The real alternative is not another privacy-preserving ML library; it is dropping one level down to Concrete, the compiler Concrete ML is built on. The difference is where the abstraction sits. Concrete presents an FHE compiler: you express computation in a form the compiler can lower into FHE circuits, and you are responsible for how the model's arithmetic is expressed. Concrete ML presents model classes. You pick LogisticRegression or a tree-based model from concrete.ml.sklearn, set n_bits, and let the library handle the lowering.

That trade is legibility against control. Concrete ML gives you a scikit-learn-shaped object and a compile call, at the cost of accepting the library's quantization scheme and its set of supported model families. Concrete gives you the ability to shape the arithmetic yourself, at the cost of knowing what the compiler needs. If your model does not fit the built-in list and cannot be expressed as a quantization-aware PyTorch or keras model exported to ONNX, Concrete ML's documented import path does not cover you, and the lower-level tool is the honest answer.

Maintenance, Upgrades and the BSD-3-Clause-Clear Licence

The repository is not archived, and the last push was on 2026-08-04, which places it inside the six-month window. The release cadence is visible in the tags: v1.7.0 in September 2024, v1.8.0 in January 2025, v1.9.0 in April 2025. The README links an UPGRADING.md file at the repository root, which is where a version-to-version migration would be documented; the README itself does not describe rollback or downgrade behaviour, so pin your version before upgrading.

Upgrade cost is driven by the dependency pins rather than by the model API. pyproject.toml pins concrete-python to version 2.10.0 from the zama-pypi-cpu source and concrete-ml-extensions to 0.1.9. Because the FHE backend is pinned rather than ranged, a Concrete ML upgrade can move the compiler underneath you, and any change in compiled-circuit behaviour will show up as a change in the clear-versus-FHE similarity score. Re-run that comparison after every upgrade.

The licence is BSD-3-Clause-Clear, per pyproject.toml and the badge in the README. The repository metadata reports the licence as NOASSERTION, so the classifier and the file disagree at the metadata level and the LICENSE file is the authority. This is a source-available licence with patent-related terms rather than a plain permissive one, which is a question for your legal team, not for this article.

Editorial conclusion

Adopt Concrete ML when your inference inputs are the sensitive part and you can accept a quantization and compile step before deployment; it is a poor fit for training on encrypted data or for Python 3.13 environments. Before committing, check that your model family appears in the built-in model list, that your OS is covered by the installation table, and that the BSD-3-Clause-Clear terms are acceptable to your legal team.

Frequently asked questions

What is Concrete ML?

It is a privacy-preserving machine learning toolkit built on top of Concrete by Zama, which converts machine learning models into homomorphic equivalents so inference can run on encrypted inputs. Its model classes follow scikit-learn and XGBoost interfaces, and PyTorch models can be converted through ONNX.

How do I install Concrete ML?

On Linux, Windows Subsystem for Linux, or macOS 11+ you can run pip install concrete-ml after upgrading pip, wheel and setuptools. Native Windows is not supported by pip, so the README points to the Docker image zamafhe/concrete-ml:latest instead.

Which Python versions does Concrete ML support?

The README states support for Python 3.8, 3.9, 3.10, 3.11 and 3.12, and pyproject.toml pins the interpreter to >=3.8.1,<3.13. Python 3.13 is outside that range.

Does Concrete ML run on GPU?

The README's installation table covers Docker and pip per operating system and does not list GPU support as an installation option, and the dependency section pins the CPU build of Concrete Python. GPU execution is not documented in the README.

Do FHE predictions in Concrete ML match the clear-text predictions?

Not necessarily. The README's example prints a similarity percentage between the clear and FHE prediction arrays, which exists because quantization to n_bits makes the two paths differ. Measure that agreement on your own data rather than assuming an exact match.

Official sources

  1. Issues
  2. README
  3. Releases
  4. zama-ai/concrete-ml on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/zama-ai-concrete-ml.svg)](https://hysenlabs.com/projects/zama-ai-concrete-ml)