NVIDIA cuML: running scikit-learn style machine learning on the GPU
NVIDIA cuML: GPU-Accelerated Machine Learning
At a glance
- What is it?
- cuML offers two paths, a GPU-native estimator API and a cuml.accel layer that redirects existing scikit-learn, UMAP and HDBSCAN code to the GPU. The interesting question is what happens when an estimator cannot be accelerated.
- Who is it for?
- Adopt cuML if you already have NVIDIA GPUs in the loop and your workload is one of the supported clustering, dimensionality reduction, regression, classification, preprocessing, model selection, time series, explanation or nearest-neighbor estimators. Do not adopt it as a general replacement for scikit-learn on CPU-only machines, and do not assume a script run through cuml.accel is fully GPU-resident.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 12 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
The two entry points and the problem each one solves
cuML addresses a narrow but expensive problem: a machine learning pipeline that fits comfortably in scikit-learn's API but does not fit comfortably in the time budget, because the CPU is doing the arithmetic. The library is a CUDA-X Data Science component, licensed Apache-2.0, and it ships two distinct ways to move that work onto an NVIDIA GPU.
The first is the `cuml` Python API. It provides GPU-native estimators that follow the fit-predict-transform pattern scikit-learn users already know, and it keeps both data and computation on the GPU. The README's example is a DBSCAN clustering run built on `cuml.datasets.make_blobs` and `cuml.cluster.DBSCAN`. Coverage spans clustering, dimensionality reduction, regression, classification, preprocessing, model selection, time series, model explanation and nearest-neighbor workflows, with the current estimator list in the API reference rather than in the README.
The second is `cuml.accel`, aimed at a different reader: someone with working scikit-learn, UMAP or HDBSCAN code who does not want to rewrite it. That is the more interesting of the two, because it changes the cost of trying cuML from a port to a command-line flag.
How cuml.accel decides what runs on the GPU
`cuml.accel` is not a compiler and not a rewrite of scikit-learn. It is a redirection layer. You run a script through the module, or load the extension in a notebook before importing scikit-learn, UMAP or HDBSCAN, and supported operations are routed to the GPU. The README is explicit that this is a per-operation decision, not an all-or-nothing switch: when an estimator or a configuration cannot be accelerated, the original CPU implementation runs instead, so the rest of the workflow continues.
That fallback is the design's best and worst property at the same time. It means a partially supported pipeline still produces correct results, which is why the approach is safe to try. It also means a script can look accelerated while spending most of its wall clock on the CPU, and nothing in the output tells you so by default. The README points to two places to resolve that ambiguity: the compatibility documentation, which lists current coverage and the conditions that trigger fallback, and the logging and profiling tools, which report which operations actually ran on the GPU. Anyone evaluating this should treat those two pages as required reading, not optional.
The performance claim in the README is worth quoting precisely rather than paraphrasing: on representative benchmarks, cuML can accelerate scikit-learn workflows by up to 50x, with the caveat that performance depends on the algorithm, dataset and hardware. The benchmarks page carries the results and methodology. An "up to" figure from a vendor's own methodology is a starting point for your own measurement, not a planning number.
Installing cuML and running a first script
The README does not print install commands inline. It sends you to the RAPIDS installation selector at docs.rapids.ai/install, which generates a command for installing nightly or release cuML packages with conda, pip or Docker. Use the selector rather than copying a package line from a blog post, because the correct channel and version differ per CUDA release. Building from source is possible and is documented separately in BUILD.md; expect a C++/CUDA build and a much longer path than the package install.
Once installed, the GPU-native API works like this. The README gives this example, which generates sample data and computes DBSCAN clusters on the GPU:
from cuml.datasets import make_blobs
from cuml.cluster import DBSCAN
# Create sample data
X, y = make_blobs(n_samples=100, centers=3, n_features=2, random_state=42)
# Fit clustering model
dbscan = DBSCAN(eps=1.0, min_samples=5)
dbscan.fit(X)
print(dbscan.labels_)The reader should see cluster labels printed, with the estimator object behaving like its scikit-learn counterpart. Note the constructor arguments: `eps` and `min_samples` are passed the same way you would in scikit-learn.
The second path needs no rewrite at all. To run an existing script through the accelerator, the README gives this command:
python -m cuml.accel script.pyIn a Jupyter notebook, the README shows loading the extension before importing scikit-learn, UMAP or HDBSCAN:
%load_ext cuml.accelAfter either form, check the logging and profiling output to confirm which operations were redirected. That step is what separates a real GPU run from a script that quietly fell back.
Where cuML is the wrong tool
The fallback behaviour has a cost that the README states plainly in the compatibility documentation rather than in the quickstart: acceleration is per estimator and per configuration. A pipeline built on estimators outside the supported set will run correctly and slowly, and the only signal is the logging and profiling output. If your work is dominated by an algorithm cuML does not cover, adding `cuml.accel` buys you nothing except an extra dependency and a debugging surface.
There are two harder boundaries. The first is hardware. This is a CUDA-X library; the stated requirement is NVIDIA GPUs. Teams on CPU-only infrastructure, or on accelerators from other vendors, are outside the target. The second is API compatibility. The README states that cuML is compatible with scikit-learn version 1.6 or higher. A codebase pinned to an older scikit-learn cannot simply switch the accelerator on; the version constraint has to be resolved first, and that may be the real cost of the migration rather than the GPU work.
Serialization deserves a separate warning, and the README gives one. cuML models can be serialized with `pickle` or `joblib` and loaded later for inference, and cuML uses cloudpickle so that models trained with `cuml.accel` can be loaded and used with scikit-learn. The README then states, in bold, that only trusted sources should be unpickled, because `pickle` and by extension `joblib` can execute arbitrary code during deserialization. That applies to `pickle.load()`, `pickle.loads()`, `joblib.load()` and any file-based model loading. This is a property of Python's serialization format, not a defect cuML introduced, but it constrains any deployment where model artifacts cross a trust boundary.
cuml.accel versus the native cuml API, and versus staying on CPU
The most useful comparison is internal. `cuml.accel` and the `cuml` API are not competing products; they are two answers to the same question at different levels of commitment. `cuml.accel` is a probe. It costs one command or one notebook line, it preserves your existing code, and it degrades gracefully. The native `cuml` API is a port. You import `cuml.cluster`, `cuml.datasets` and their siblings, you accept the estimator set the API reference lists, and in exchange you get direct control over the workflow instead of a redirection layer you have to profile to understand.
The practical sequence is to run `cuml.accel` first, read the profiling output, and only port the estimators that actually dominated runtime. Porting a pipeline whose bottleneck turned out to be an unsupported estimator is wasted effort.
The external comparison is with scikit-learn on CPU. The difference is not API shape, since cuML deliberately mirrors it, but where the arrays live and what happens when the GPU cannot help. scikit-learn has no fallback concept because it has nothing to fall back from. cuML's fallback is a feature for migration and a hazard for measurement. If your datasets are small, the transfer overhead and the profiling work may exceed anything you gain, and the README's own framing of the speedup as dependent on algorithm, dataset and hardware is the honest signal there.
For work that outgrows one device, `cuml.dask` provides distributed implementations of selected algorithms for multi-GPU and multi-node execution with Dask. The word to hold onto is "selected": the multi-GPU guide documents supported algorithms, and it is a subset, not the whole API.
Maintenance cadence, release train and licence
The repository is not archived and the last push was on 2026-09-10, which places it inside an active release train rather than at the end of one. Releases follow a dated scheme: v26.08.00 on 2026-08-05, v26.06.00 on 2026-06-04, v26.04.00 on 2026-04-09. That is roughly a two-month cadence, and the documentation URLs embed the release identifier, so a link to `docs.nvidia.com/cuml/26.08/...` is versioned to a specific release. Plan for documentation that moves with the release rather than a single evergreen page.
That cadence is the main upgrade cost. Versioned docs mean bookmarked links and internal runbooks need updating each cycle, and the scikit-learn 1.6 or higher floor means a cuML upgrade can be blocked by an unrelated scikit-learn pin. The repository carries a CHANGELOG.md, a dependencies.yaml and a RAPIDS_BRANCH file at the top level, which is where to check what moved between releases before upgrading.
The licence is Apache-2.0, stated in the repository and in the SPDX headers of files such as pyproject.toml. Apache-2.0 is a permissive licence with an explicit patent grant, which matters for a library that implements published algorithms. It does not resolve the separate question of the NVIDIA CUDA toolkit and driver terms your deployment depends on, and it says nothing about the pickle security issue above. Read the licence text and your own distribution obligations rather than treating permissive as equivalent to unencumbered; this is a description of the licence identifier, not legal advice.
Editorial conclusion
Adopt cuML if you already have NVIDIA GPUs in the loop and your workload is one of the supported clustering, dimensionality reduction, regression, classification, preprocessing, model selection, time series, explanation or nearest-neighbor estimators. Do not adopt it as a general replacement for scikit-learn on CPU-only machines, and do not assume a script run through cuml.accel is fully GPU-resident. Before committing, run the compatibility check on your own estimator list, turn on the logging and profiling tools to see which operations actually ran on the GPU, and confirm your scikit-learn version is 1.6 or higher.
Frequently asked questions
How do I install NVIDIA cuML?
The README does not list install commands directly. It points to the RAPIDS installation selector at docs.rapids.ai/install, which generates a command for installing nightly or release cuML packages with conda, pip or Docker. Building from source is covered separately in BUILD.md.
Can scikit-learn be used on GPUs with NVIDIA cuML?
Yes, through cuml.accel. Running a script with python -m cuml.accel script.py, or loading the extension in a notebook before importing scikit-learn, UMAP or HDBSCAN, routes supported operations to the GPU. When an estimator or configuration cannot be accelerated, the original CPU implementation is used so the workflow can continue.
Does NVIDIA cuML require a specific scikit-learn version?
The README states that cuML is compatible with scikit-learn version 1.6 or higher. That floor applies before you can use cuml.accel against an existing codebase.
How can I tell whether NVIDIA cuML actually ran an operation on the GPU?
The README points to the cuml.accel logging and profiling tools, which report which operations ran on the GPU. It also refers to the cuml.accel compatibility documentation for current coverage and the conditions that trigger fallback to CPU.
Can NVIDIA cuML models be saved and loaded later?
The README states that cuML models can be serialized with pickle or joblib and loaded later for inference, and that cuML uses cloudpickle so models trained with cuml.accel can be loaded and used with scikit-learn. It warns that pickle and joblib are not secure and that models should only be loaded from trusted sources.
Does NVIDIA cuML work across more than one GPU?
The cuml.dask API provides distributed implementations of selected algorithms for multi-GPU and multi-node execution with Dask. The README points to the multi-GPU guide for cluster setup, supported algorithms and examples.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvidia-cuml)
Community notes