Library / SDK
rapidsai/cudf avatar
rapidsai/cudf

NVIDIA cuDF: GPU DataFrame Library for Pandas-Compatible and Polars Workloads

cuDF - GPU DataFrame Library. cudf.pandas With a Python file containing pandas code: Use cudf.pandas by invoking python with -m cudf.pandas If running the pandas code in an interactive Jupyter environment, call %load_ext cudf.pandas before importing pandas.

9,764 stars1,119 forksC++Apache-2.0

At a glance

What is it?
NVIDIA cuDF is an Apache-licensed GPU DataFrame library that provides a pandas-compatible Python API, a zero-code-change accelerator called cudf.pandas, and a GPU execution engine for Polars. It is part of the NVIDIA RAPIDS ecosystem and requires a CUDA-capable NVIDIA GPU.
Who is it for?
cuDF is the right tool for Python data workloads on tabular data that are bottlenecked by CPU pandas performance, specifically on NVIDIA GPU hardware. The cudf.pandas path offers the least friction: existing pandas scripts run faster without code changes by adding -m cudf.pandas to the Python command.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What cuDF Solves and Who It Targets

Pandas is the standard Python library for tabular data manipulation. It runs on the CPU and is single-threaded by default, which becomes a bottleneck when processing large datasets. GPU hardware offers parallel processing over thousands of cores, but using it for DataFrame operations requires either specialized code or a library that maps familiar operations to GPU-native kernels.

cuDF provides that mapping. The library exposes a DataFrame API that mirrors pandas, so engineers familiar with pandas can work with cuDF with minimal learning curve. The underlying computation runs on the GPU via CUDA, which the README describes as providing significant speed improvements for large tabular datasets. The README also notes that cuDF is part of the NVIDIA CUDA-X suite of GPU accelerated libraries.

The target audience is data engineers and machine learning practitioners who process large tabular datasets in Python and have NVIDIA GPU hardware available. Notable downstream users listed in the README include Spark RAPIDS (a GPU accelerator plugin for Apache Spark), Velox-cuDF (a Velox extension for GPU query execution), and Sirius DB (a GPU-native SQL engine).

Component Architecture: Five Libraries, One Ecosystem

The README describes cuDF as composed of multiple libraries with distinct layers. libcudf is a CUDA C++ library that provides Apache Arrow compliant data structures and fundamental algorithms for tabular data. It is the foundation on which the higher-level Python libraries are built.

pylibcudf provides Python Cython bindings directly to libcudf. It sits between libcudf and the higher-level cudf Python library, and the documentation at docs.nvidia.com/cudf covers it as a separate API surface for users who need lower-level control.

cudf is the main Python library with two access paths: a standalone DataFrame library with a pandas-like API (import cudf), and cudf.pandas, a zero-code-change accelerator that intercepts pandas calls and redirects them to cuDF when possible. cudf-polars provides a GPU engine for the Polars lazy API, activated by passing engine="gpu" to Polars' collect() call. dask-cudf provides a GPU backend for Dask DataFrames, enabling distributed GPU processing across multiple nodes.

Installing cuDF and Matching the CUDA Version

cuDF packages are available on PyPI with a CUDA version suffix. The suffix must match the major version of CUDA installed on the system. For CUDA 13:

bash
pip install libcudf-cu13
pip install pylibcudf-cu13
pip install cudf-cu13
pip install cudf-polars-cu13
pip install dask-cudf-cu13

For CUDA 12, replace cu13 with cu12 in each package name. A conda-based install via the rapidsai channel does not require manual CUDA version selection:

bash
conda install -c rapidsai cudf
conda install -c rapidsai cudf-polars
conda install -c rapidsai dask-cudf

The README refers to the RAPIDS Installation Guide at docs.nvidia.com/datascience/install for system requirements including supported operating systems, GPU driver versions, and CUDA versions. Building from source requires following the contribution guide, which documents a more involved environment setup.

cudf.pandas: Accelerating Existing Code Without Rewriting It

cudf.pandas is the most accessible entry point for engineers who already have working pandas code. It requires no API changes. Given an existing Python script that uses pandas:

python
import pandas as pd
df = pd.read_parquet("data.parquet")
df.dropna().groupby(["A", "B"]).mean()

Running it with the cudf.pandas module flag redirects the pandas operations to cuDF on the GPU:

bash
python -m cudf.pandas script.py

For Jupyter notebooks, the extension must be loaded before importing pandas:

python
%load_ext cudf.pandas
import pandas as pd

For direct cuDF usage without the compatibility layer, importing cudf provides the GPU DataFrame API directly:

python
import cudf
df = cudf.read_parquet("data.parquet")
df.dropna().groupby(["A", "B"]).mean()

For Polars users, cudf-polars integrates through Polars' existing lazy API. The engine parameter activates GPU execution:

python
import polars as pl
lf = pl.scan_parquet("data.parquet")
lf.drop_nulls().group_by(["A", "B"]).mean().collect(engine="gpu")

Hard Constraints: NVIDIA GPU and Unsupported pandas Operations

cuDF requires an NVIDIA CUDA-capable GPU. There is no CPU fallback in the cudf library itself. Teams running on AMD GPUs, Apple Silicon, or CPU-only instances cannot run cudf or cudf-polars. The RAPIDS Installation Guide specifies which NVIDIA GPU generations are supported.

cudf.pandas does not accelerate every pandas operation. When cuDF does not support a specific operation, cudf.pandas falls back to the CPU-based pandas implementation transparently. The README describes this as the zero-code-change guarantee: code continues to work even when a GPU operation is unsupported. The trade-off is that a script relying heavily on unsupported operations may not see meaningful GPU acceleration despite the overhead of the cudf.pandas interceptor.

The repository's Python layer is organized under python/cudf, python/pylibcudf, python/dask_cudf, and python/cudf_polars directories. The C++ layer is under cpp/. Building the C++ library from source requires a CUDA toolchain, CMake, and the build script documented in CONTRIBUTING.md.

cuDF vs pandas on CPU: When GPU Processing Wins

Pandas on CPU performs well for datasets that fit comfortably in RAM and involve straightforward operations like groupby aggregation, merge, or filter on a few million rows. For these sizes, the overhead of data transfer between CPU and GPU memory often offsets GPU parallelism gains.

GPU acceleration with cuDF becomes worthwhile when datasets are large enough that CPU-bound processing is measurably slow. The parallel architecture of a GPU provides advantages on operations that can be decomposed into independent per-row or per-column computations: groupby aggregation over hundreds of millions of rows, large-scale joins, and sorting over wide datasets are examples where the GPU's throughput exceeds what pandas achieves on a single CPU core.

For the dask-cudf path, datasets too large to fit in a single GPU's memory can be partitioned across multiple GPUs or a distributed cluster. This extends the viable dataset size beyond single-device limits. Dask-cudf is listed separately in the conda package index from cudf, so it installs independently for users who need the distributed path.

The repository's notebooks/ directory contains example Jupyter notebooks demonstrating cuDF usage on real datasets. The ci/ and conda/ directories contain the conda recipe and CI configuration, which document the exact package versions and platform constraints for each release. The CHANGELOG.md records changes per release for teams tracking version-to-version differences. The java/ directory provides a Java API for libcudf, covering teams that use JVM-based data pipelines with Spark RAPIDS.

Editorial conclusion

cuDF is the right tool for Python data workloads on tabular data that are bottlenecked by CPU pandas performance, specifically on NVIDIA GPU hardware. The cudf.pandas path offers the least friction: existing pandas scripts run faster without code changes by adding -m cudf.pandas to the Python command. Teams without NVIDIA GPUs, or running on CPU-only cloud instances, cannot use cuDF at all. Before adopting, verify that the CUDA version in the deployment environment matches one of the pip suffix variants (cu12 or cu13). The latest release is v26.08.01, published on 2026-08-25.

Frequently asked questions

How does cuDF work?

cuDF provides a pandas-like DataFrame API implemented in CUDA C++ (libcudf) and exposed to Python via Cython bindings (pylibcudf) and a higher-level Python layer. Data lives in GPU memory, and operations execute on GPU cores. The cudf.pandas module intercepts standard pandas API calls and redirects them to cuDF when the operation is supported, falling back to CPU pandas when it is not.

Can I use cuDF without an NVIDIA GPU?

No. cuDF requires a CUDA-capable NVIDIA GPU. There is no CPU fallback for the core cudf library. The RAPIDS Installation Guide at docs.nvidia.com/datascience/install specifies supported GPU models, driver versions, and CUDA versions.

How do I install cuDF?

Install via pip with a CUDA version suffix matching your system: pip install cudf-cu12 for CUDA 12 or pip install cudf-cu13 for CUDA 13. A conda install via conda install -c rapidsai cudf selects the correct CUDA version automatically. Nightly builds are available from the rapidsai-nightly channel.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/rapidsai-cudf.svg)](https://hysenlabs.com/projects/rapidsai-cudf)