CV-CUDA: GPU Image Processing That Skips the Host Round Trip
CV-CUDA™ is an open-source, GPU accelerated library for cloud-scale image processing and computer vision.
At a glance
- What is it?
- CV-CUDA is NVIDIA's Apache-licensed library of GPU-accelerated image operators with C++ and Python APIs. It is aimed at cloud-scale AI pipelines, and its main constraint is that it only runs where CUDA runs.
- Who is it for?
- Adopt CV-CUDA if your preprocessing already runs on NVIDIA GPUs and you want decode, resize and color conversion to stay on device; the pip install cvcuda-cu12 wheel is the fastest way to check that. Skip it if you need native Windows, Volta-class hardware or a CUDA-free deployment, and verify your driver version against the compatibility table before committing.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 14 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Who CV-CUDA Is Actually For
The pitch is narrow and worth stating plainly: CV-CUDA is a library of GPU-accelerated computer vision algorithms built for throughput on NVIDIA hardware. The README describes it as delivering "high-throughput, low-latency image/video processing for AI pipelines" across cloud, desktop and edge platforms, and it exposes both a C/C++ and a Python API. If your inference server spends measurable time decoding JPEGs, resizing to model input dimensions and converting color spaces on the CPU before a tensor ever reaches the GPU, that is the workload this library targets.
It is not a general replacement for a CPU image library. There is no fallback path described in the README for machines without a supported NVIDIA GPU, and the compatibility table lists CUDA compute capability SM7.5 or higher as a floor. Teams running mixed hardware fleets, or preprocessing on ARM CPUs before shipping tensors to a remote GPU, get nothing from it. The relevant audience is the group already committed to NVIDIA accelerators end to end.
The Mechanism: Tensors, Operators and nvImageCodec
CV-CUDA organizes work around tensors and operators rather than around images as objects. The README's example shows the shape of a pipeline: an nvImageCodec decoder reads a JPEG directly into GPU memory, the result is wrapped with cvcuda.as_tensor using an "HWC" layout string, and cvcuda.resize then produces a new tensor at the target dimensions with cvcuda.Interp.LINEAR as the interpolation mode.
The detail that matters is where the data lives. Decoding to GPU, converting to a CV-CUDA tensor and resizing are all described as happening without a host round trip, which is the whole point of the design. The architecture diagram in the repository (docs/sphinx/content/cvcuda_arch.svg) is the place to check how operators are layered, and the online documentation is where the operator list lives; the README itself does not enumerate them. Treat the tensor layout string as a contract you have to get right, because the layout you declare is what the operator assumes when it indexes memory.
Installing CV-CUDA from PyPI and Running a First Resize
The README gives pre-built Python wheels on pypi.org for Python 3.10 through 3.14 on x86_64 and aarch64 Linux. There are two package names, one per CUDA major version. Pick the one that matches the CUDA toolkit on the machine:
pip install cvcuda-cu12If the target machine runs CUDA 13, the README lists the alternative as pip install cvcuda-cu13. Only one CUDA version of CV-CUDA packages can be installed at a time, so do not mix them in the same environment. After installation, the first real use is the decode, wrap, resize sequence from the README:
import cvcuda
from nvidia import nvimgcodec
decoder = nvimgcodec.Decoder()
image = decoder.read("input.jpg")
cvcuda_tensor = cvcuda.as_tensor(image, "HWC")
resized = cvcuda.resize(cvcuda_tensor, (224, 224, 3), cvcuda.Interp.LINEAR)What you should see is a tensor whose spatial dimensions are 224 by 224 with three channels, produced without an intermediate host buffer. Note that nvImageCodec is imported separately and is not part of the cvcuda package, so the decoder is a second dependency to account for.
For C++ work the README points at a different route. The repository ships Debian packages and tar archives alongside the wheels, and building from source is documented in the online installation guide. One change worth knowing before you try: CV-CUDA no longer uses git submodules. Build dependencies such as googletest, nvbench, dlpack and pybind11 are pre-installed in the Docker devel images described in docker/README.md and resolved through CMake's find_package, so running git submodule update --init finds nothing. Outside Docker you install those packages yourself or pull them in with CMake's FetchContent.
Where CV-CUDA Stops Being the Right Tool
The README is unusually direct about the boundaries. Native Windows is not supported; the only Windows path is WSL2, which rules the library out for teams shipping a Windows desktop product. CUDA 11, SM7 (Volta) and Ubuntu 20.04 lost official support starting with v0.16, so hardware that was fine two releases ago now sits outside the supported matrix.
The packaging split on aarch64 is a second trap. Starting with v0.14, the aarch64 packages on GitHub release assets and PyPI are the SBSA-compatible ones; Jetson builds live in explicitly named Jetson archives in the GitHub release assets. Installing the default aarch64 wheel on a Jetson Orin and expecting it to work is the mistake this note exists to prevent. Building from source for Orin has its own flag, -DCVCUDA_AARCH64_JETSON=ON, which restricts the build to Orin-relevant GPU architectures to cut build time.
There is also a testing caveat that affects anyone reading the test suite as a quality signal. The C++ test module builds with gcc 10 with partial coverage; full coverage needs gcc 11 with full C++20 support for NTTP. A build that compiles and passes on gcc 10 is not exercising everything. Finally, the CV-CUDA samples are only officially supported with CUDA 12, so a CUDA 13 setup should not expect the sample applications to be a supported path. The README does not document a rollback procedure for a failed wheel upgrade, and it does not describe CPU fallbacks for any operator.
CV-CUDA Against OpenCV's CUDA Module and Plain OpenCV
The comparison people reach for is OpenCV, and the difference is architectural rather than a matter of which has more functions. OpenCV is a CPU-first library with a GPU module bolted on; the cv::cuda::GpuMat type is a distinct container, and moving between cv::Mat and GpuMat is an explicit upload and download step you manage yourself. CV-CUDA starts from the GPU side. Its tensors are the primary representation, and the README's example never constructs a host image at all because the decoder writes straight to device memory.
That matters most when the pipeline is decode, preprocess, infer. With OpenCV's CUDA path, a typical flow is CPU decode, upload, resize on GPU, then hand the buffer to the inference runtime, and the upload is a synchronization point. With CV-CUDA plus nvImageCodec, the decode is already on device, so the boundary moves. The trade is ecosystem breadth: OpenCV covers far more operations and runs on CPUs and non-NVIDIA accelerators, while CV-CUDA is tied to the CUDA stack and the operator set documented online. If you need one obscure filter, check that operator list before planning a migration. If your bottleneck is the host round trip in a fixed preprocessing chain, the architectural difference is the reason to switch.
Licence and the Cost of Tracking Releases
The README carries Apache-2.0 headers and the repository's LICENSE.md is the file to read for the actual terms; the GitHub metadata reports the licence as NOASSERTION, which means the automated classifier did not recognize it, not that the project is unlicensed. Apache-2.0 includes an explicit patent grant, which is usually the reason a team picks it over a permissive licence without one. This is a description of what the files say, not legal advice, and anything redistributed in a product should go past whoever handles licensing at your organization.
The upgrade cadence is visible in the release history. v0.15.0 landed on 2025-05-19, v0.16.0 on 2025-11-15 and v0.17.0 on 2026-08-11, with the most recent push to the repository on the same day as v0.17.0. That is roughly two releases a year, which is slow enough that pinning a version and moving deliberately is reasonable, but the v0.16 notes show why you cannot ignore the changelog: a single release dropped CUDA 11, Volta and Ubuntu 20.04 support at once. Budget for reading release notes before every minor bump, and check the driver floor in the compatibility table. CUDA 12 x86_64 and aarch64 SBSA packages need driver r525 or newer, the samples need r535 or newer, Jetson Orin packages follow JetPack 6 and r535, and CUDA 13 needs r580 or newer. A driver that satisfies the library can still be too old for the samples.
Editorial conclusion
Adopt CV-CUDA if your preprocessing already runs on NVIDIA GPUs and you want decode, resize and color conversion to stay on device; the pip install cvcuda-cu12 wheel is the fastest way to check that. Skip it if you need native Windows, Volta-class hardware or a CUDA-free deployment, and verify your driver version against the compatibility table before committing.
Frequently asked questions
What is CV-CUDA?
It is an open-source library of GPU-accelerated computer vision algorithms from NVIDIA, with C/C++ and Python APIs. The README describes it as built for high-throughput, low-latency image and video processing in AI pipelines.
How do I install OpenCV with CUDA?
That is a different project. For CV-CUDA itself, the README lists two wheels on pypi.org: pip install cvcuda-cu12 for CUDA 12 and pip install cvcuda-cu13 for CUDA 13, and only one CUDA version can be installed at a time.
Does OpenCV require a GPU?
OpenCV does not require one, and neither does CV-CUDA's use case force the comparison: CV-CUDA itself has no described fallback path for machines without a supported NVIDIA GPU, and the compatibility table sets a floor of compute capability SM7.5.
What is better, OpenCL or CUDA?
The README does not compare the two. What it does say is that CV-CUDA is built on CUDA 12.2 or newer and CUDA 13.x, so choosing CV-CUDA means choosing the CUDA stack rather than a portable compute layer.
Does NVIDIA still use CUDA?
The repository answers this indirectly: CV-CUDA is an NVIDIA project whose compatibility table is organized entirely around CUDA versions, from CUDA 12.2 upward and CUDA 13.x, with driver floors of r525, r535 and r580 depending on the package.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/cvcuda-cv-cuda)