# tensorrt-cpp-api: A C++20 TensorRT Wrapper With Safe Caching and Zero-Copy Python Bindings

> tensorrt-cpp-api is a MIT-licensed C++20 library that wraps NVIDIA TensorRT 10+ for CNN inference on Linux with CUDA 12. It adds safe content-hash-keyed engine caching, a thread-safe EnginePool for concurrent inference, dynamic shape profiles, and optional Python bindings that let CuPy and PyTorch arrays pass directly to the engine without host round-trips. Windows and LLM-style transformer models are explicitly out of scope.

**cyrusbehr/tensorrt-cpp-api** — TensorRT C++ API Tutorial

- Repository: https://github.com/cyrusbehr/tensorrt-cpp-api
- Stars: 814 · Forks: 106
- Language: C++
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/cyrusbehr-tensorrt-cpp-api

## What tensorrt-cpp-api Solves: TensorRT Without Boilerplate or Leaky Abstractions

Using TensorRT's raw C++ API requires managing engine build-or-load logic, input and output tensor bookkeeping, CUDA stream ownership, concurrent execution context pooling, and error propagation. The API exposes nvinfer1 types in headers, which means any consuming project must link TensorRT at compile time, not just at runtime.

tensorrt-cpp-api addresses these points. The public headers contain no nvinfer1 types, no OpenCV types, and no spdlog types; consumers link TensorRT at runtime rather than compile time. The error model uses a Status/Result<T> pair with no exceptions, which is consistent with C++ codebases that avoid exception-based error handling. Engine caching is keyed by the ONNX file's content hash combined with build options, TensorRT version, and GPU UUID, so a stale cache (changed model, different driver, or different GPU) is detected and rebuilt rather than silently reused.

The library targets TensorRT 10 and above, CUDA 12, C++20, and Linux. Those constraints are hard: Windows support and LLM-specific features are explicitly described as out of scope.

## Safe Engine Caching, Dynamic Shapes, and Thread-Safe Concurrency

The engine cache uses a JSON sidecar file and atomic writes. When buildAndLoad is called, it checks whether a cached engine exists that matches the current ONNX content hash, build options, TensorRT version, and GPU UUID. If any of those four factors differ from what the cached engine was built with, the cache is invalidated and a new engine is built. This prevents the silent model-serving errors that can occur when a cached engine from a previous TensorRT version or a different GPU is loaded without validation.

Dynamic shapes are handled through per-input min/opt/max optimization profiles with -1-aware shape inference. Each execution context gets one optimization profile, which supports concurrent dynamic-shape inference without context conflicts.

The EnginePool class leases execution contexts across threads. The pool is thread-safe; the underlying engine is described as thread-compatible. Multi-stream inference is supported by passing a caller-provided Stream to each call:

```cpp
#include <tensorrt_cpp_api/all.h>
using namespace trtcpp;

Stream stream;
auto engine = EngineBuilder{}.buildAndLoad("model.onnx", opt);
```

Stream ownership is explicit: a new Stream allocates a CUDA stream, or Stream::wrap takes an existing handle.

## Building and Installing the Library

TensorRT and CUDA are system-provided dependencies, not managed by the build. The build takes three commands:

```sh
cmake -S . -B build -DTRT_CPP_API_BUILD_PREPROC=ON
cmake --build build -j$(nproc)
cmake --install build --prefix /opt/trtcpp
```

The flag -DTRT_CPP_API_BUILD_PREPROC=ON enables the optional fused preprocessing kernel. For a tarball TensorRT installation, add -DTensorRT_DIR=<root> to the first command to point CMake at the tarball location.

After installation, downstream projects consume the library with standard CMake find_package:

```cmake
find_package(tensorrt_cpp_api REQUIRED)
target_link_libraries(myapp PRIVATE tensorrt_cpp_api::tensorrt_cpp_api tensorrt_cpp_api::preproc)
```

The README documents the full build options in docs/install.md, including the apt versus tarball TensorRT paths. An upgrade guide from v6 is in docs/upgrading_from_v6.md. The project uses pre-commit hooks for clang-format and cmake-format; contributors run `pre-commit install` to set those up.

## Building an FP16 Engine and Running Inference

The core workflow is constructing BuildOptions, calling buildAndLoad, and running inference on a Stream. The README's example builds an FP16 engine from an ONNX file and loads from the on-disk cache if one already exists:

```cpp
BuildOptions opt;
opt.precision = Precision::kFp16;
opt.engineCacheDir = "engines";
auto engine = EngineBuilder{}.buildAndLoad("model.onnx", opt);
if (!engine) {
    std::fprintf(stderr, "%s\n", engine.status().message().c_str());
    return 1;
}
```

The Result<T> error model means no exceptions are thrown. Callers check the returned result and retrieve the error message from the status.

The optional preprocessing module (tensorrt_cpp_api::preproc) fuses letterbox-resize, BGR-to-RGB or RGB-to-BGR conversion, per-channel normalization, HWC-to-NCHW layout conversion, and type casting into a single CUDA kernel with no intermediate buffers. This is useful for vision pipelines where preprocessing latency adds up at high throughput.

The README's benchmark section reports single-stream inference latency on an RTX 3080 Laptop GPU under TensorRT 10. According to that benchmark, YOLOv8n in FP16 runs at 1.07 ms per inference, YOLOv8n in FP32 at 2.00 ms, and MobileNetV2 in FP16 at 0.31 ms. These numbers reflect TensorRT engine inference time only; the wrapper's name-keyed IO and no-throw API add no measurable overhead according to the same benchmark.

## Optional Python Bindings: Zero-Copy With CuPy and PyTorch

The trtcpp Python package exposes the same engine through bindings built with pybind11 and scikit-build-core. Install is:

```bash
pip install .
```

The Python bindings accept CuPy, PyTorch, and Numba GPU arrays via __cuda_array_interface__ or DLPack. Inputs stay on the GPU throughout inference; there are no host round-trips. The GIL is released during inference, allowing Python threads to run other work while the GPU processes a batch.

The README reports that the Python bindings run within approximately 13% of the C++ path on a parity benchmark from examples/python/benchmark_parity.py. This overhead comes from the Python-to-C++ crossing and memory management, not from extra data copies.

The Python wheel is built from the same CMakeLists.txt as the C++ library, with TRT_CPP_API_BUILD_PYTHON=ON set by the scikit-build-core configuration. Pointing the build at a tarball TensorRT installation uses the cmake.define.TensorRT_DIR config setting during pip install.

## Scope Boundaries: CNN Models on Linux, Not Windows or Transformers

The README defines the scope explicitly: Linux, CUDA 12, TensorRT 10 or later, and CNN-style vision models. Windows support is out of scope. Transformer-specific inference features are out of scope. The four runnable examples in examples/ cover the supported use cases: image classification on ImageNet top-5, object detection with YOLOv8n and NMS, semantic segmentation with DeepLabV3, and the zero-copy Python demo.

Comparing tensorrt-cpp-api to using the raw TensorRT C++ API directly: the raw API is the direct alternative and gives full access to nvinfer1 types, builder configurations, and network definitions. It does not include the content-hash engine cache, the EnginePool, or the Status/Result error model. Projects that need to manipulate the TensorRT network graph directly before building an engine, or that have existing code that already wraps TensorRT, will find the raw API more flexible. Projects that want a clean interface for the common inference case will save the boilerplate by using tensorrt-cpp-api.

The library is the inference backend for two related repositories: YOLOv8-TensorRT-CPP and YOLOv9-TensorRT-CPP. The last push was on 2026-05-30. The project is MIT-licensed. The repository has no GitHub releases; the current version is 7.0.0 as stated in pyproject.toml.

## Conclusion

tensorrt-cpp-api is the right starting point for C++ teams targeting TensorRT 10+ on Linux with CNN-style vision models who want safe engine caching, name-keyed IO, and no-throw error handling without writing the boilerplate themselves. It is not suitable for Windows deployment, for LLM or transformer-based models, or for teams still on TensorRT 9 or earlier. Before integrating it into a downstream project, run the benchmark examples from examples/benchmark on your target GPU and TensorRT version to verify that the reported latency numbers match your hardware.

## FAQ

### Does tensorrt-cpp-api work with YOLOv8 models?

Yes. The library is used as the inference backend in the YOLOv8-TensorRT-CPP and YOLOv9-TensorRT-CPP sister projects. The examples/ directory includes an object detection example using YOLOv8n with NMS. The README's benchmark reports YOLOv8n FP16 latency on an RTX 3080 Laptop GPU.

### Can tensorrt-cpp-api be used from Python?

Yes. The trtcpp Python package is built via scikit-build-core with pip install . It accepts CuPy, PyTorch, and Numba GPU arrays without host round-trips. According to the README's parity benchmark, the Python bindings run within approximately 13% of the C++ path.

### Which TensorRT and CUDA versions does tensorrt-cpp-api require?

The library targets TensorRT 10 or later (built to the TensorRT 11 API surface) and CUDA 12. Both TensorRT and CUDA are system-provided dependencies, not managed by the library's build. Windows is explicitly out of scope; the library targets Linux.

## Sources

- [cyrusbehr/tensorrt-cpp-api on GitHub](https://github.com/cyrusbehr/tensorrt-cpp-api)
- [Issues](https://github.com/cyrusbehr/tensorrt-cpp-api/issues)
- [License: MIT](https://github.com/cyrusbehr/tensorrt-cpp-api/blob/main/LICENSE)
- [README](https://github.com/cyrusbehr/tensorrt-cpp-api/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/cyrusbehr-tensorrt-cpp-api
