Open-source project
cyrusbehr/YOLOv8-TensorRT-CPP avatar
cyrusbehr/YOLOv8-TensorRT-CPP

YOLOv8-TensorRT-CPP: A C++ Inference Path for YOLOv8 on NVIDIA GPUs

YOLOv8 TensorRT C++ Implementation

737 stars84 forksC++MIT

At a glance

What is it?
The repository wraps a TensorRT C++ API into object detection, segmentation and pose estimation executables. The trade-off is a manual toolchain: CUDA, cuDNN, CUDA-enabled OpenCV and a hand-edited CMake path before anything compiles.
Who is it for?
Adopt it if you already ship CUDA code and want YOLOv8 inference inside a C++ process without a Python runtime, and if you can accept Ubuntu only. Skip it if you need Windows, a packaged installer, or a maintained upstream: the README asks for maintainers and the last push was on 2026-05-30.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 124 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap this fills between PyTorch training and GPU deployment

YOLOv8 is trained and exported through Ultralytics in Python. Deploying that model into a C++ service usually means either embedding a Python interpreter or rewriting preprocessing, decoding and non-maximum suppression by hand. This repository takes the second route and publishes the result as a small set of command line executables. The README describes it as "a C++ Implementation of YoloV8 using TensorRT" and states that it supports object detection, semantic segmentation, and body pose estimation.

The intended reader is an engineer who already builds against CUDA and TensorRT and wants the model to run inside the same process as the rest of the application. It is not a library you link into an existing pipeline: the repository ships executables such as detect_object_image, detect_object_video and benchmark, and the inference work is delegated to the author's separate tensorrt-cpp-api project, pulled in as a submodule. The README is explicit that you should be familiar with that other project first, which is a fair signal about the expected skill level.

How inference is split: preprocessing, engine, C++ postprocessing

The design keeps the model graph narrow. The conversion script scripts/pytorch2onnx.py produces an ONNX file, and the README warns that if you convert with a different script, end2end must be disabled, because that flag bakes bounding box decoding and NMS into the model. Here those steps are done "external to the model using good old C++". That choice matters: it keeps the TensorRT engine free of custom plugins for decoding, and it lets you change confidence thresholds and IoU behaviour without rebuilding the engine. The cost is that postprocessing time is real and visible in the benchmark tables, where yolov8n-seg spends 12.353 ms of a 15.194 ms total in postprocess.

At runtime the executables accept either an ONNX model through --model or a prebuilt engine through --trt_model. The first run with an ONNX file triggers engine construction, which the README says may take five minutes or more; the engine is then saved to disk and loaded on later runs. Preprocessing, inference and postprocessing are separately timed when the ENABLE_BENCHMARKS flag is compiled in, which is the only way to see where time goes on your own hardware.

Installing on Ubuntu: prerequisites before the first build

There is no package or container path documented. The README lists the prerequisites as Ubuntu 20.04 or 22.04, CUDA (recommended 12.0 or newer), cuDNN (8 or newer), build-essential, python3-pip, cmake via pip3, OpenCV built with CUDA support (4.8 or newer, with a build_opencv.sh script referenced in the tensorrt-cpp-api repository), and TensorRT 10 or newer. Windows is stated as not supported at this time.

The build tools come from apt and pip:

bash
sudo apt install build-essential
sudo apt install python3-pip
pip3 install cmake

The repository must be cloned with submodules, because the inference backend lives in one:

bash
git clone https://github.com/cyrusbehr/YOLOv8-TensorRT-CPP --recursive

After extracting TensorRT, the README says to open CMakeLists.txt and replace the TODO with the path to your TensorRT installation. That edit is manual and there is no documented fallback if the path is wrong. Then the project builds with the usual sequence:

bash
mkdir build
cd build
cmake ..
make -j

A first run: converting a model and detecting objects in an image

Model conversion happens in Python, once, before any C++ code runs. Install Ultralytics, download the model you want from the official YOLOv8 repository, then run the script from the scripts directory:

bash
pip3 install ultralytics
python3 pytorch2onnx.py --pt_path <path to your pt file>

The README says you should end up with an ONNX file. The same script is documented as handling segmentation models such as YOLOv8x-seg and pose models, so the executable you choose later should match the model type.

For image inference, the documented invocation is:

bash
./detect_object_image --model /path/to/your/onnx/model.onnx --input /path/to/your/image.jpg

The annotated image is saved to disk. The first execution will spend minutes building the TensorRT engine; later runs load the cached engine. For a live camera instead, the README gives:

bash
./detect_object_video --model /path/to/your/onnx/model.onnx --input 0

Running any executable with no arguments prints the full argument list, which is the practical way to discover options such as --trt_model and the INT8 flags.

INT8 calibration and the accuracy cost it buys

INT8 is opt-in and needs data. The README states that calibration data representative of real inference input must be supplied, and advises 1K or more calibration images. The documented source is the COCO validation set:

bash
wget http://images.cocodataset.org/zips/val2017.zip

With data in place, the executables take two extra arguments:

bash
--precision INT8 --calibration-data /path/to/your/calibration/data

The README is honest that this trades accuracy for speed because of the reduced dynamic range. It also documents one concrete failure: an "out of memory in function allocate" error means Options.calibrationBatchSize is too large for the GPU, and the fix is to reduce it. That option is not exposed as a command line flag in the README, so lowering it means editing the source and rebuilding.

Where the project stops short

The most visible boundary is the platform. Windows is explicitly not supported, and there is no mention of Jetson, ARM or Docker images, so anyone deploying to an edge device is on their own. The build assumes you can compile OpenCV with CUDA, which is a heavier task than the rest of the setup combined and is only linked out to another repository's script.

Maintenance is the second boundary. The README carries a "Looking for Maintainers" section asking for help guiding the project's growth, which is unusual to see in a repository that is not archived. The last push was on 2026-05-30. That is recent enough that the code is not abandoned, but a project soliciting maintainers is telling you something about the bus factor, and the README does not document a release process, a versioning scheme, or a rollback path for a bad engine build.

The third boundary is scope. This is a set of demonstration executables, not a serving framework. There is no batching server, no dynamic shape handling described, and no API stability promise. If you need a long-running inference service with request queuing, you will be writing that layer yourself around the tensorrt-cpp-api submodule.

How it compares with the Ultralytics Python path

The obvious alternative is to stay in Python and call Ultralytics directly, which also runs on TensorRT and supports Windows. The difference is where the work happens. Ultralytics owns the whole pipeline: export, inference, decoding and NMS are library calls, and the same code runs on CPU, CUDA or TensorRT with a device argument. This repository takes the opposite stance and hands decoding and NMS to you in C++, with the model exported as a plain ONNX graph. That gives tighter control over thresholds and avoids a Python runtime in production, at the price of writing and maintaining the preprocessing and postprocessing code yourself.

A second alternative is the author's own tensorrt-cpp-api project, which this repository depends on. If your model is not YOLOv8, that project is the more general starting point; this one adds the YOLOv8-specific decode and the three task heads on top of it.

Licence and the cost of keeping it current

The repository is MIT licensed, which permits commercial use and modification provided the copyright notice and permission notice are retained. That is permissive enough for most products, but note that the licence covers this repository, not the submodule it pulls in or the Ultralytics models you export; those carry their own terms, and the README does not discuss them. This is not legal advice.

Upgrade cost is dominated by two moving parts. TensorRT 10 is the documented minimum, and the README's benchmark tables were produced on an RTX 3080 Laptop GPU with an i7-10870H, so those numbers describe one machine rather than a range. Engine files are tied to the TensorRT version and GPU they were built on, which means a TensorRT upgrade invalidates your cached engines and forces a rebuild plus a fresh five-minute-or-longer engine generation per model. The MIGRATION.md file at the repository root suggests upgrades have required code changes before, though its contents are not part of the README.

Editorial conclusion

Adopt it if you already ship CUDA code and want YOLOv8 inference inside a C++ process without a Python runtime, and if you can accept Ubuntu only. Skip it if you need Windows, a packaged installer, or a maintained upstream: the README asks for maintainers and the last push was on 2026-05-30. Verify first that your TensorRT 10 install path is set in CMakeLists.txt, that OpenCV was built with CUDA, and that your exported ONNX model has end2end disabled, because the bounding box decoding and NMS happen in the C++ code rather than in the graph.

Frequently asked questions

Does YOLOv8-TensorRT-CPP run on Windows?

No. The README lists Ubuntu 20.04 and 22.04 as tested and working and states that Windows is not supported at this time.

Which YOLOv8 tasks does YOLOv8-TensorRT-CPP support?

The README says it supports object detection, semantic segmentation, and body pose estimation, and that the executables work out of the box with Ultralytics pretrained models for all three.

Why does the first run of a YOLOv8-TensorRT-CPP executable take so long?

The README notes that the first run may take five minutes or more because TensorRT must generate an optimized engine file from the ONNX model. That engine is saved to disk and loaded on subsequent runs.

What is TensorRT used for?

In this project it is the inference runtime: the C++ API runs the exported YOLOv8 ONNX model on the GPU, and the README requires TensorRT 10 or newer.

Is TensorRT a part of NVIDIA?

The README links to NVIDIA's CUDA and cuDNN download pages and to a TensorRT 10 download page, and the prerequisites are all NVIDIA toolchain components. The repository itself is a separate MIT-licensed project by cyrusbehr.

Official sources

  1. cyrusbehr/YOLOv8-TensorRT-CPP on GitHub
  2. Issues
  3. License: MIT
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/cyrusbehr-yolov8-tensorrt-cpp.svg)](https://hysenlabs.com/projects/cyrusbehr-yolov8-tensorrt-cpp)