Lidar AI Solution: four perception pipelines over one set of CUDA building blocks
A project demonstrating Lidar related AI solutions, including three GPU accelerated Lidar/camera DL networks (PointPillars, CenterPoint, BEVFusion) and the related libs (cuPCL, 3D SparseConvolution, YUV2RGB, cuOSD,).
At a glance
- What is it?
- NVIDIA-AI-IOT/Lidar_AI_Solution is described as three GPU accelerated networks, ships four CUDA sub-projects, and layers a small inference engine, an on-screen display library and a point cloud toolkit underneath all of them. The interesting part is the split between TensorRT and hand-written kernels.
- Who is it for?
- Lidar AI Solution is best read as a set of engineering notes about where hand-written CUDA kernels pay for themselves in a lidar perception stack, rather than as a library you install. The 3D sparse convolution engine is the strongest of the pieces, because running the backbone outside TensorRT is what keeps the 422MB FP16 and 426MB INT8 memory figures on the table, and cuOSD is unusually specific about its own API surface.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 72 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 8, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Four CUDA sub-projects, not three
The repository description names three GPU accelerated networks, PointPillars, CenterPoint and BEVFusion, plus a set of related libraries. The file tree tells a slightly fuller story. At the top level sit four directories whose names all begin with CUDA:
- `CUDA-BEVFusion/` - `CUDA-CenterPoint/` - `CUDA-PointPillars` - `CUDA-V2XFusion/`
Alongside them are `libraries/`, `dependencies/`, `assets/`, a `CLA.md`, a `LICENSE.md`, and a `.gitmodules` file. So V2XFusion is implemented here too, and the description simply did not get updated to mention it. The `.gitmodules` file is the reason the install is a recursive clone rather than a plain one.
Each sub-folder holds a self-contained CUDA and TensorRT solution for one network, and the README is explicit that you should read the readme in the sub-folder for the specific task you care about. There is no top-level build script, no unified CMake entry point and no installed package. The root is a directory of projects, not a framework.
What the four have in common is the same shape: a preprocessing stage implemented as a CUDA kernel, a backbone and head exported to ONNX and run through TensorRT, a postprocessing stage as a CUDA kernel, and quantization work on top. The pipelines are described as optimised for self-driving 3D lidar, and the README claims each one reproduces the accuracy of the corresponding torch implementation when you run preparation, inference and evaluation through the provided flow.
The backbone is the part that leaves TensorRT
The most interesting technical claim in the repository concerns 3D sparse convolution, described as a tiny inference engine for 3D sparse convolutional networks running in int8 or fp16. Four properties are listed.
The first is independence from TensorRT. The README calls it a tiny lidar-backbone inference engine independent of TensorRT, and says the execution graph is built from ONNX, with an ONNX export solution included. That is the unusual part, because the surrounding pipelines are all built on TensorRT, so this one component deliberately opts out.
The second is memory. The stated figures are 422MB at FP16 and 426MB at INT8 for the sparse convolutional network. Those two numbers being within four megabytes of each other is the tell: the INT8 path is not halving activation memory, which suggests the footprint is dominated by something other than the quantized weights themselves.
The third is fidelity, claimed as a low accuracy drop on nuScenes validation. The fourth is that it is built on CUDA kernels directly and does not depend on cutlass, which cuts one build dependency out of a stack that otherwise leans on tuned template libraries.
This split is what makes BEVFusion and CenterPoint worth reading. Both describe a lidar encoder running on the tiny engine while the camera encoder and feature fusion run through TensorRT with ONNX export. PointPillars, being 2.5D, uses TensorRT for its backbone throughout and reserves CUDA kernels for voxelization with feature extending and for parsing the bounding box, class type and direction at the end.
Quantization support differs per network
The four sub-projects treat quantization differently, and the distinction is post-training quantization against quantization-aware training rather than a uniform effort.
CenterPoint is the one with QAT, plus quantization solutions for the spconv implementation used by the traveller59 project. Its preprocessing is voxelization through a CUDA kernel and its postprocessing is decoding with non-maximum suppression through another, with the 3D backbone using NVIDIA's spconv-scn and the region proposal network and CenterHead handled through TensorRT with an ONNX export solution.
BEVFusion gets post-training quantization instead, described as quantization solutions for the spconv used by mmdet3d. Its stage list is longer, because a camera and lidar fusion model has more boundaries: a ResNet50 camera encoder, a finetuned BEV pooling step, the tiny lidar backbone, a camera and lidar feature fuser, and a pre and postprocess stage covering interval precomputing, lidar voxelization and a feature decoder, all in CUDA kernels.
V2XFusion adds the one capability none of the others claim. It supports 4:2 structural sparsity, and it takes a different tack on quantisation error by using a PointPillars based backbone with pre-normalization, which the README says reduces quantisation error. Its DeepStream sample runs inference with CUDA, TensorRT or Triton inside DeepStream SDK 7.0, which dates that particular sample and means the pinned SDK version matters if you copy it.
PointPillars, the simplest of the four, has no quantisation story at all. It is the baseline the other three extend, and reading it first makes the others easier to follow.
cuOSD and the shape of the utility layer
Underneath the networks sits a set of libraries, and cuOSD is the most fully specified one. It is a CUDA on-screen display library that draws all elements using a single CUDA kernel, which is the whole architectural claim: one launch for an entire overlay instead of one per primitive.
The element list is explicit: line, with nearest or linear interpolation; rotatebox with configurable border and fill colours; circle; rectangle; text; arrow, built from a combination of three lines; point, with nearest or linear interpolation; and a clock plotted on the text support.
Text is the part with a dependency decision worth noting. It supports the stb_truetype backend and the pango-cairo backend, which means fonts can be read from a TTF file or resolved by font-family through the system. Supporting both costs something in setup, but it covers the two real cases, an embedded device with no font stack and a workstation that already has one.
cuPCL is the point cloud counterpart, offering GPU accelerated operations with high accuracy and high performance at the same time:
cuICP, cuFilter, cuSegmentation, cuOctree, cuCluster, cuNDTThat expands to CUDA accelerated iterative corresponding point registration, PassThrough and VoxelGrid filtering, random sample consensus against a plane model, approximate nearest and radius search, distance based clustering, and 3D normal distribution transform registration. Voxelization is listed as incoming, so the registration and search primitives are the finished part of this library.
The conversion libraries are where the documentation gets repetitive. YUVToRGB and ROI Conversion are each described as combining resize, padding, conversion and normalisation into a single kernel function, and both share the same input formats, NV12BlockLinear, NV12PitchLinear and YUV422Packed_YUYV, the same nearest and bilinear interpolation, and the same Uint8, Float32 and Float16 output types. Both claim bit alignment with OpenCV when the scaling factor is a rational number, and better performance when the stride divides by four. ROI Conversion adds a Gray output layout and the DLA input layouts that YUVToRGB already lists. Read them as two front ends over largely the same kernel work.
Installing means a recursive clone and reading four readmes
The entire install procedure is two commands, and the flag is the whole story:
git clone --recursive https://github.com/NVIDIA-AI-IOT/Lidar_AI_Solution
cd Lidar_AI_Solution`--recursive` pulls the submodules declared in `.gitmodules`, which is how the four CUDA sub-projects obtain their content. Without it you get four directory names and nothing inside them, which is a failure mode that costs people an afternoon because nothing errors loudly.
After that the README offers no further help at the root level, because the architecture is per-task. You are expected to move into the sub-folder for your network and read the readme there. There are no release artefacts, no version tags and no container images referenced in the documentation, so nothing pins the code. GitHub reports no releases at all for this repository.
The supporting pieces are all present but all partial. `dependencies/` is where third-party code sits, with the README crediting stb_image for PNG and JPEG support and pybind11 for C++ and Python interop, and pointing at the folder for the rest. `assets/` holds the title image and the pipeline diagram, which is the fastest way to see how the stages connect. `CLA.md` means contributions are covered by a contributor licence agreement, worth knowing before you send a patch.
The last push to the default branch landed on 2026-07-28, which puts the repository in active development rather than a frozen archive. That is the right time to read it, and also the reason to expect the sub-folder instructions to move.
Signals in the metadata worth checking before you build on it
Three things about how this repository presents itself deserve a second look before you plan around it.
The first is the description drift noted earlier. It says three GPU accelerated lidar and camera networks, and the tree contains four sub-projects including V2XFusion with its sparsity support and its DeepStream 7.0 sample. Nothing is broken by this, but it tells you the description is maintained by hand and may lag the tree, which is a reason to read the directory listing rather than the summary.
The second is that the README promises per-task readmes in the sub-folders and this listing cannot confirm they are all present. The tree exposes the sub-project directories at the top level, and the per-folder documentation is the only place where the actual build commands, TensorRT versions, ONNX export steps and expected accuracy numbers live. Anyone writing about this repository from the outside is working from a summary, and that summary stops at the README.
The third is the license signal. GitHub reports no recognised license identifier for the repository, while a `LICENSE.md` file sits at the root. For a codebase you intend to compile and link into a product, the file in the tree is the authority, and NVIDIA's licensing posture on its AI IoT samples is the thing to confirm for yourself rather than infer from the metadata field.
Set against that, the technical content is unusually concrete where it can be. Memory footprints are given to the megabyte, the toolkit components are named one by one, and the supported pixel formats are enumerated rather than gestured at. This is a repository whose value is in its specifics, which is precisely why the missing versions and the absent release history matter more here than they would in a project that was mostly documentation.
Editorial conclusion
Lidar AI Solution is best read as a set of engineering notes about where hand-written CUDA kernels pay for themselves in a lidar perception stack, rather than as a library you install. The 3D sparse convolution engine is the strongest of the pieces, because running the backbone outside TensorRT is what keeps the 422MB FP16 and 426MB INT8 memory figures on the table, and cuOSD is unusually specific about its own API surface. What you cannot do from here is pin anything: GitHub reports no releases for this repository, the version a driver looks for has to come from each sub-folder, and the last push landed on 2026-07-28. Clone it with `--recursive`, read the readme inside the CUDA sub-folder for the network you care about, and start with `libraries/cuOSD` if you only want the drawing kernels.
Frequently asked questions
What networks does Lidar AI Solution actually implement?
The file tree shows four CUDA sub-projects: CUDA-PointPillars, CUDA-CenterPoint, CUDA-BEVFusion and CUDA-V2XFusion. The repository description names only the first three, so V2XFusion is present but missing from the summary. Each is a separate CUDA and TensorRT solution with preprocessing and postprocessing kernels of its own, and the README directs you to the readme inside each sub-folder for build instructions.
What is the 3D sparse convolution engine used for?
It is a small inference engine for 3D sparse convolutional networks that runs in int8 and fp16 and is independent of TensorRT, building its execution graph from ONNX. The README reports memory of 422MB at FP16 and 426MB at INT8 and a low accuracy drop on nuScenes validation. It is used as the lidar backbone inside the CenterPoint and BEVFusion pipelines, while the remaining stages of those pipelines run through TensorRT.
How do I install the code and why does the clone flag matter?
The documented procedure is `git clone --recursive https://github.com/NVIDIA-AI-IOT/Lidar_AI_Solution` followed by `cd Lidar_AI_Solution`. The recursive flag is required because a `.gitmodules` file at the root declares the CUDA sub-projects as submodules. Without the flag the four sub-folders are present but empty, and there is no later step that would tell you they were never populated.
Are there tagged releases I can pin a build to?
No. GitHub reports no releases for this repository, and the README references no version tags or container images. Because the four sub-projects live behind submodules with their own build requirements, including TensorRT versions and ONNX export steps documented only in their sub-folder readmes, you are tracking a moving default branch rather than a fixed point.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvidia-ai-iot-lidar-ai-solution)