# onnx-tensorrt: parsing ONNX models into TensorRT engines

> onnx-tensorrt is the C++ parser library and Python backend that turns an ONNX file into a TensorRT engine. It is for teams already committed to NVIDIA hardware who need to control how that conversion happens.

**onnx/onnx-tensorrt** — ONNX-TensorRT: TensorRT backend for ONNX

- Repository: https://github.com/onnx/onnx-tensorrt
- Stars: 3,239 · Forks: 552
- Language: C++
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/onnx-onnx-tensorrt

## The gap between an ONNX file and a TensorRT engine

An ONNX file describes a computation graph. A TensorRT engine is a serialized, hardware-specific plan. Nothing in the ONNX specification says how a graph should be lowered onto a particular GPU, and nothing in TensorRT knows how to read protobuf. onnx-tensorrt sits in that gap. The README states its purpose plainly: it parses ONNX models for execution with TensorRT.

That one-sentence job description hides the reason the project exists as a separate repository. TensorRT ships a parser, but the parser is also the place where operator coverage, shape inference, constant folding and plugin selection get decided. Keeping it open source means an engineer whose model fails to import can read the importer for that operator, see which TensorRT layer it maps to, and either patch it or work around it. The repository layout reflects that: onnxOpImporters.cpp and onnxOpCheckers.cpp sit at the top level, next to ModelImporter.cpp and ImporterContext.cpp. This is not a thin wrapper. It is the translation layer itself.

The audience is narrower than the topic list suggests. You need a TensorRT 11.2 installation, a CUDA toolkit, and a reason to build C++ against it. If you only want to run an ONNX model on a GPU, you are not the target user of this repository; you are the target user of the tools described further down.

## How the parser turns operators into TensorRT layers

The architecture is a single pass over the ONNX graph. ModelImporter.cpp walks the model, and for each node it dispatches to an importer registered for that operator type. Those importers live in onnxOpImporters.cpp. Each one reads the node's attributes through OnnxAttrs, converts its inputs and outputs into TensorOrWeights objects, and emits TensorRT layers into the network being built.

Two pieces of state make this more than a switch statement. ImporterContext carries the per-parse context, including the flag set that decides optional behaviour. WeightsContext and WeightsContextMemoryMap handle the constant data attached to the graph, which matters because a large model's weights should not be copied more than necessary. ShapeTensor.cpp and ShapedWeights.cpp deal with the awkward cases where a shape is itself a computed value rather than a constant, which is what makes dynamic and full-dimension shapes possible.

Subgraph handling is split into its own files. LoopHelpers.cpp, ConditionalHelpers.cpp and RNNHelpers.cpp exist because ONNX control flow and recurrent operators contain nested graphs, and those nested graphs have to be imported recursively rather than as flat sequences of layers. The presence of these three files separately from the main importer is a reasonable signal of where the complexity actually lives.

ModelRefitter.cpp is the other notable entry. It addresses the case where you have a built engine and want to swap weights without re-importing the whole graph, which is a different operation from parsing and worth knowing about before you assume every weight change requires a full rebuild.

The output of all this is a TensorRT network, which TensorRT then builds into an engine. The parser does not build the engine. That distinction matters when something fails: an import error and a build error come from different stages and have different fixes.

## Building onnx-tensorrt and running a first model

There is no package on PyPI for the parser itself. The README directs you to clone the repository and build with CMake, and it points at the main TensorRT repository for Docker and Windows build instructions. The dependencies listed are TensorRT 11.2, the TensorRT 11.2 open source libraries, and optionally Protobuf 3.20.3 or newer.

The build sequence from the README, with the TensorRT path substituted:

```bash
cd onnx-tensorrt
mkdir build && cd build
cmake .. -DTENSORRT_ROOT=<path_to_trt> && make -j
export LD_LIBRARY_PATH=$PWD:$LD_LIBRARY_PATH
```

The export line is not optional decoration. The README notes it explicitly so that the newly built library is found at runtime. By default CMake looks in /usr/local/cuda for CUDA; if your toolkit is elsewhere, pass -DCUDA_TOOLKIT_ROOT_DIR=<path_to_cuda_install> instead. There is also a -DUSE_ONNX_LITE_PROTO=1 flag for protobuf-lite builds.

For the Python path, the README pins the ONNX version: TensorRT 11.2 supports ONNX release 1.21.0. Install that first, then the backend.

```bash
python3 -m pip install onnx==1.21.0
python3 setup.py install
```

The setup.py file lists pycuda, numpy and onnx as required packages, so pycuda is pulled in whether or not you use the parser directly from C++.

A first real use, taken from the README's Python example:

```python
import onnx
import onnx_tensorrt.backend as backend
import numpy as np

model = onnx.load("/path/to/model.onnx")
engine = backend.prepare(model, device='CUDA:1')
input_data = np.random.random(size=(32, 3, 224, 224)).astype(np.float32)
output_data = engine.run(input_data)[0]
print(output_data)
print(output_data.shape)
```

The device string selects the GPU, so 'CUDA:1' means the second device. If prepare returns without raising, the model imported and built. The shape printed at the end is the engine's output shape, which is the quickest way to confirm that dynamic dimensions resolved the way you expected.

If you would rather not write code at all, the README names two officially supported tools. C++ users get trtexec, usually in <tensorrt_root_dir>/bin, invoked as trtexec --onnx=model.onnx. Python users get polygraphy, invoked as polygraphy run model.onnx --trt. Both are useful as a first check before you invest in a custom integration.

## The InstanceNormalization flag and what it costs you

The README documents two implementations of InstanceNormalization, and the default is the native TensorRT one. To benchmark the plugin implementation instead, you unset the parser flag kNATIVE_INSTANCENORM before parsing.

```python
parser.clear_flag(trt.OnnxParserFlag.NATIVE_INSTANCENORM)
```

The C++ equivalent is parser->unsetFlag(nvonnxparser::OnnxParserFlag::kNATIVE_INSTANCENORM).

Here is the part worth reading twice. The README states that the plugin implementation cannot be used for building version compatible or hardware compatible engines, and that attempting to do so results in an error. So the flag is not a free performance knob. Choosing the plugin path means giving up engine compatibility, which is the mechanism that lets an engine built on one machine run on another with a different TensorRT version or a different GPU. If your deployment ships prebuilt engines to heterogeneous hardware, the default native implementation is the one you can actually use, regardless of which one benchmarks faster on your development box.

This is a good example of the kind of trade-off the project exposes rather than hides. The flag exists, it is documented, and the constraint attached to it is stated in the same paragraph.

## Where onnx-tensorrt is the wrong choice

Operator coverage is the first constraint and the one most likely to stop you. The README points to a separate operator support matrix in docs/operators.md rather than claiming complete coverage. A model containing an operator that is not in that matrix will not import. The failure is not silent, but it is also not something you can fix by changing a flag; you either rewrite that part of the graph, implement a plugin, or pick a different path. Check the matrix against your model before you plan any work around this repository.

The second constraint is version coupling. Development on the main branch targets TensorRT 11.2, and the README says previous TensorRT versions have their own branches. An ONNX model is portable; the engine built from it is not. If your production environment runs an older TensorRT, you are on a different branch with a different operator set, and the release branches in the repository (11.0-GA, 11.1-GA, 11.2-GA) are the visible trace of that maintenance model. Upgrading TensorRT means moving branches and re-validating your models.

The third constraint is the build itself. There is no wheel for the parser. You need TensorRT, CUDA, CMake and a working compiler toolchain, and the README defers Docker and Windows specifics to the TensorRT repository rather than covering them here. Teams without an existing TensorRT build pipeline should treat this as real setup work, not an afternoon.

Finally, if your goal is simply to run an ONNX model on whatever hardware is available, this is the wrong layer. The Python backend's device argument is a CUDA device string. There is no CPU path here.

## onnx-tensorrt against the ONNX Runtime TensorRT execution provider

The most common alternative is the ONNX Runtime TensorRT execution provider. The difference is where the boundary sits.

With onnx-tensorrt, your program owns the model, the parser and the engine. You call backend.prepare, you get an engine object back, and you call run on it. Everything between the ONNX file and the CUDA kernel is in your process and, because the parser is open source, inspectable. That control is the point. If an operator imports badly, you can read the importer. If you need a specific flag set before parsing, you set it. If you want to build the engine once and serialize it, the parser gives you the network to build from.

With the execution provider, ONNX Runtime owns the session. You create an InferenceSession with the TensorRT provider and run it like any other ONNX Runtime session. The provider handles partitioning: operators it can take go to TensorRT, and the rest fall back to other providers. That fallback is the substantive difference. A model with partial operator coverage can still run, because unsupported nodes do not have to be handled by TensorRT. With the standalone parser, partial coverage is an import failure.

The cost of that fallback is that you have less control over where the boundary lands, and a graph split across providers can involve transfers between them. Which approach wins depends on whether your model fits cleanly in TensorRT's operator set. If it does, the standalone parser gives you a tighter, more predictable engine. If it does not, the execution provider is the more pragmatic starting point, and you can move to the parser later for the subgraphs that matter.

## Licence, release branches and the cost of keeping up

onnx-tensorrt is Apache-2.0, and the licence identifier appears as an SPDX header at the top of the README and setup.py. Apache-2.0 is permissive and includes an explicit patent grant, which is the usual reason projects in this space choose it over MIT. It is not a copyleft licence, so linking the parser into a proprietary product is not, on its face, a source-disclosure trigger. That is a description of the licence text, not legal advice; if the distinction matters to your organisation, have counsel read it.

The practical maintenance cost is version tracking rather than patch cadence. Releases follow TensorRT: 11.0-GA, 11.1-GA and 11.2-GA, with the most recent push to the repository on 2026-08-03. The versioning scheme tells you what to expect. There is no independent semantic version for the parser; the branch you are on corresponds to a TensorRT release. When TensorRT moves, you move, and moving means re-importing and re-validating your models because the operator importers and the flag behaviour can change between branches.

One upgrade detail worth planning for: the README pins ONNX 1.21.0 for TensorRT 11.2. If your pipeline produces ONNX files with a newer opset, check that combination before upgrading either side. The two version numbers are coupled in the documentation, and treating them as independent is how you end up debugging an import error that is really a version mismatch.

## Conclusion

Adopt onnx-tensorrt if you are building against TensorRT 11.2 and need the parser as a library rather than a tool, or need to inspect and change how individual ONNX operators map onto TensorRT layers. Do not adopt it if you want a portable runtime: it requires an NVIDIA GPU, a matching TensorRT build and a source build with CMake. Before committing, check the operator support matrix against your model's operator set, confirm which TensorRT version your target deployment carries, and decide whether the native or plugin InstanceNormalization path is the one you can ship.

## FAQ

### Is TensorRT faster than vLLM?

The README does not compare TensorRT with vLLM, and it gives no throughput or latency figures for either. It only documents how ONNX models are parsed for execution with TensorRT, so this question cannot be answered from the project material.

### Is TensorRT a part of NVIDIA?

TensorRT is developed by NVIDIA: the README links to developer.nvidia.com for the download and lists researchinquiries@nvidia.com for business inquiries. This repository, onnx-tensorrt, is the ONNX parser for it and is published under the onnx organisation.

### What is the difference between CUDA and TensorRT?

The README treats them as separate dependencies: TensorRT 11.2 is required to build the parser, and CUDA is a build dependency that CMake looks for in /usr/local/cuda by default. The Python backend takes a CUDA device string such as 'CUDA:1', so CUDA is the device layer the engine runs on.

### What is onnx-tensorrt?

It is a parser that reads ONNX models and builds them into TensorRT networks, plus a Python backend for running them. The README describes it as the TensorRT backend for ONNX, and the C++ library is libnvonnxparser.so with its API declared in NvOnnxParser.h.

### What is the difference between onnx-tensorrt and CUDA?

CUDA is a dependency of the build rather than an alternative to it: the README lists it as required and lets you point at a non-default toolkit path with -DCUDA_TOOLKIT_ROOT_DIR. onnx-tensorrt sits above that layer, turning an ONNX graph into a TensorRT network that TensorRT then builds into an engine to run on the CUDA device.

## Sources

- [Issues](https://github.com/onnx/onnx-tensorrt/issues)
- [License: Apache-2.0](https://github.com/onnx/onnx-tensorrt/blob/main/LICENSE)
- [onnx/onnx-tensorrt on GitHub](https://github.com/onnx/onnx-tensorrt)
- [README](https://github.com/onnx/onnx-tensorrt/blob/main/README.md)
- [Releases](https://github.com/onnx/onnx-tensorrt/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/onnx-onnx-tensorrt
