# TensorRT: NVIDIA's Inference SDK and What the Open Source Repository Actually Ships

> TensorRT is an SDK for high-performance deep learning inference on NVIDIA GPUs, and the repository holds only its open source components: plugins, the ONNX parser, samples and demos. The 11.x line removed weakly-typed networks, implicit quantization, IPluginV2 and TREX, so version choice is a migration decision, not a preference.

**NVIDIA/TensorRT** — NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.

- Repository: https://github.com/NVIDIA/TensorRT
- Website: https://developer.nvidia.com/tensorrt
- Stars: 13,374 · Forks: 2,411
- Language: C++
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/nvidia-tensorrt

## What TensorRT solves, and who the open source repository is actually for

TensorRT takes a trained network and produces an inference engine tuned for a specific NVIDIA GPU. The pitch is not training. It is the serving side: you have a model that already works in PyTorch or TensorFlow, and you want it to answer requests faster on hardware you own. The README frames the repository as "the Open Source Software (OSS) components of NVIDIA TensorRT," a subset of the General Availability release with some extensions and bug-fixes. That sentence matters more than it looks. What you clone is not the whole SDK. It is the ONNX parser, the plugin sources, sample applications and demos, plus build glue. The engine builder, the runtime libraries and the kernels largely arrive as the GA build you download separately or get preinstalled in the container.

So the audience splits in two. The first group installs the Python package and never builds anything; the README says you can skip the Build section entirely and enjoy TensorRT with Python. The second group needs the sources: people writing custom layers as plugins, patching the ONNX parser for an operator it does not handle, or reading the samples to learn the API. If you are in the first group, the repository is documentation and reference material. If you are in the second, it is your build tree, and the GA version you pair it with becomes a hard constraint.

## How the pieces fit: ONNX parser, plugins, engine, runtime

The data flow the repository exposes is narrow and worth stating plainly. A model arrives as ONNX. The parser in parsers/ converts that graph into a TensorRT network. Plugins from plugin/ supply layer implementations the core library does not provide, including custom layers you write yourself. The network is then built into an engine, which is what the runtime executes on the GPU. The README points to a separate Import Workflows Guide for "the TensorRT import paths (ONNX, Torch-TensorRT, HuggingFace/Optimum, Network Definition API)," which tells you ONNX is one of four doors, not the only one.

The 11.x release notes in the README describe an API that has been "streamlined," which is a diplomatic word for removals. Weakly-typed networks are gone, replaced by strongly typed networks. Implicit quantization is gone, replaced by explicit quantization. IPluginV2 is gone, replaced by IPluginV3. TREX is gone, replaced by Nsight Deep Learning Designer. Python bindings for Python 3.9 and older are gone, and RPM packages for RHEL/Rocky Linux 8 and 9 now depend on Python 3.12.

That is a coherent design direction. Strong typing and explicit quantization both move decisions from the builder's guesswork into your graph, where you can inspect them. IPluginV3 replaces an interface whose lifetime and resource handling were a recurring source of bugs. But the cost is real: any plugin you wrote against IPluginV2 does not compile against 11.x, and any calibration-based quantization workflow has to be re-expressed as explicit quantization. The README links migration guides for each of these, which is the honest signal that migration is expected work, not an edge case.

## Installing TensorRT and running a first model

The shortest path is the prebuilt Python package. The README gives one command and states that after it you can skip the Build section.

```bash
pip install tensorrt
```

Python support in this line is version-bounded: the build prerequisites list python >= v3.10 and <= v3.14.x, and the release notes say bindings for 3.9 and older were removed. Check your interpreter before you start, because a 3.9 environment will not get a usable wheel from this series.

If you need the sources, the README's first build step is a clone with submodules. The submodule update is not optional; onnx-tensorrt, cub and protobuf are downloaded along with TensorRT OSS rather than installed separately.

```bash
git clone -b main https://github.com/nvidia/TensorRT TensorRT
cd TensorRT
git submodule update --init --recursive
```

Before configuring, you need the matching TensorRT GA build. The README names TensorRT v11.2.1.2, available as direct download links, and offers two CUDA pairings: cuda-13.3.0 and cuda-12.9.0 are the recommended versions. If you build inside the TensorRT OSS container, the libraries are preinstalled under /usr/lib/x86_64-linux-gnu and you can skip the download. Otherwise you extract the GA tarball and point the build at it. The remaining prerequisites are GNU make >= v4.1, cmake >= v3.31, pip >= v19.0, and the usual git, pkg-config and wget.

Multi-device support is opt-in. The README lists NCCL >= v2.19, < v3.0 as required only when building with -DTRT_BUILD_ENABLE_MULTIDEVICE=ON, which is what the sampleDistCollective sample needs. Leave that flag off and NCCL is not part of your build.

For a first real use, the repository's own entry points are the samples/ directory and the demo/ directory, which contains DeBERTa, Diffusion and EfficientDet examples. The README also mentions an .agents/skills directory with skills related to TensorRT usage and benchmarking, installed according to your preferred coding agent's instructions. None of these are substitutes for reading the Import Workflows Guide if your model comes from a framework rather than from ONNX directly.

## The version coupling is the real limitation

The hardest constraint in this repository is not performance. It is that the OSS components are pinned to a specific GA build. The README names TensorRT v11.2.1.2 and two recommended CUDA versions. Build against a different GA release and you are outside the supported combination, with the failure modes you would expect from mismatched headers and libraries. This is not a repository you can float on latest-of-everything.

The second limitation is the removal cadence. Three releases landed between 2026-06-02 and 2026-08-04: v11.0, v11.1, v11.2. The README describes 11.X as a major version bump that removed legacy features from 10.X. If your deployment depends on TREX for visualizing engines, on implicit quantization, or on IPluginV2 plugins, 11.x is the wrong line for you until those are ported, and the README's migration guides are the place to judge how much work that is.

Third, the repository is not the product. If you file an issue expecting the engine builder to change, you are in the wrong tracker; the README directs business inquiries to researchinquiries@nvidia.com and points to NVIDIA AI Enterprise for enterprise support. The open source tree covers plugins, the ONNX parser and samples. Everything else is the GA build.

Finally, hardware scope. TensorRT targets NVIDIA GPUs. There is no CPU fallback path described here, and no non-NVIDIA accelerator path. If your serving fleet is mixed, TensorRT is one branch of it, not the runtime for all of it.

## TensorRT against vLLM and against CUDA

The comparison people actually search for is TensorRT against vLLM, and the two sit at different layers. TensorRT is an SDK: you produce an engine, then you serve it, and the serving layer is a separate concern. vLLM is a serving system with its own scheduling and memory management for language models. Choosing TensorRT means you own the engine build step and the serving wrapper. Choosing vLLM means you inherit its scheduler and its paged attention implementation, and you accept its model support surface. For a single model on a fixed GPU, an engine you built yourself gives you control over exactly what runs. For a fleet serving many models with bursty traffic, a serving system that manages memory across requests is doing work you would otherwise write.

The TensorRT against CUDA comparison is a category error that the search data keeps producing. CUDA is the programming model and toolkit. TensorRT is a library that runs on top of it. You do not pick one instead of the other; you pick how much of the graph-level optimization you want TensorRT to do for you versus how much kernel-level work you want to write in CUDA yourself. The README's CUDA version recommendations (13.3.0 or 12.9.0) are a build requirement, not an alternative.

A more useful axis is the import path. The README lists ONNX, Torch-TensorRT, HuggingFace/Optimum and the Network Definition API, with a per-model support matrix in documents/supported_models.md covering LLM, encoder-NLP, vision, audio, diffusion and multimodal categories. If your model is a HuggingFace transformer, going through Optimum or Torch-TensorRT is a shorter road than exporting to ONNX and patching the parser. If your model has custom operators, the Network Definition API plus your own IPluginV3 layers is the road, and it is longer.

## Licence, maintenance and upgrade cost

The repository is Apache-2.0, and there is a NOTICE file at the top level alongside LICENSE, which is standard for a project that incorporates third-party code (the README notes onnx-tensorrt, cub and protobuf are downloaded along with the OSS tree). Apache-2.0 covers the open source components in this repository. It does not automatically describe the TensorRT GA build you download from NVIDIA Developer Zone, which is a separate artifact with its own terms. If you are shipping a product, the licence question is not answered by the badge at the top of the README, and this is a question for your own counsel rather than for an article.

On maintenance: the last push to the default branch was on 2026-08-25, and the most recent release is v11.2 from 2026-08-04. The repository is not archived. The README also links a roadmap PDF for Q3 2026, which is the closest thing to a forward-looking statement in the README.

Upgrade cost is where you should budget. Moving from 10.X to 11.X means porting plugins to IPluginV3, re-expressing quantization explicitly, and confirming your Python environment is 3.10 through 3.14 with 3.12 required on the RHEL/Rocky RPM path. Each of those has a linked migration guide in the README, which is a good sign for tractability and no sign at all that it is free. If you maintain private plugins, the port is on you.

## Conclusion

Adopt TensorRT if you already target NVIDIA GPUs and need to serve a trained model with lower latency than a framework runtime gives you, and start from pip install tensorrt or the TensorRT OSS build container before touching the source build. Do not adopt it if you need a vendor-neutral runtime, if your models depend on weakly-typed networks, implicit quantization or IPluginV2 plugins, or if you cannot tolerate a migration every major version. Verify first: which exact GA build your CUDA toolkit pairs with, whether your plugins are already IPluginV3, and whether the ONNX operators your graph uses appear in the parser sources in this repository.

## FAQ

### What does TensorRT do?

It is an SDK for high-performance deep learning inference on NVIDIA GPUs. It takes a trained network and produces an engine tuned for a specific GPU, with the repository supplying the ONNX parser, plugin sources and samples.

### What is the difference between CUDA and TensorRT?

CUDA is the toolkit and programming model TensorRT runs on top of, not a substitute for it. The README lists CUDA 13.3.0 or 12.9.0 as build prerequisites for the OSS components.

### How do I install TensorRT?

For Python, the README gives pip install tensorrt and notes you can skip the Build section. Building the OSS components instead requires the TensorRT v11.2.1.2 GA build, cmake >= v3.31, GNU make >= v4.1 and a clone with git submodule update --init --recursive.

### Is TensorRT faster than vLLM?

The README contains no benchmark comparing the two, so no speed claim can be made here. They also sit at different layers: TensorRT is an inference SDK, while vLLM is a serving system, so the comparison depends on what you would otherwise have to build around the engine.

### How do I use TensorRT with PyTorch?

The README lists Torch-TensorRT as one of the four import paths, alongside ONNX, HuggingFace/Optimum and the Network Definition API, and points to the Import Workflows Guide for step-by-step walkthroughs with examples.

## Sources

- [License: Apache-2.0](https://github.com/NVIDIA/TensorRT/blob/main/LICENSE)
- [NVIDIA/TensorRT on GitHub](https://github.com/NVIDIA/TensorRT)
- [Project website](https://developer.nvidia.com/tensorrt)
- [README](https://github.com/NVIDIA/TensorRT/blob/main/README.md)
- [Releases](https://github.com/NVIDIA/TensorRT/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nvidia-tensorrt
