Torch-TensorRT: compiling PyTorch models to TensorRT on NVIDIA GPUs
PyTorch/TorchScript/FX compiler for NVIDIA GPUs using TensorRT
At a glance
- What is it?
- Torch-TensorRT is a PyTorch compiler that lowers TorchScript and torch.export graphs to TensorRT engines on NVIDIA hardware. It is a one-line torch.compile backend, but the version matrix and the platform table decide whether it fits your deployment.
- Who is it for?
- Adopt Torch-TensorRT if you ship PyTorch inference on Linux AMD64, Linux SBSA or Windows GPUs and can pin CUDA 13.2, TensorRT 11.2.1.2 and a matching torch build; the export path plus libtorch deployment is the reason to pick it over a serving framework. Skip it if you are on ppc64le, if you need an ahead-of-time path on Windows (the platform table lists Dynamo only there), or if your graph breaks often enough that partitioning overhead eats the gain.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap Torch-TensorRT fills between PyTorch and TensorRT
TensorRT is a separate inference runtime with its own engine format and its own API. Getting a trained PyTorch model into it normally means rewriting the model in the TensorRT network API or going through an intermediate format, then re-validating numerics. Torch-TensorRT exists to remove that rewrite. It takes a PyTorch or TorchScript module, partitions the graph, sends the supported subgraphs to TensorRT, and leaves the rest in PyTorch, so the result is still a PyTorch module you can call with ordinary tensors.
The intended audience is narrow and specific: engineers who already have a working PyTorch model, already run on NVIDIA GPUs, and care about inference latency or throughput rather than training. The README frames the payoff as inference latency reduced by up to 5x compared to eager execution, described as achievable in one line of code. That claim is a vendor figure from the project's own README, and it is a ceiling for favorable models, not a default. Models with dynamic shapes, heavy control flow or unsupported operators will land well below it, and the partitioning machinery is what determines where you land.
How the compiler partitions a graph and where TensorRT takes over
The mechanism is graph capture followed by partitioning. In the torch.compile path, Torch-TensorRT registers itself as a backend named tensorrt. Dynamo captures the model into a graph, the backend decides which segments TensorRT can execute, compiles those segments into TensorRT engines, and wraps the remainder back into PyTorch calls. Compilation happens on the first invocation, which is why the README annotates the first call as compiled on first run and the second as fast.
The export path is different in timing, not in principle. torch_tensorrt.compile with ir="dynamo" performs the same lowering ahead of time, then torch_tensorrt.save serializes the result. The README notes that PyTorch only supports a Python runtime for an ExportedProgram, so a C++ deployment needs the TorchScript output format instead. That distinction matters more than it looks: choosing the .ep format locks you into a Python runtime, while .ts is what the C++ loading example uses through torch::jit::load.
Both paths depend on the same underlying machinery, so a graph break is not an error, it is a split. The cost is a boundary crossing between the TensorRT engine and the PyTorch interpreter. A model that breaks once around a single unsupported op will do well. A model that breaks in a loop will spend its time at boundaries.
Installing Torch-TensorRT and running a first compiled model
Stable builds are on PyPI. The README gives a single command, with no version pin, which is worth noticing because the dependency list is explicit about what the tests verify.
pip install torch-tensorrtNightly builds come from the PyTorch package index instead. The README uses the cu130 index for nightlies, while the dependency table lists CUDA 13.2 for verified test cases, so the two are not the same channel.
pip install --pre torch-tensorrt --index-url https://download.pytorch.org/whl/nightly/cu130 --extra-index-url https://pypi.org/simpleThe README also points to the NVIDIA NGC PyTorch container as a ready-to-run distribution that already carries matching dependencies and example notebooks. If you do not want to resolve the version matrix yourself, that is the path the project itself recommends for advanced setups.
For a first real run, the torch.compile route needs no export step. Import torch_tensorrt, define the model in eval mode on CUDA, build an example input with the shape you will actually serve, and pass the backend name.
import torch
import torch_tensorrt
model = MyModel().eval().cuda()
x = torch.randn((1, 3, 224, 224)).cuda()
optimized_model = torch.compile(model, backend="tensorrt")
optimized_model(x)
optimized_model(x)The first call triggers compilation and will be slow; the second is the one that reflects steady-state latency. If you need ahead-of-time compilation or a C++ target, the README's export example compiles with ir="dynamo", saves both a .ep and a .ts file, and the C++ side loads the TorchScript file with torch::jit::load. Note the README's own warning in that snippet: the Python runtime is the only option for an ExportedProgram.
Where Torch-TensorRT is the wrong tool
The platform table is the first hard boundary. Linux ppc64le is listed as not supported. Windows is listed as supported for Dynamo only, which means the ahead-of-time export workflow described in the README is not the Windows story. Jetson GPU and DLA support is source compilation on JetPack-4.4 and later, not a wheel install, so the one-line pip command does not cover that platform.
The second boundary is version coupling. The dependency list names Bazel 8.1.1, Libtorch 2.15.0.dev, CUDA 13.2 and TensorRT 11.2.1.2 as the versions used to verify test cases, with an explicit note that other versions may work but tests are not guaranteed to pass. The build backend in pyproject.toml pins torch to >=2.15.0.dev,<2.16.0. That is a narrow window for a production environment. If your organization is on an older CUDA or a stable torch release, you are outside the verified set and the fallback is source compilation, which the repository supports through setup.py, Bazel files and a justfile but which is a build project rather than an install.
Third, this is an inference-only compiler for NVIDIA hardware. If your target is CPU inference, AMD GPUs, or a non-NVIDIA accelerator, nothing here applies. There is also no documented rollback story in the README: once a model is compiled and serialized, the README does not describe how to revert to an eager artifact, so keep the original checkpoint and the compilation step as a reproducible pipeline rather than a one-off.
Torch-TensorRT against vLLM and against hand-written TensorRT
vLLM appears often in the same searches, and the difference in approach is structural. vLLM is a serving system: it owns the model, the scheduler, the KV cache and the HTTP surface, and it optimizes for batched LLM throughput. Torch-TensorRT is a compiler that produces a module you keep using inside PyTorch. It gives you no server, no request batching policy and no continuous-batching scheduler. If your problem is serving many concurrent LLM requests, a serving framework is the closer fit. If your problem is a fixed model where you want the graph itself to run faster, and you want to keep calling it as a PyTorch module or load it from C++, Torch-TensorRT addresses that and vLLM does not.
The other real alternative is going to TensorRT directly with the network API or an ONNX intermediate. That gives you full control over layer fusion and engine configuration, and it is the only option when a model has operators Torch-TensorRT cannot lower. The cost is that you reimplement the forward pass, you own the numerics comparison against the PyTorch reference, and every model change becomes a rewrite. Torch-TensorRT's partitioning is the trade: less control, but the unsupported parts keep running in PyTorch instead of blocking the whole conversion.
Release cadence, deprecation window and licence terms
The repository is not archived, and the last push was on 2026-09-10, the same day v2.14.0 was released. The two prior releases were v2.13.0 on 2026-07-28 and v2.12.1 on 2026-06-09. That is a roughly six-to-eight-week cadence across the three most recent tags, which is the concrete number to plan around rather than any general statement about activity.
The upgrade cost is dominated by the version matrix, not by API churn. Each release is tied to a torch range, a CUDA version and a TensorRT version, so an upgrade is usually a coordinated move of all four. The deprecation policy, which the README says begins with version 2.3, gives a six-month migration period: deprecated APIs warn at runtime, keep working through the window, and are removed in a manner consistent with semantic versioning. Six months is a workable budget, but it assumes you can move torch and CUDA on that schedule, which many production environments cannot. Plan the compilation step as something you can rebuild, not something you install once.
Licensing is BSD-3-Clause, per the LICENSE file and the pyproject.toml classifier. That is permissive and carries no copyleft obligation on your own code. It says nothing about the licences of CUDA, TensorRT or the NGC container you pull in alongside it, which are separate terms you should read yourself. Nothing here is legal advice.
Editorial conclusion
Adopt Torch-TensorRT if you ship PyTorch inference on Linux AMD64, Linux SBSA or Windows GPUs and can pin CUDA 13.2, TensorRT 11.2.1.2 and a matching torch build; the export path plus libtorch deployment is the reason to pick it over a serving framework. Skip it if you are on ppc64le, if you need an ahead-of-time path on Windows (the platform table lists Dynamo only there), or if your graph breaks often enough that partitioning overhead eats the gain. Before committing, verify your exact torch and TensorRT versions against the dependency list, and check whether your ops fall back to PyTorch rather than running as TensorRT engines.
Frequently asked questions
How do I use TensorRT with PyTorch?
Install torch-tensorrt, then pass backend="tensorrt" to torch.compile. The README shows the model being set to eval mode on CUDA, compiled on the first call, and running fast on the second. For ahead-of-time compilation, use torch_tensorrt.compile with ir="dynamo" and serialize the result.
How do I install TensorRT for PyTorch?
The README gives pip install torch-tensorrt for the stable release, and a --pre variant against the PyTorch nightly index for nightly builds. It also notes that Torch-TensorRT ships in the NVIDIA NGC PyTorch container with dependencies and example notebooks included.
How does Torch-TensorRT work?
It captures a PyTorch or TorchScript graph, partitions it, compiles the supported segments into TensorRT engines and leaves the rest executing in PyTorch. Compilation happens on the first invocation in the torch.compile path, which is why the README marks the first call as compiled and the second as fast.
Is TensorRT faster than vLLM?
The two are not the same kind of tool, so the comparison does not resolve cleanly. vLLM is a serving system with its own scheduler and batching, while Torch-TensorRT compiles a model into a module you keep calling from PyTorch or libtorch. The README's performance claim of up to 5x is measured against eager PyTorch execution, not against vLLM.
How do I install TensorRT on Windows?
The README gives the same pip install torch-tensorrt command without a platform distinction, but the platform support table lists Windows as supported for Dynamo only. That means the ahead-of-time export and C++ deployment workflow described in the README is not the Windows path.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/pytorch-tensorrt)