Model or dataset
NVIDIA/TensorRT-LLM avatar
NVIDIA/TensorRT-LLM

TensorRT-LLM: NVIDIA's Inference Stack for LLMs and Visual Gen Models

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

14,743 stars2,781 forksPythonNOASSERTION

At a glance

What is it?
TensorRT-LLM is NVIDIA's Python API and runtime for optimized LLM inference on NVIDIA GPUs. It rewards teams with NVIDIA hardware and a fixed model set, and punishes anyone who wants a portable install.
Who is it for?
Adopt TensorRT-LLM if you serve a fixed set of models on NVIDIA GPUs and can pin Python 3.10 or 3.12, CUDA 13.2.1 and torch 2.12.0 to 2.14.0a0 in a container. Do not adopt it if you need CPU-only inference, a vendor-neutral stack, or a release branch that is not a release candidate.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem TensorRT-LLM solves, and for whom

Serving a large language model on a GPU is not the same as running one. A PyTorch forward pass gets you tokens; it does not get you batched continuous decoding, paged KV cache, quantized GEMM kernels, or an execution graph that avoids per-step launch overhead. TensorRT-LLM exists to supply those pieces as one package. The README describes it as optimizing inference for LLMs and Visual Gen models with specialized kernels for common operations, an efficient runtime, and a pythonic framework for customization.

The audience is narrow and specific. You need NVIDIA GPUs, and the repository topics name Blackwell explicitly, alongside cuda, llm-serving, moe and pytorch. If your deployment target is a single consumer card, a CPU box, or a mixed accelerator fleet, the project's own framing does not address you. If you are running mixture-of-experts models at scale, or serving video generation models, the technical blog list shows where NVIDIA is spending effort: GEMM quantization and skip softmax attention for video generation, expert parallelism across NVL72 racks, and DeepSeek variants on Blackwell.

One distinction trips people up constantly, and it is worth stating plainly. TensorRT is the general inference compiler and runtime for neural networks. TensorRT-LLM is the layer above it, aimed at large language models and visual generation, with its own Python API and its own C++ and Python runtimes for orchestrating execution. They are related, not interchangeable, and the repository is the second one.

How the runtime is put together

The repository layout tells you most of the architecture. There is a tensorrt_llm/ Python package that holds the API, a cpp/ directory for the C++ runtime, and a triton_backend/ directory for serving through NVIDIA's Triton Inference Server. The README says the project contains components to create Python and C++ runtimes that orchestrate the inference execution. So the flow is: define or convert a model, let the stack apply its kernels and quantization, then execute through one of those runtimes.

The examples directory is the clearest map of what is actually supported. It contains subdirectories for llm-api, quantization, disaggregated serving, wide expert parallelism (wide_ep), DWDP, sparse attention, KV cache compression, ngram speculative decoding, visual_gen, scaffolding, ray_orchestrator, trtllm-eval and opentelemetry. That list is the feature surface. Disaggregated serving means splitting prefill and decode across separate workers. wide_ep and dwdp are the multi-node parallelism strategies for very large models. The opentelemetry example indicates tracing hooks exist for production observability.

Dependencies in requirements.txt reveal the intended environment. The file pins CUDA Python at >=13, torch between 2.12.0a0 and 2.14.0a0, transformers at exactly 5.5.4, and NCCL between 2.29.7 and 2.30.7. Several pins carry inline comments explaining why they cannot move yet. The torch floor comment notes that torch==2.13.0 pins triton==3.7.1, which conflicts with another pin, so the floor stays low until public triton catches up. This is a stack that moves with the NGC container cadence, not with PyPI.

Installing TensorRT-LLM and running a first model

The README points to a quick-start guide under the documentation site rather than embedding install steps, and the repository ships a docker/ directory and a CONTAINER_SOURCE.md file. The container path is the one the project's own dependency comments assume, because the version pins are written against NGC PyTorch release notes. The README badges state Python 3.12 and 3.10, CUDA 13.2.1, torch 2.12.0, and release 1.3.0rc26.

If you install from source, the packaging is standard setuptools. The build system requires setuptools >= 64 and pip >= 24, and requirements.txt carries an extra index for the CUDA 13 PyTorch wheels. A minimal invocation looks like this:

bash
pip install -r requirements.txt
pip install -e .

The first command pulls the pinned dependency set, including the PyTorch CUDA 13 wheels from the extra index URL declared at the top of the file. The second installs the local tensorrt_llm package in editable mode. Expect a long resolve; transformers is pinned to an exact version and the torch range spans alpha builds.

Once installed, the llm-api example is the shortest path to a working generation. It is a Python entry point, and the model weights come from Hugging Face checkpoints that the stack converts. The exact model list lives in examples/models, and the README's blog list names DeepSeek-V4, DeepSeek-V3.2, DeepSeek R1, GPT-OSS-120B, Kimi K3 and Gemma among the validated targets. Pick a model from that set rather than assuming an arbitrary checkpoint will convert cleanly.

For a server rather than a script, the examples/serve directory and the triton_backend directory are the two documented routes. The opentelemetry example shows how to wire tracing if you need per-request visibility. The README does not document a rollback procedure for a failed conversion, so keep your original checkpoint.

Where the version pinning becomes your problem

The dependency file is honest about its own fragility, and that honesty is the most useful thing in the repository for an adopter. Comments in requirements.txt record that the torch floor cannot rise because of a triton conflict, that the NCCL window is bounded on both ends by specific NGC and wheel combinations, and that diffusers is capped below 0.41 pending a recheck of MiniMax H3's experimental imports. Each of those is a constraint you inherit.

The practical consequence is that TensorRT-LLM is not a library you add to an existing environment. It is an environment. If your application already pins a different transformers version, or a torch build outside that alpha-to-alpha range, you will be resolving conflicts rather than serving models. The same applies to CUDA: the badge says 13.2.1, and the requirements pull CUDA 13 wheels. A host with CUDA 12 drivers is not the target.

There is a second limitation that has nothing to do with dependencies. The recent releases listed are v1.3.0rc26, v1.3.0rc25 and v1.3.0rc24, all release candidates, with the newest dated 2026-09-09. The repository is not archived and the last push was on 2026-09-10, so development is ongoing, but the release channel an adopter would land on is a candidate build. Teams that require a stable tagged release should verify what the current stable line is before planning a deployment around it.

Finally, the licence metadata is inconsistent. The repository metadata reports NOASSERTION, while the README badge links to the LICENSE file and labels it Apache 2. Read LICENSE yourself. Do not take a badge as the answer, and do not treat this paragraph as legal advice.

TensorRT-LLM versus vLLM: different starting points

The comparison people search for most is TensorRT-LLM against vLLM, and the difference is architectural rather than a matter of tuning. vLLM is a Python serving engine built around paged attention, and it is designed to run broadly across GPU vendors and to be installed like a normal Python package. TensorRT-LLM starts from NVIDIA's kernel and runtime stack, which is why its dependency file reads like a container manifest and why its parallelism features (wide_ep, DWDP, disaggregated serving) are described in terms of NVLink and NVL72 racks.

That means the two projects optimize for different risks. Choosing vLLM minimizes the risk that your environment will not resolve and the risk that a model will not be supported. Choosing TensorRT-LLM accepts those risks in exchange for kernel-level control over the execution path, including quantization schemes, sparse attention and speculative decoding variants that the project documents in its own blog series. If your workload is a handful of well-known models on a fixed NVIDIA fleet, the second trade is reasonable. If your workload is many models, changing weekly, on heterogeneous hardware, it is not.

A third option worth naming is the Triton Inference Server route, which TensorRT-LLM itself supports through triton_backend/. That is not an alternative to TensorRT-LLM so much as a different deployment surface for it, and it is the right choice when you already run Triton for other models and want one serving layer.

Maintenance cost and what to verify before you commit

The upgrade cost is the container cadence. Because requirements.txt ties torch, NCCL and CUDA versions together with comments referencing specific NGC releases, moving forward generally means moving the whole base image rather than bumping one package. The last push was on 2026-09-10 and the newest listed release candidate is dated 2026-09-09, so the project ships frequently; that frequency is also the cost, because each release candidate is a new pin set to validate against your models.

On licensing, the repository metadata says NOASSERTION and the README badge says Apache 2. Those two statements do not agree, and only the LICENSE file resolves it. If your organization requires an approved licence identifier before adoption, that file is the artifact to review, and the ATTRIBUTIONS-Python.md and ATTRIBUTIONS-CPP-x86_64.md files in the repository root list bundled third-party components that may carry their own terms.

What to verify first, concretely: confirm your target model appears in examples/models or in the blog list; confirm your GPU generation is covered, since Blackwell is named in the topics and the recent blogs; and confirm you can build the pinned environment in a container before you write application code against the API. The pyproject.toml shows the project lints with ruff at a 100-character line length and maintains a legacy-files.txt baseline, which tells you the codebase is mid-migration on style, not that it is unstable. Judge it on the pins and the release channel instead.

Editorial conclusion

Adopt TensorRT-LLM if you serve a fixed set of models on NVIDIA GPUs and can pin Python 3.10 or 3.12, CUDA 13.2.1 and torch 2.12.0 to 2.14.0a0 in a container. Do not adopt it if you need CPU-only inference, a vendor-neutral stack, or a release branch that is not a release candidate. Before committing, check the LICENSE file because the repository metadata reports NOASSERTION while the README badge says Apache 2, and confirm which v1.3.0rc build your model was validated against.

Frequently asked questions

What is TensorRT-LLM?

It is NVIDIA's Python API and set of runtimes for running large language models and visual generation models efficiently on NVIDIA GPUs. The README describes specialized kernels for common operations, an efficient runtime, and a pythonic framework for customization.

What is the difference between TensorRT and TensorRT-LLM?

TensorRT is the general inference compiler and runtime for neural networks. TensorRT-LLM sits above it and targets large language models and visual generation, adding its own Python API plus Python and C++ runtimes for orchestrating execution.

Is TensorRT-LLM faster than vLLM?

The repository does not publish a head-to-head comparison against vLLM, so no number can be quoted here. The two differ in approach: TensorRT-LLM builds on NVIDIA kernels and multi-GPU strategies such as wide expert parallelism and disaggregated serving, while vLLM is a Python serving engine aimed at broad portability.

Is TensorRT-LLM free?

The README badge labels the project Apache 2 and links to the LICENSE file, but the repository metadata reports NOASSERTION. Read the LICENSE file and the ATTRIBUTIONS files before relying on either label.

How do I install TensorRT-LLM?

The README points to a quick-start guide on the documentation site and the repository ships a docker/ directory plus CONTAINER_SOURCE.md. Installing from source uses the pinned requirements.txt and then an editable install of the package.

How do I use TensorRT-LLM?

Start from the examples directory: llm-api for a script-level generation path, examples/serve for a server, and triton_backend for Triton Inference Server. The opentelemetry example covers request tracing.

Official sources

  1. Issues
  2. NVIDIA/TensorRT-LLM on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nvidia-tensorrt-llm.svg)](https://hysenlabs.com/projects/nvidia-tensorrt-llm)