TensorRT Edge-LLM: NVIDIA's C++ Inference Runtime for Jetson, DRIVE and DGX Spark
High-performance, light-weight C++ LLM and VLM Inference Software for Physical AI
At a glance
- What is it?
- TensorRT Edge-LLM exports Hugging Face checkpoints to ONNX and builds C++ TensorRT engines for edge devices. It is a deployment toolchain, not a general-purpose serving stack, and its support matrix decides whether you can use it at all.
- Who is it for?
- Adopt TensorRT Edge-LLM if you are deploying a supported model to a supported NVIDIA edge platform and you can commit to the pinned Python stack in requirements.txt and the release cadence. Do not adopt it for a cloud GPU fleet, for a model outside the supported-models matrix, or if you need a stable server API, because the OpenAI-compatible server is labelled experimental.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 26 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What TensorRT Edge-LLM is for, and who should care
The README describes TensorRT Edge-LLM as NVIDIA's C++ inference runtime for text, vision, audio, speech, and action models on NVIDIA Jetson, NVIDIA DRIVE, and NVIDIA DGX Spark. That sentence is the whole scope. This is not a general serving framework that happens to run on small machines. It is a deployment path for a fixed set of devices, and the device list is the first thing that decides whether the project is relevant to you.
The intended user is an engineer shipping a model onto a device that sits in a robot, a vehicle, or a workstation on a desk. The repository layout reflects that: cpp/ holds the deployment runtimes, examples/ is split into llm, multimodal, omni, python, accuracy and utils, and tensorrt_edgellm/ is the Python side that builds the artifacts. The Python code exists to produce engines, not to serve traffic in production. The runtime that matters is C++.
If your target is a cloud GPU and your problem is throughput across many concurrent users, this is the wrong layer. If your target is a Jetson module inside a product and your problem is getting a specific checkpoint to run at acceptable latency with a small memory footprint, the project is aimed exactly at you.
The ONNX export path and the experimental direct engine builder
The README states that the supported frontend exports Hugging Face checkpoints to ONNX for C++ engine building, and that an experimental direct frontend builds engines from checkpoints without ONNX. Both paths use the same C++ deployment runtimes. That is the central architectural fact: two ways to produce the artifact, one runtime to consume it.
The default path is a three-stage pipeline. A checkpoint is exported to ONNX, TensorRT compiles that graph into an engine for the specific device, and the C++ runtime loads the engine and executes inference. Quantization sits alongside this, with a dedicated page in the user guide for creating quantized checkpoints for tensorrt_edgellm. The pyproject.toml mirrors the split: the base package depends only on cuda-python and numpy, while the export and tools extras pull in torch, transformers, onnx, onnxscript, safetensors and onnx-graphsurgeon, and tools adds nvidia-modelopt for quantization work.
The direct builder removes the ONNX file from the middle of that chain. The documentation calls it experimental, and the word is doing real work here. An ONNX intermediate is inspectable: you can open the graph, compare it against the reference implementation, and find where a conversion went wrong. A direct build gives you fewer artifacts to debug. For a model already on the supported list, the ONNX path is the one the project recommends.
Installing tensorrt_edgellm and running a first inference
The README does not put install commands in the repository root. It points to the Installation page in the hosted documentation, which covers setting up quantization, tensorrt_edgellm, and the C++ runtime, and to a Quick Start Guide that the README says runs your first inference in about 15 minutes. Treat those two pages as the source of truth, because the exact steps depend on your JetPack or DriveOS version.
What the repository does pin is the Python dependency set. requirements.txt is short and fully version-locked, which tells you the export tooling is tested against one combination rather than a range.
torch==2.13.0
transformers==5.14.1
onnx==1.19.0
onnxscript==0.7.1
safetensors==0.8.0
numpy==2.2.6
onnx-graphsurgeon==0.6.1The package itself is named tensorrt-edgellm on the Python side and declares requires-python >=3.10, with classifiers for 3.10 through 3.12. The base install is deliberately thin, and pyproject.toml lists no console script, so the export and build entry points come from the modules described in the documentation rather than from an installed command.
The extras are where the weight is. export pulls the ONNX and PyTorch stack, tools adds nvidia-modelopt, datasets, peft, librosa and others for quantization and evaluation, and server pulls fastapi, uvicorn, huggingface-hub and av for the experimental HTTP server.
Before running anything, check the Support Matrix for your platform, JetPack, DriveOS, CUDA and TensorRT versions, then look up your checkpoint ID in Supported Models. The user guide also documents request and chat template formats, which you will need once you move past the quick start.
Where the toolchain will fight you
The obvious limitation is the support matrix. Platforms, JetPack versions, DriveOS versions, CUDA and TensorRT versions are enumerated, and a model is only usable if it appears in the supported-models matrix with a checkpoint ID. This is not a framework where you point it at an arbitrary Hugging Face repository and expect a working engine. The release notes show the model list growing release by release, which also means the list is a moving target rather than a stable contract.
The second limitation is the dependency pinning. torch==2.13.0, transformers==5.14.1, numpy==2.2.6 and onnx==1.19.0 are exact pins, not floors. If your existing environment holds a different transformers or numpy, the export extra will conflict with it. Installing into a dedicated environment is the practical answer, but it means the build tooling cannot share a venv with unrelated work.
The third is the experimental surface. The OpenAI-compatible server and the direct engine builder are both labelled experimental in the README and the news entries. The 0.10.1 release redesigned the server for faster cold launches and lower memory usage, which is a good sign for the direction and a warning about API stability at the same time. If you need a server whose request schema will not move under you, this is not it yet. The README does not document rollback or downgrade procedures between releases.
How it differs from TensorRT-LLM and vLLM
TensorRT-LLM is the closest name and the easiest confusion. TensorRT Edge-LLM is a separate repository with its own runtime, its own documentation site, and its own support matrix centred on Jetson, DRIVE and DGX Spark. The shared element is TensorRT itself as the compilation layer. If you have used TensorRT-LLM on data-centre GPUs, the mental model of build-then-serve transfers, but the code, the model list and the platform constraints do not.
Against vLLM the difference is sharper. vLLM is a Python serving engine built around paged attention and continuous batching on server GPUs. TensorRT Edge-LLM compiles a model into a device-specific engine and executes it from C++. The Python layer is a build tool, and the server component is explicitly experimental. Choosing between them is mostly a question of where the model runs: a rack of GPUs favours vLLM, a Jetson module inside a product favours this project.
The comparison that matters most is against doing nothing. If your model already runs acceptably on the target device through another runtime, the cost of adopting this one is the pinned Python environment, the export and quantization steps, and the release cadence. The benefit is a C++ runtime and TensorRT engines tuned for the device. Whether that trade is worth it depends on your latency and memory budget, which the project's own performance benchmarks page is the place to check.
Maintenance, releases and licence
The repository is not archived, and the last push was on 2026-09-03, the same day v0.10.1 was released. The release history in the README shows a steady cadence: 0.9.0 and 0.9.1 in July 2026, 0.10.0 in August, 0.10.1 in September. Each release adds models or features, and the news entries name them specifically, from the Gemma 4 family and Qwen3-Omni to Nemotron-3.5 Lightning with MTP and DFlash, Cosmos3-Edge, DiffusionGemma and Nemotron-3.5-ASR.
That cadence is the upgrade cost. The pinned dependency set moves with releases, and the experimental components are being redesigned between versions, as the 0.10.1 server rewrite shows. Budget for re-running your export and quantization pipeline when you move versions, and keep the exact package versions recorded alongside the engines you build.
The licence is Apache-2.0, declared in both the LICENSE file and the pyproject.toml license field, with license-files listing LICENSE and 3rdParty/nlohmannJson/LICENSE.MIT. The presence of a bundled MIT-licensed third-party component means the distribution carries more than one licence text. Apache-2.0 is permissive and includes a patent grant, but this is not legal advice; if you are shipping a product, have your own counsel review the notices.
Editorial conclusion
Adopt TensorRT Edge-LLM if you are deploying a supported model to a supported NVIDIA edge platform and you can commit to the pinned Python stack in requirements.txt and the release cadence. Do not adopt it for a cloud GPU fleet, for a model outside the supported-models matrix, or if you need a stable server API, because the OpenAI-compatible server is labelled experimental. Verify first that your exact JetPack or DriveOS version appears in the support matrix, then run the Quick Start Guide end to end on the target device before porting any application code.
Frequently asked questions
Is TensorRT Edge-LLM free?
The repository is licensed under Apache-2.0, declared in both the LICENSE file and the pyproject.toml license field, so the software itself is free to use under those terms. The hardware it targets, such as Jetson, DRIVE and DGX Spark platforms, is separate.
What is the difference between TensorRT and TensorRT Edge-LLM?
TensorRT is the compilation layer that turns a graph into a device-specific engine. TensorRT Edge-LLM is NVIDIA's C++ inference runtime built on top of it for text, vision, audio, speech, and action models on Jetson, DRIVE and DGX Spark, with its own documentation site and support matrix.
Is TensorRT Edge-LLM faster than vLLM?
The repository points to its own Performance Benchmarks page rather than making a comparison against vLLM. The two take different approaches: vLLM is a Python serving engine for server GPUs, while TensorRT Edge-LLM compiles models into TensorRT engines executed from C++ on edge devices.
Is TensorRT a part of NVIDIA?
Yes. TensorRT Edge-LLM is published under the NVIDIA organisation on GitHub, and the README describes it as NVIDIA's C++ inference runtime for NVIDIA edge platforms.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvidia-tensorrt-edge-llm)