# TensorRT-LLM Backend: Serving LLMs Through Triton Inference Server with NVIDIA Acceleration

> The TensorRT-LLM Backend connects NVIDIA's TensorRT-LLM inference optimization library to Triton Inference Server, enabling HTTP/gRPC serving of large language models with inflight batching, paged attention, speculative decoding, and multi-node deployment. A PyTorch LLM API path added recently removes the engine compilation step for users who want to serve HuggingFace models without pre-building TensorRT engines.

**triton-inference-server/tensorrtllm_backend** — The Triton TensorRT-LLM Backend

- Repository: https://github.com/triton-inference-server/tensorrtllm_backend
- Stars: 945 · Forks: 146
- Language: Unknown
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/triton-inference-server-tensorrtllm-backend

## What the TensorRT-LLM Backend Does and How It Fits the Triton Ecosystem

Triton Inference Server is NVIDIA's model serving platform. It handles HTTP and gRPC traffic, manages request queuing, and coordinates model instances across GPUs. A Triton backend is the plugin layer that connects a specific inference library to Triton's serving infrastructure.

The TensorRT-LLM Backend connects TensorRT-LLM, NVIDIA's library for optimized LLM inference, to Triton. With this backend installed, operators can serve TensorRT-LLM-optimized models through Triton's standard endpoints, gaining Triton features such as model ensembles, dynamic batching, and metrics alongside TensorRT-LLM's inference optimizations like inflight batching, paged KV-cache attention, and quantization.

An important note in the README: the core backend source code and tests have moved to the TensorRT-LLM repository under the triton_backend/ directory. The tensorrtllm_backend repository is the integration layer and documentation entry point; the implementation is in NVIDIA/TensorRT-LLM.

## Two Paths: PyTorch LLM API Versus Compiled TRT-LLM Engines

The README describes two distinct serving paths.

The PyTorch backend (LLM API path) serves any HuggingFace model directly without requiring engine compilation. This is the faster path for getting started. The model configuration is a YAML file where the model key accepts any HuggingFace model ID or local path. All keys in model.yaml map to LLM() constructor arguments, covering KV cache, quantization, and parallelism configuration. The README lists TRT-LLM's supported model architectures, which include Llama, Mistral, Mixtral, Falcon, GPT-2, GPT-J, Gemma, Phi, Qwen, and others, all usable through this path.

The traditional compiled TRT-LLM path requires building a TensorRT engine from the model weights first. This is more involved but unlocks the full set of TensorRT-LLM optimization features including speculative decoding algorithms (Medusa, ReDrafter, Lookahead, Eagle), advanced quantization, and LoRA adapter loading. The inflight_batcher_llm directory in the triton_backend/ subdirectory of the TensorRT-LLM repository contains the C++ implementation of this traditional backend, which handles inflight batching and paged KV-cache directly.

The README recommends the PyTorch path for new users and notes that the Triton model configuration for this path lives at TensorRT-LLM/triton_backend/all_models/llmapi/. Users who need features not yet available in the PyTorch path, such as chunked prefill or specific quantization formats, should look at the compiled engine path and its model configuration documentation in the docs/ directory.

## Quick Start: Launching a Triton Server with the PyTorch Backend

The recommended starting point is the NVIDIA Triton container from NGC. The container tag in the README is 25.12, and the README notes to replace it with the latest tag from NGC:

```bash
docker run --rm -it --net host --shm-size=2g --ulimit memlock=-1 --gpus all \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    nvcr.io/nvidia/tritonserver:25.12-trtllm-python-py3 bash
```

Inside the container, clone TensorRT-LLM and set the model:

```bash
git clone https://github.com/NVIDIA/TensorRT-LLM.git
```

Edit the model.yaml configuration to point to a HuggingFace model:

```yaml
model: TinyLlama/TinyLlama-1.1B-Chat-v1.0
```

All keys in model.yaml map directly to LLM() constructor arguments. For gated models such as Llama, set the HuggingFace token before launching: `export HF_TOKEN=hf_...`.

Then launch the server from the parent of the TensorRT-LLM/ folder:

```bash
python3 TensorRT-LLM/triton_backend/scripts/launch_triton_server.py \
    --model_repo=TensorRT-LLM/triton_backend/all_models/llmapi/
```

The README warns explicitly to run this command from the parent directory, not from inside TensorRT-LLM/. Running from inside causes ModuleNotFoundError: No module named 'tensorrt_llm.bindings'. Once the server is up, test it with a generation request:

```bash
curl -X POST localhost:8000/v2/models/tensorrt_llm/generate \
    -d '{"text_input": "The future of AI is", "sampling_param_max_tokens": 50}' | jq
```

To cancel a request that is still running, send a second POST with stop set to true and the same triton-request-id header value from the original request. The in-flight cancellation is available only in the PyTorch backend path.

## Advanced Deployment: Parallelism, Speculative Decoding, and Multi-Node

The backend supports tensor parallelism, pipeline parallelism, and expert parallelism for distributing large models across multiple GPUs. Multi-node deployments work in two modes: Leader Mode, where one instance coordinates all instances, and Orchestrator Mode, where a dedicated orchestrator process manages worker instances. The README includes a worked example for running multiple LLaMA instances across nodes.

Speculative decoding algorithms are supported through the compiled TRT-LLM engine path. The supported algorithms include Medusa, ReDrafter, Lookahead, and Eagle, covering both draft-model-based and single-model speculative approaches. Decoding mode configuration in the model configuration files selects between top-k, top-p, beam search, and speculative variants.

MIG (Multi-Instance GPU) support allows a single A100 or H100 to run as multiple independent GPU instances, each serving its own model replica. Slurm cluster deployment is also documented: a set of preparation scripts and a Slurm job submission example are in the repository for teams running Triton within HPC environments.

In-flight request cancellation works by sending a second POST to the generate endpoint with 'stop': true and the same request_id used in the original request. The Slurm integration also includes preparation scripts and a job submission example for teams operating within HPC cluster environments.

## Where the TensorRT-LLM Backend Differs from vLLM

vLLM is the most commonly compared alternative. Both serve LLMs on NVIDIA GPUs with PagedAttention and continuous batching. The key difference is integration model: vLLM is a self-contained serving framework with its own HTTP API, while the TensorRT-LLM Backend is a plugin that brings TensorRT-LLM into the Triton ecosystem.

Teams already running Triton for other model types (vision, embeddings, preprocessing) can add LLM serving without adopting a separate server process. Triton's model ensemble, metrics, and management APIs then apply uniformly across all models. vLLM offers a simpler initial setup for teams who only need LLM serving and do not have an existing Triton deployment to build on.

The TensorRT-LLM compiled engine path also gives access to NVIDIA-specific optimizations that go beyond what vLLM exposes, including the speculative decoding algorithms mentioned above and integration with TensorRT's quantization workflow.

## License, Maintenance, and Repository Structure

The TensorRT-LLM Backend is copyright NVIDIA CORPORATION and Affiliates and is released under the Apache-2.0 license, which permits commercial use, modification, and distribution without requiring proprietary dependencies. The last push to the repository was on 2026-09-16. There are no formal GitHub releases for this backend repository; versioning follows Triton's container release cycle (for example, the 25.12 tag in the NGC container URL).

The top-level repository contains the README, a dockerfile/ directory for building the backend from source, a docs/ directory, and a tensorrt_llm symlink or directory pointing to the moved source code. Build-from-source instructions, supported model lists, and configuration documentation are in the README and the docs/ directory. The .github/ directory holds CI configuration. A .pre-commit-config.yaml is present, meaning the project enforces code style checks at commit time.

Metrics are available through Triton's standard metrics endpoint, covering request counts, latency, and GPU utilization. Benchmarking instructions and a testing guide are included in the README's table of contents. The Benchmarking section covers performance measurement tooling, and the Testing section covers how to run the backend's own test suite against a running Triton instance.

## Conclusion

The TensorRT-LLM Backend is the right path for teams deploying LLMs on NVIDIA hardware through Triton's HTTP and gRPC serving infrastructure, especially when they need inflight batching, speculative decoding, or multi-node scaling. The PyTorch backend path is the lowest-friction entry point; the compiled TensorRT-LLM engine path gives access to more advanced optimization features. Before starting, note that the core backend source code has moved to the TensorRT-LLM repository's triton_backend/ directory, so keep both repositories in sync when building from source or reading implementation details.

## FAQ

### What is TensorRT LLM used for?

TensorRT-LLM is NVIDIA's library for optimizing and accelerating LLM inference on NVIDIA GPUs. The TensorRT-LLM Backend makes those optimized models available through Triton Inference Server's HTTP and gRPC endpoints.

### What is the difference between TensorRT and TensorRT LLM?

TensorRT is NVIDIA's general deep learning inference optimizer and runtime. TensorRT-LLM is a higher-level library built specifically for large language models, adding features like inflight batching, paged attention, and speculative decoding that are not part of the base TensorRT toolkit.

### Is TensorRT LLM open source?

Yes. Both TensorRT-LLM and the TensorRT-LLM Backend are released under the Apache-2.0 license, which permits commercial use, modification, and redistribution. The source is on GitHub under the NVIDIA and triton-inference-server organizations.

## Sources

- [Issues](https://github.com/triton-inference-server/tensorrtllm_backend/issues)
- [License: Apache-2.0](https://github.com/triton-inference-server/tensorrtllm_backend/blob/main/LICENSE)
- [README](https://github.com/triton-inference-server/tensorrtllm_backend/blob/main/README.md)
- [triton-inference-server/tensorrtllm_backend on GitHub](https://github.com/triton-inference-server/tensorrtllm_backend)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/triton-inference-server-tensorrtllm-backend
