# NVIDIA Dynamo: the orchestration layer above vLLM, SGLang and TensorRT-LLM

> Dynamo is a Rust-and-Python inference serving stack that coordinates multiple GPUs and nodes instead of replacing your inference engine. It is worth reading only if your bottleneck is cluster-level, not single-GPU.

**ai-dynamo/dynamo** — A Datacenter Scale Distributed Inference Serving Framework

- Repository: https://github.com/ai-dynamo/dynamo
- Website: https://docs.nvidia.com/dynamo/latest
- Stars: 8,190 · Forks: 1,643
- Language: Rust
- License: NOASSERTION
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/ai-dynamo-dynamo

## What Dynamo solves, and the cluster size where it starts to matter

A single inference engine optimizes one GPU or one node. The README is explicit that Dynamo is the orchestration layer above inference engines and does not replace SGLang, TensorRT-LLM or vLLM. The problem it targets is what happens once you have many GPUs: prefill and decode have different compute profiles, replicas repeat the same prefill work, and scaling one phase forces you to scale the other.

The README lists the situations where this matters: serving LLMs across multiple GPUs or nodes, wanting KV-aware routing to avoid redundant prefill computation, needing to scale prefill and decode independently, wanting autoscaling that meets latency SLAs at minimum total cost of ownership, and needing fast cold starts for new replicas. The audience is platform and inference engineers who already operate a serving stack.

The same list contains the boundary. If you run a single model on a single GPU, the README states your inference engine alone is probably sufficient. That is unusually direct for a project README, and it is the right frame: Dynamo is infrastructure for people whose problem has already outgrown one machine.

## Disaggregated serving, KV-aware routing and the Rust workspace behind them

The architecture separates prefill and decode into independently scalable GPU pools, which the README calls disaggregated prefill/decode. Each phase then runs on hardware tuned for its workload. A router sits in front of those pools and makes decisions using KV cache state rather than round-robin, so a request that shares a prefix with an earlier one can be sent where the cache already lives.

The repository layout matches that description. Cargo.toml defines a workspace whose members include lib/kv-router, lib/kv-hashing, lib/kvbm-common, lib/kvbm-config, lib/kvbm-engine, lib/kvbm-kernels, lib/kvbm-consolidator, lib/kvbm-logical and lib/kvbm-physical. The KVBM modules are the multi-tier KV cache manager; the router and hashing crates are the routing engine named in the repository topics. Sidecar crates exist per backend: lib/sidecar/vllm, lib/sidecar/sglang and lib/sidecar/trtllm, with matching mock servers under lib/mocker/servers/.

There is also a planner component, referenced in the README as the SLA-Based Planner, and a global_planner example under examples/. The Rust workspace uses edition 2024 and resolver 3, and the workspace version is 1.6.0, matching the version in pyproject.toml. Python is the extensibility surface: the wheel is named ai-dynamo and depends on ai-dynamo-runtime of the same version.

## Installing Dynamo from PyPI and running a first deployment

The Python package is published as ai-dynamo, and pyproject.toml declares requires-python >=3.10. The dependencies listed there include ai-dynamo-runtime, aiohttp, transformers, kubernetes, prometheus_client, msgspec, zstandard and pyzmq. The package name and the version pin on the runtime are the two facts to carry into an environment file:

```toml
[project]
name = "ai-dynamo"
version = "1.6.0"
requires-python = ">=3.10"
dependencies = [
    "ai-dynamo-runtime==1.6.0",
]
```

Backend support is an optional extra rather than a separate package. The trtllm extra in pyproject.toml pulls in uvloop and tensorrt-llm, and the same file pins tensorrt-llm==1.3.0rc26:

```toml
[project.optional-dependencies]
trtllm =[
    "uvloop",
    "tensorrt-llm==1.3.0rc26",
]
```

Beyond the wheel, the README points at three places for runnable material: the recipes directory, the examples directory, and NGC containers under the nvidia/ai-dynamo team. The examples directory contains deployments/, backends/, router/, diffusers/, rl/ and power-aware-budget/ subdirectories, and the router examples include custom policy crates such as soft-pin-repin and simple-filter-score-pick. The README does not print a single canonical launch command, so the first real use is to pick the example matching your engine and follow its local README rather than improvising flags.

## Where Dynamo is the wrong tool, and what the support matrix does not promise

The clearest failure mode is scope. A team running one model on one GPU gets nothing from disaggregation, routing or a planner; the README says so itself. Adding Dynamo there means operating a router, a planner and sidecars for a workload that has no coordination problem.

The second limit is per-backend feature parity. The README's support table marks disaggregated serving, KV-aware routing and the SLA-based planner as available for SGLang, TensorRT-LLM and vLLM. KVBM is different: it is marked available for TensorRT-LLM and vLLM but shown as under construction for SGLang. If your deployment is SGLang-based and you depend on the multi-tier KV cache, the table is telling you to wait or to change engines.

The third is that the README's headline results (7x throughput per GPU on GB200 NVL72, 7x faster model startup via ModelExpress weight streaming, 2x faster time to first token on Qwen3-Coder 480B, 80% fewer SLA breaches, 750x throughput on GB300 NVL72) are attributed to external sources such as InferenceX, a Baseten benchmark and an Alibaba APSARA talk. They are cited claims about specific hardware and models, not guarantees that transfer to your cluster.

## Dynamo against running vLLM or SGLang on their own

The real alternative is not a different orchestration framework; it is the inference engine by itself. vLLM and SGLang both ship continuous batching, paged attention and their own serving front ends, and both are listed as backends Dynamo drives. The difference is where coordination lives. Running vLLM alone means each replica is independent: it does not know what another replica has cached, it cannot move prefill work to a different pool, and scaling is something you do at the replica count.

Dynamo moves those decisions into a routing and planning layer. KV-aware routing becomes a cluster-level decision rather than a per-replica one, prefill and decode become separately scalable pools, and the planner adjusts capacity against latency SLAs. The trade is operational: you now run router, planner and sidecar components in addition to the engine, and you inherit the version coupling between the ai-dynamo wheel, ai-dynamo-runtime and whichever engine backend you selected.

A narrower alternative exists inside the repository itself. The lib/mocker crates and the examples/router custom policy crates let you exercise routing policies without a full GPU deployment, which is a cheaper way to evaluate the router before committing cluster capacity.

## Maintenance cadence, release naming and what the licence label does not say

The repository is not archived, and the last push was on 2026-09-20. Recent releases carry unusual names: v1.4.1-k-exaone-2.0-750b-post.1 on 2026-09-17, v1.4.1-solar-open2-250b-post.1 on 2026-09-16 and v1.4.1-a.x-k2-post.1 on 2026-09-15. These look like per-model or per-partner build tags rather than a single linear release line, while pyproject.toml and Cargo.toml both declare 1.6.0. If you pin versions, pin the wheel version and the runtime version together, since the dependency is written as ai-dynamo-runtime==1.6.0.

Upgrade cost is dominated by the backend extras. The trtllm extra pins tensorrt-llm==1.3.0rc26, a release candidate, so that path moves with TensorRT-LLM's own cadence. The base install is light by comparison.

On licensing: the README header and pyproject.toml both state Apache-2.0, and the package metadata declares the Apache Software License classifier. GitHub reports the repository licence as NOASSERTION, which usually means the detector could not match the LICENSE file to a known template. The README does not document rollback or a downgrade path between releases, so treat the LICENSE file itself, not the classifier, as the thing to read before you ship.

## Conclusion

Adopt Dynamo if you already run vLLM, SGLang or TensorRT-LLM across several GPUs or nodes and your problem is prefill/decode imbalance, redundant prefill work or replica cold-start time. Do not adopt it for a single model on a single GPU: the README says your inference engine alone is probably sufficient, and nothing in the repository contradicts that. Before committing, verify three things: that your engine appears in the backend support matrix for the specific feature you need (KVBM is still marked as work in progress for SGLang), that the Python wheel installs against your interpreter, and that the LICENSE file resolves the NOASSERTION label GitHub reports for this repository.

## FAQ

### What is NVIDIA Dynamo used for?

It is an orchestration layer above inference engines such as SGLang, TensorRT-LLM and vLLM, turning a cluster of GPUs into a coordinated multi-node inference system. The README lists disaggregated prefill/decode, KV-aware routing, multi-tier KV caching and automatic scaling as its core capabilities.

### How do I install Dynamo?

The Python package is published as ai-dynamo and requires Python 3.10 or newer. Backend support comes through extras, for example the trtllm extra, and the README also points to NGC containers and the recipes and examples directories for runnable deployments.

### When should I not use Dynamo?

The README states that if you are running a single model on a single GPU, your inference engine alone is probably sufficient. Dynamo targets workloads spread across multiple GPUs or nodes where prefill and decode need independent scaling.

### Does Dynamo support SGLang, TensorRT-LLM and vLLM equally?

Not entirely. The README support table lists disaggregated serving, KV-aware routing and the SLA-based planner for all three, but marks KVBM as available for TensorRT-LLM and vLLM and still under construction for SGLang.

## Sources

- [ai-dynamo/dynamo on GitHub](https://github.com/ai-dynamo/dynamo)
- [Issues](https://github.com/ai-dynamo/dynamo/issues)
- [Project website](https://docs.nvidia.com/dynamo/latest)
- [README](https://github.com/ai-dynamo/dynamo/blob/main/README.md)
- [Releases](https://github.com/ai-dynamo/dynamo/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ai-dynamo-dynamo
