# AWS Deep Learning Containers: nine tracks, two tag schemes

> This repository builds and releases pre-built AI/ML images rather than hosting them, so consuming it means pulling from ECR. Nine autorelease workflows cover PyTorch, TensorFlow training and inference, vLLM, vLLM-Omni, SGLang, Ray and two CUDA base tracks, and the release notes are where you find out what actually changed inside an image.

**aws/deep-learning-containers** —  One stop shop for running AI/ML on AWS.

- Repository: https://github.com/aws/deep-learning-containers
- Website: https://aws.github.io/deep-learning-containers/
- Stars: 1,190 · Forks: 558
- Language: Python
- License: NOASSERTION
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/aws-deep-learning-containers

## Pre-built images for AI/ML, and where the images actually live

The first thing to get straight is that this repository is not the product. AWS Deep Learning Containers are pre-built Docker images for AI/ML workloads on AWS, and they live in a container registry. The repository builds them, tests and patches them, and documents them. If you are looking for something to add to your requirements, there is nothing here; you consume this project by pulling a tag.

The stated value proposition is narrow and specific: each image is tested and patched for security vulnerabilities. That is the whole argument for a pre-built stack over assembling your own, and it is also the part that transfers to you, because a patched image is only patched at the moment it was built.

Three links carry the documentation, and they are worth bookmarking separately. The documentation site is at aws.github.io/deep-learning-containers, the Available Images reference is the page that tells you which tags exist, and the tutorials page has worked examples. Individual tracks also have their own gallery pages under the deep-learning-containers path in ECR, one for vLLM, one for SGLang and one for Ray.

So the workflow is unglamorous: find the track, find the tag, pull it, and treat the release notes as a changelog for a dependency you did not write.

## Nine autorelease workflows, one per track

The build matrix is visible in the badges at the top of the README, one workflow file per track, and the version is in the filename. That is the mechanism: a track is a workflow, and bumping a framework means editing that workflow and letting it run.

The PyTorch track is `pytorch.autorelease-2.13-ec2.yml`, so PyTorch 2.13 images for EC2. TensorFlow splits by role, with `tensorflow-training.autorelease-2.21-sagemaker.yml` and `tensorflow-inference.autorelease-2.20-sagemaker.yml`, which is a useful distinction to remember because the two are versioned independently and neither tracks PyTorch's cadence. Then `vllm.autorelease-ec2-amzn2023.yml` and `vllm-omni.autorelease-ec2.yml` for the two vLLM lines, `sglang.autorelease-ec2-amzn2023.yml`, and `ray.autorelease-ec2.yml`.

The last two are the base tracks, `base.autorelease-cu130.yml` and `base.autorelease-cu132.yml`, which is where the catalogue's CUDA floor shows up: 13.0 and 13.2 are the two generations being maintained, and every GPU track sits on one of them. That is a genuine constraint for anyone still on an older CUDA.

What this arrangement buys is that a track can move without dragging the others. The SGLang highlights from 2026/09/07 and the vLLM highlights from 2026/09/11 landed in different workflows on different days, and neither had to wait for the other. The cost is that you now have nine version streams to track instead of one dependency, which is the trade the release notes exist to cover.

## Two tag schemes, so a tag cannot be inferred

Here is the friction that will cost you an afternoon if nobody warns you. There are two naming conventions in this catalogue and they do not match.

The first is the release tag on the repository itself, which encodes framework, platform, framework version, CPU or GPU, and Python: `v2.9-pt-ec2-2.10.0-tr-gpu-py313` and `v2.5-pt-ec2-2.9.0-tr-py312`. That reads as `pt` for PyTorch, `ec2` for the platform, the framework version, `tr` for training, `gpu` or nothing for accelerator, and the Python version.

The second is the image tag you actually pull, and the serving tracks ignore all of that:

```text
0.29.0-gpu-py312-ec2
0.29.0-gpu-py312
0.5.19-gpu-py312-ec2
serve-llm-cuda-v1.0
train-ml-cuda-v1.1
omni-cuda-v1.6
server-cuda-v2.4
```

The rule that does hold across the serving tracks is the platform suffix. `-ec2` at the end means the EC2 or EKS flavour, and the same tag without it is the SageMaker flavour, which is how vLLM 0.29.0 exists as both `0.29.0-gpu-py312-ec2` and `0.29.0-gpu-py312`. Ray uses the same idea with `serve-llm-cuda-v1.0` and `train-ml-cuda-v1.1` for EC2 and EKS.

The base operating system also varies by track rather than being uniform. The Ray Train and Ray LLM images and the vLLM Server and vLLM-Omni lines are on AL2023, the vLLM and SGLang runtime images are on Ubuntu, and vLLM 0.28.0's release note records that the runtime base moved to Ubuntu 24.04. So the only reliable procedure is the Available Images reference. Do not construct a tag from a framework version and hope.

## Picking an image: read the reference, then check the base

A first use goes like this. Go to the Available Images reference, filter by framework and by whether you need a GPU, and copy the tag rather than composing it. Then confirm four things before you build on it, all of which are visible in the tag or the release note.

Python version, because a model repository written for 3.12 will not necessarily run on a 3.13 image, and the PyTorch tags make that explicit. Accelerator, because a training image and a CPU-only image are different tags and the training tracks carry a `gpu` marker while the CPU ones do not. Base OS, because Ubuntu and AL2023 differ in package availability, and the vLLM line has already moved its runtime base once inside a version. And CUDA generation, because the base workflows say which CUDA the track sits on and 13.0 and 13.2 are both in flight.

Beyond the tags, the repository carries the Dockerfiles under `docker/`, the build and release scripts under `scripts/`, and two worked example directories, `examples/ray/` and `examples/vllm-omni/`. If you are standing up Ray training or a vLLM-Omni deployment, those are the closest thing to a starting template, and reading them tells you which environment variables the images expect.

One more check worth making early: whether the track you need has a SageMaker variant at all. TensorFlow training and inference are SageMaker-only in this list, while PyTorch and the serving tracks are on EC2, so a plan that assumes one tag works everywhere is wrong from the start.

## What the images add, and the default that changed under you

The release highlights are the useful part of this README, because they say what is inside rather than what the tag is called. Two entries show the value and the risk at the same time.

The vLLM v0.28.0 entry from 2026/08/26 is the clearest example. It lists stack-wide Kimi-K3 work, including Decode Context Parallel, fused FlashKDA kernels and GEMM-RS sequence parallelism, end-to-end DeepSeek V4 sparse MLA with MTP and DSpark speculative decoding, tiered KV cache offloading to disk, and new models. It also states that the runtime base moves to Ubuntu 24.04, that Transformers goes to 5.15.0, and that the `max_num_batched_tokens` default changes from 8192 to 16384. That last change doubles the default batched token budget in a minor release of an image whose version number did not move. It is a throughput and memory behaviour change delivered inside a patch.

The Ray Train v1.0 entry from the same day is the other direction. It ships a stack rather than a feature: Ray Train, Tune and Data on PyTorch 2.13.0 with CUDA 13.0.2 and Python 3.13, plus EFA 1.47.0, the AWS NCCL OFI plugin, GDRCopy, flash-attn, Transformer Engine and DeepSpeed all pre-installed. One image runs as a Ray head or a worker under KubeRay on any EKS cluster, including SageMaker HyperPod-EKS, or standalone on EC2. Ray Train v1.1 two days later moves EFA to 1.49.0 and nothing else.

The small entries are still worth reading. vLLM Server v2.4 and vLLM-Omni v1.6 both carry the same fix, where SageMaker `SM_VLLM_*` JSON-array environment variables expand into multiple argv values so multi-value flags parse. SGLang 0.5.19 adds beam search and DeepEP v2 MoE all-to-all, and 0.5.18 adds diffusion models plus an overlapped checkpoint staging mode for faster startup.

## Bare engine, engine behind Ray Serve, or a different engine

Inside this catalogue there is a genuine architectural choice, and it is not visible from the tag alone.

On the plain vLLM images you get the engine and nothing else. You run it, it serves, and anything in front of it is your code. That is the right shape when you have one model on one GPU and want the shortest path between a checkpoint and an endpoint.

The Ray LLM image adds a layer on top: the 2026/09/01 release note says `ray[llm]`'s `build_openai_app` runs vLLM behind Ray Serve, giving OpenAI-compatible serving, single GPU on EC2 or multi-node on EKS via KubeRay. The difference is that request routing, scaling and deployment become Ray's problem instead of yours, and the cost is a second framework in the path plus a Ray version to track alongside the engine.

SGLang is the third option and not a Ray variant of anything. It carries its own model list, its own release train, and features vLLM's images do not list in the same notes, such as beam search in 0.5.19 and the DeepEP v2 MoE all-to-all. The engine choice determines which model and kernel work lands in your image first, so it is worth reading two release trains rather than picking on tag freshness.

Whichever you take, the failure mode is the same and it is not subtle: an image that does not contain the framework version your model code needs. That is a resolution problem, not a tuning problem, and no amount of flags fixes it.

## The repository itself is ruff, pre-commit and mkdocs

The Python in this project is tooling, not runtime, and the files make that unambiguous. pyproject.toml carries a Ruff configuration and nothing else:

```toml
[tool.ruff]
line-length = 100
target-version = "py312"

[tool.ruff.format]
quote-style = "double"
indent-style = "space"
line-ending = "lf"
```

requirements.txt contains exactly one package, `pre-commit`, and there is a `.pre-commit-config.yaml` to go with it. A `.markdownlint.json` and a `_typos.toml` extend the same idea to the documentation, which is most of what this repository contains. The docs site is MkDocs, driven by `mkdocs.yaml` and the `docs/` tree, and `DEVELOPMENT.md` and `CONTRIBUTING.md` cover working on the build.

Two things follow for anyone filing a problem. First, a bug in your model code is not a bug here; the images carry upstream frameworks and the release notes are where framework fixes are recorded. Second, if you are reading this repository to work out what is inside an image, you will not find out from the tree. You need the release notes and the Available Images reference.

There is also a `.claude/` directory at the top level and a `test/` directory, so the project has an agent configuration and a test suite for its own build, both of which are about the pipeline rather than the artefacts.

## Weekly image churn, a custom licence, and when to skip it

The maintenance picture is fast and consistent. The last push was on 2026-09-20, and three tagged releases landed inside two weeks: v2.9-pt-ec2-2.10.0-tr-gpu-py313 on 2026-09-16, v2.5-pt-ec2-2.9.0-tr-py312 on 2026-09-15, and v2.8-pt-ec2-2.10.0-tr-py312 on 2026-09-09. The release highlights run at a similar cadence, with nine entries between 2026/08/25 and 2026/09/11.

That cadence is a feature for training and inference images, where a framework fix matters, and a tax for anything with a reproducibility requirement. Pinning a tag gives you a fixed artefact, but the interesting content moves weekly, so the version you validated last month is three framework releases behind.

The licence needs care and this review will not interpret it. The project page shows a custom licence that GitHub does not classify as a standard one, and the tree carries both a LICENSE file and a NOTICE file. The terms are in the file. For an internal evaluation, reading it is a formality. For anything you ship, it is the first thing to check, and it is not answered anywhere in the README.

Three cases where you would be better off building your own image. If you need a framework version no track has published, you are pinning to nothing. If your base OS is neither Ubuntu nor AL2023, the installed system libraries will not be what you tested. And if you cannot pull from ECR, the patched-and-tested argument does not apply to you either, because you would be rebuilding the stack from a Dockerfile in `docker/` and owning every layer from that point on.

## Conclusion

Use AWS Deep Learning Containers when you want an AI/ML stack that has already been tested and patched, and you are willing to accept a base image and a set of defaults you did not choose. Read the release notes before pinning a tag, because vLLM 0.28.0 moved the runtime base to Ubuntu 24.04 and raised the `max_num_batched_tokens` default from 8192 to 16384 inside the same image version. Read the LICENSE file in the tree before you build a product on these images, since the project uses a custom licence that GitHub does not classify as a standard open-source one.

## FAQ

### Where are the AWS Deep Learning Container images published?

To AWS Elastic Container Registry, with a gallery page per track under the deep-learning-containers path, including separate pages for vllm, sglang and ray. The Available Images reference in the documentation site is the authoritative list of tags, and the repository itself only builds and publishes them.

### What is the difference between an EC2 tag and a SageMaker tag?

The same image built for a different platform. vLLM 0.29.0 exists as `0.29.0-gpu-py312-ec2` for EC2 and `0.29.0-gpu-py312` for SageMaker, and the Ray tracks use `serve-llm-cuda-v1.0` and `train-ml-cuda-v1.1` for EC2 and EKS. Note that TensorFlow training and inference tracks here are SageMaker-only.

### Which frameworks does AWS Deep Learning Containers publish?

PyTorch through a 2.13 EC2 auto-release, TensorFlow Training 2.21 and TensorFlow Inference 2.20 for SageMaker, vLLM, vLLM-Omni and vLLM Server, SGLang, Ray LLM and Ray Train, plus base images for CUDA 13.0 and 13.2. Each track has its own workflow, so the version streams move independently.

### How do I run distributed Ray training from an AWS Deep Learning image?

Use the `train-ml-cuda-v1.1` image on EC2 or EKS, which ships Ray Train, Tune and Data on PyTorch 2.13.0 with CUDA 13.0.2 and Python 3.13, with EFA 1.49.0, the AWS NCCL OFI plugin, GDRCopy, flash-attn, Transformer Engine and DeepSpeed pre-installed. One image runs as a Ray head or worker under KubeRay on any EKS cluster, including SageMaker HyperPod-EKS, or standalone on EC2.

### Can an AWS Deep Learning Container serve LLMs through an OpenAI-compatible API?

Yes. Ray LLM v1.0 ships as `serve-llm-cuda-v1.0`, where `ray[llm]`'s `build_openai_app` runs vLLM behind Ray Serve on PyTorch 2.11.0 with CUDA 13.0.2 and Python 3.13. The same release supports a single GPU on EC2 and multi-node serving on EKS through KubeRay.

### Are the AWS Deep Learning Container images patched for security vulnerabilities?

The project states that each image is tested and patched for security vulnerabilities. The release notes record what changed, including fixes such as the vLLM Server v2.4 and vLLM-Omni v1.6 change where SM_VLLM_* JSON-array environment variables expand into multiple argv values so multi-value flags parse correctly.

## Sources

- [aws/deep-learning-containers on GitHub](https://github.com/aws/deep-learning-containers)
- [Issues](https://github.com/aws/deep-learning-containers/issues)
- [Project website](https://aws.github.io/deep-learning-containers/)
- [README](https://github.com/aws/deep-learning-containers/blob/main/README.md)
- [Releases](https://github.com/aws/deep-learning-containers/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/aws-deep-learning-containers
