HugeCTR: NVIDIA's GPU framework for CTR and large-embedding recommender training
HugeCTR is a high efficiency GPU framework designed for Click-Through-Rate (CTR) estimating training
At a glance
- What is it?
- HugeCTR is a C++ recommender framework from NVIDIA Merlin for training and inference of models with very large embedding tables. Its Python interface is the practical entry point, and since version 25.03 you build the Docker image yourself.
- Who is it for?
- Adopt HugeCTR if you train CTR or ranking models with embedding tables too large for a single GPU and you already have NVIDIA GPUs and a Docker workflow. Do not adopt it if you need a CPU-only path, a small tabular model, or a framework whose container image is published for you, because since version 25.03 the README states that only the Dockerfile source is provided and users must build the image themselves.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 59 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem HugeCTR targets: embedding tables that do not fit
A click-through-rate model is mostly an embedding lookup. The dense part of a DCN or DeepFM network is small; the parameters live in the slot tables, and those tables grow with the cardinality of user IDs, item IDs and categorical features. On a single GPU the table eventually exceeds memory, and the usual workaround is to shard the model by hand or to shrink the vocabulary, which changes the model you can train. HugeCTR is built around that constraint. The README lists model parallel training, multi-node training, mixed precision training and an optimized GPU workflow as core features, and the design goal it states is to let you deploy recommender models with very large embedding. The audience is narrow and specific: engineers training CTR or ranking models on NVIDIA hardware, not people looking for a general-purpose deep learning framework. If your model fits comfortably on one card, the machinery here is overhead you do not need.
How a HugeCTR training job is assembled
The Python interface is a builder, not a declarative model definition. You create a solver, a data reader and an optimizer, pass them to hugectr.Model, then add layers one at a time with model.add(...). The README's DCN example shows exactly this shape: CreateSolver carries max_eval_batches, batchsize_eval, batchsize, lr, vvgpu and repeat_dataset; CreateOptimizer takes an optimizer_type such as hugectr.Optimizer_t.Adam and an update_type such as hugectr.Update_t.Global; DataReaderParams takes the reader type, the training and evaluation source lists and a slot_size_array. The slot_size_array is the interesting part. It is a per-slot cardinality list, and it is what allows the framework to allocate and shard embedding tables across the GPUs named in vvgpu. Because the reader, the solver and the layer list are separate objects, a HugeCTR script reads more like a configuration assembled in code than like a PyTorch module. That is a deliberate trade: it gives the runtime enough static information to plan memory and communication before training starts, and it costs you the flexibility of defining layers dynamically at runtime.
Building the image and running a first DCN job
The README's getting-started path assumes Docker and at least one visible NVIDIA GPU. From version 25.03, HugeCTR only provides the Dockerfile source, and the README says users need to build the image themselves from tools/dockerfiles/Dockerfile.base. Run this from the repository root:
docker build --build-arg RELEASE=true -t hugectr:release -f tools/dockerfiles/Dockerfile.base .Then start a container with your host directory mounted. The README notes that HugeCTR uses NCCL to share data between ranks and that NCCL may require shared memory for IPC and pinned system memory, so it recommends raising those limits:
docker run --gpus=all --rm -it --cap-add SYS_NICE -shm-size=1g -ulimit memlock=-1 -v /your/host/dir:/your/container/dir -w /your/container/dir -u $(id -u):$(id -g) hugectr:releaseInside the container, generate a synthetic Parquet dataset before touching real data. The README's dcn_parquet_generate.py builds DataGeneratorParams with label_dim 1, dense_dim 13, num_slot 26 and a 26-entry slot_size_array, then calls DataGenerator(...).generate(). Running python dcn_parquet_generate.py writes training and evaluation files into ./dcn_parquet. The training script then points DataReaderParams at ./dcn_parquet/file_list.txt and ./dcn_parquet/file_list_test.txt and reuses the same slot_size_array, which must match the generator. If those two lists disagree, the reader is working from cardinalities the data does not have.
Where HugeCTR stops being the right tool
The framework is tied to NVIDIA GPUs and to a container workflow. There is no CPU training path described in the documentation, and the getting-started instructions begin with docker build and --gpus=all. Two consequences follow. First, the build step is now yours: the README states that from version 25.03 only the Dockerfile source ships, so an environment that cannot build a large C++ image, or that pins an older base image, is a real blocker rather than an inconvenience. Second, the API surface is domain-specific. HugeCTR gives you the essentials for recommender models with very large embeddings, and the samples directory reflects that with bst, criteo, dcn, deepfm, din, dlrm, ftrl, mmoe, ncf and wdl. A sequence model for text, a vision backbone or a small gradient-boosted problem has nothing to gain here. The slot_size_array requirement is another boundary: your categorical features have to be enumerable and their cardinalities known up front, which rules out pipelines where the vocabulary is discovered online.
HugeCTR compared with the Sparse Operation Kit in the same repository
The most direct alternative is not in another project; it is in this one. The repository contains sparse_operation_kit alongside HugeCTR, with its own documentation site, and the README lists it as a core feature. The difference in approach is where the embedding lives. HugeCTR is a standalone framework: you build the model with hugectr.Model and its layer calls, and the embedding tables are managed by the framework's own planner, which is what makes model parallel and multi-node training possible. Sparse Operation Kit takes the opposite route, plugging sparse embedding operations into an existing framework so the rest of the model stays where it is. If your team already has a TensorFlow or PyTorch training stack and only the embedding layer is the bottleneck, SOK is the smaller change. If you are willing to write the model in HugeCTR's builder API, you get the reader, solver and optimizer integration in one place. Choosing between them is mostly a question of how much of the existing training code you are prepared to rewrite.
Maintenance, releases and the Apache-2.0 licence
The repository is not archived, and the last push was on 2026-08-03, so work is ongoing. Releases are not frequent: v26.03.00 landed on 2026-03-12, v25.03.00 on 2025-03-14, and v24.06.00 on 2024-06-14. That cadence matters for planning. A pinned release may sit for a year, so the upgrade cost is concentrated in occasional large jumps rather than continuous small ones, and the 25.03 change to Dockerfile-only distribution is exactly the kind of shift that lands in one of those jumps. Check release_notes.md before moving between versions. HugeCTR is licensed under Apache-2.0, which permits commercial use and modification and requires that you preserve the licence and attribution notices; the repository also vendors third-party code under third_party/, and those components may carry their own terms. Read the LICENSE file and the third-party notices rather than assuming a single licence covers the whole build. This is a description of the licence text, not legal advice.
Editorial conclusion
Adopt HugeCTR if you train CTR or ranking models with embedding tables too large for a single GPU and you already have NVIDIA GPUs and a Docker workflow. Do not adopt it if you need a CPU-only path, a small tabular model, or a framework whose container image is published for you, because since version 25.03 the README states that only the Dockerfile source is provided and users must build the image themselves. Before committing, verify that your GPU generation is supported by the release you intend to pin, that the build of tools/dockerfiles/Dockerfile.base completes in your environment, and that your data can be expressed as the Parquet or other reader formats the DataReaderParams accepts.
Frequently asked questions
What is NVIDIA Merlin?
Merlin is the NVIDIA umbrella the HugeCTR repository belongs to, and its release tags carry the Merlin prefix, such as v25.03.00 labelled Merlin: HugeCTR 25.03. HugeCTR itself is the GPU-accelerated recommender framework for training and inference of large deep learning models, and Sparse Operation Kit is distributed from the same repository.
How do I install HugeCTR?
The README directs you to build the Docker image yourself from tools/dockerfiles/Dockerfile.base, because from version 25.03 only the Dockerfile source is provided. You then run a container with --gpus=all and the shared memory and memlock options the README recommends for NCCL.
Does HugeCTR require an NVIDIA GPU?
Yes. The getting-started instructions build a CUDA-based image and start the container with --gpus=all, and the solver's vvgpu parameter names the GPUs the job uses. The documentation describes no CPU training path.
What is the slot_size_array in a HugeCTR data reader?
It is a per-slot cardinality list passed to DataReaderParams, and it must match the slot_size_array used when the dataset was generated. In the README's DCN example both the generator and the trainer carry the same 26-entry list.
Can HugeCTR export models to another runtime?
The README lists a HugeCTR to ONNX Converter among the core features, which is the path it documents for moving a trained model out of the framework. Details of the conversion are in the linked core features documentation rather than the README itself.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvidia-merlin-hugectr)