NVIDIA Merlin: GPU-Accelerated Recommender Systems, Component by Component
NVIDIA Merlin is an open source library providing end-to-end GPU-accelerated recommender systems, from feature engineering and preprocessing to training deep learning models and running inference in production.
At a glance
- What is it?
- NVIDIA Merlin is an umbrella repository that ties together six open source libraries for building recommender systems on NVIDIA GPUs. The README documents the architecture and the benefits; the installation section is truncated, so the entry point below is the repository's own requirements.txt and examples.
- Who is it for?
- Adopt NVIDIA Merlin if your recommender pipeline already runs on NVIDIA GPUs and you need to move terabyte-scale feature engineering and large embedding tables through training and into Triton serving without leaving the CUDA stack. Do not adopt it if you are CPU-only, if you need a single pip-installable package rather than six coordinated ones, or if you want a managed service.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 71 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What NVIDIA Merlin solves, and who the README addresses
The README names its audience directly: data scientists, machine learning engineers, and researchers. The problem it targets is scale. Feature engineering, training, and inference for recommender systems each have different bottlenecks, and the README claims each stage of the pipeline is optimized to support hundreds of terabytes of data. That is a statement of intent rather than a measured result, and no benchmark appears in the repository description.
The more concrete claim is about memory. Deep learning recommender models are dominated by embedding tables, and the README says Merlin can distribute large embedding tables that exceed available GPU and CPU memory. If your model fits comfortably in a single GPU, this is not the problem you have. If your embedding table does not fit, the whole design starts to make sense.
A secondary goal is that you do not have to abandon your existing framework. The README describes accelerating existing TensorFlow, PyTorch, or FastAI training pipelines with custom data loaders, rather than requiring you to rewrite training in a Merlin-specific framework. That framing matters for adoption cost, and it is also where the component boundaries get complicated.
Six libraries, one pipeline: how the Merlin components fit together
Merlin is not a single library. The README enumerates six open source components, and the split is by pipeline stage rather than by convenience.
NVTabular handles feature engineering and preprocessing for tabular data. It offers a high-level API for defining transformation workflows and is designed to process datasets that exceed GPU and CPU memory. HugeCTR is the training framework, and it is the component responsible for distributing training across multiple GPUs and nodes and for scaling embedding tables beyond available memory. The README describes loading a subset of an embedding table into a GPU in a coarse-grained, on-demand manner during training.
Merlin Models provides standard model implementations, from classic machine learning to deep learning architectures, and its stated role is to map datasets created with NVTabular into a model input layer automatically. That mapping is the integration point that lets you change features or change models without breaking the other side. Transformers4Rec covers sequential and session-based recommendation, with modular building blocks compatible with standard PyTorch modules, supporting next-item prediction as well as binary classification and regression.
Merlin Systems is the serving layer. It combines models with feature stores, nearest neighbor search, and exploration strategies into graphs served with Triton Inference Server. Merlin Core sits underneath all of them, providing a shared dataset abstraction and a common schema that identifies key dataset features.
The dependency direction is worth noting. Merlin Core is described as functionality used throughout the ecosystem, so upgrading it affects every other component. That is the structural reason version alignment matters here more than in a typical single-package project.
Installing NVIDIA Merlin and running the first example
The README's installation section is truncated in the repository description, so the reliable starting point is the repository's own requirements.txt, which lists the components and their minimum versions. Note that the top-level setup.py declares an empty packages list and an empty install_requires, so installing the repository itself does not pull in the stack.
pip install -r requirements.txtThat command installs merlin-core, merlin-dataloader, nvtabular, merlin-models, merlin-systems and transformers4rec at version 23.4.0 or newer, plus feast pinned at exactly 0.31. The exact pin on feast is the one line that can conflict with an existing environment, since it does not float.
The repository ships worked examples rather than a single quickstart script. The examples directory contains getting-started-movielens, quick_start, ranking, scaling-criteo, traditional-ml, Next-Item-Prediction-with-Transformers, Building-and-deploying-multi-stage-RecSys, and sagemaker-tensorflow. The README does not document the commands to run inside those directories, so read each example's own README before assuming an entry point.
ls examples/For a first real use, getting-started-movielens is the smallest end-to-end path in the repository layout, and scaling-criteo is the one that exercises the terabyte-scale preprocessing claims. Start with the former to confirm your environment resolves, then move to the latter if data scale is your actual problem.
Where NVIDIA Merlin stops being the right tool
The strongest limitation is stated by the project itself: this is GPU acceleration on NVIDIA GPUs. Nothing in the README describes a CPU path for the training and embedding-table scaling that define the project. If your infrastructure is CPU-only, or if you run on accelerators from another vendor, the core value proposition does not apply.
The second limitation is the packaging model. Because the functionality is split across six libraries with a shared core, you inherit six version constraints. The requirements.txt floors are all 23.4.0, which suggests the components are expected to move together, but the repository also carries a CHANGELOG.md and a docker directory, implying container images are part of how the project is expected to be consumed. The README does not document an upgrade or rollback procedure for the component set.
The third limitation is scope. Merlin Models aims to provide standard models, and Transformers4Rec covers sequential and session-based recommendation. Neither is presented as a general-purpose framework for arbitrary architectures. If your recommender is a custom architecture that does not resemble the building blocks described, you are using the data loaders and preprocessing rather than the modeling layer, and the value you get is correspondingly narrower.
Finally, the top-level setup.py classifies the project as Development Status 4 - Beta, with version 0.0.1. That classifier describes the umbrella repository rather than the individual components, but it is the only status signal the repository itself provides.
Merlin versus a CPU feature store plus a general training framework
The obvious alternative approach is to assemble the pipeline yourself: a feature store for preprocessing and serving features, a general deep learning framework for training, and a separate model server for inference. The difference is where the work happens. In that assembly, feature transformation and embedding lookup are typically CPU-bound or split awkwardly across the boundary, and the scaling strategy for embedding tables is something you implement.
Merlin's approach is to push each stage onto the GPU and keep the data there. NVTabular does preprocessing on GPU, the data loaders feed TensorFlow, PyTorch, or HugeCTR without a CPU round trip, HugeCTR distributes embedding tables across GPUs and nodes, and Merlin Systems expresses the serving graph for Triton. The common schema in Merlin Core is what makes the handoffs mechanical rather than bespoke.
The trade-off is real. You give up framework neutrality at the hardware level, and you take on six dependencies that are versioned together. You also get a serving story that assumes Triton Inference Server, which is a specific choice. If your team already runs a different serving stack, Merlin Systems is the component you would skip, and skipping it removes the part of the pipeline the README describes as deployable in a few lines of code.
Maintenance, licensing, and what upgrading actually costs
The repository is not archived, and the last push was on 2026-07-22. The most recent tagged release listed is v24.06.00 from 2024-06-14, with v23.12.00 and v23.09.00 before it. There is a visible gap between the release cadence and the commit activity, which is worth understanding before you plan an upgrade: commit activity on the umbrella repository does not necessarily correspond to new component releases, and the components live in separate repositories.
The upgrade cost is driven by the shared core. Because Merlin Core provides the dataset abstraction and schema used across the ecosystem, and because the requirements.txt floors move in lockstep at 23.4.0, a partial upgrade is the risky path. The repository provides a docker directory and a CHANGELOG.md, but the README does not document a supported upgrade sequence or a rollback procedure, so treat the changelog as the source of truth for what changed.
On licensing, the repository is Apache-2.0, and the setup.py declares the same. Apache-2.0 is a permissive licence with an explicit patent grant and requires preservation of notices. The components are separate repositories with their own licence files, and the README does not state that all of them share a single licence. Check each component's LICENSE before redistribution. This is a description of what the repository states, not legal advice.
Editorial conclusion
Adopt NVIDIA Merlin if your recommender pipeline already runs on NVIDIA GPUs and you need to move terabyte-scale feature engineering and large embedding tables through training and into Triton serving without leaving the CUDA stack. Do not adopt it if you are CPU-only, if you need a single pip-installable package rather than six coordinated ones, or if you want a managed service. Before committing, verify which component versions satisfy the requirements.txt floor of 23.4.0 for merlin-core, merlin-dataloader, nvtabular, merlin-models, merlin-systems and transformers4rec, and confirm the pinned feast==0.31 is compatible with whatever feature store you already run, because that pin is the one dependency in the file with an exact version rather than a minimum.
Frequently asked questions
What is NVIDIA Merlin, and is it one library or several?
It is an umbrella repository for an end-to-end GPU-accelerated recommender system stack. The README enumerates six open source libraries: NVTabular, HugeCTR, Merlin Models, Transformers4Rec, Merlin Systems and Merlin Core.
How do I install NVIDIA Merlin?
The README's installation section is truncated in the repository description, but the repository ships a requirements.txt listing each component at version 23.4.0 or newer plus feast==0.31, so pip install -r requirements.txt is the documented dependency set. The top-level setup.py declares an empty install_requires, so installing the repository itself pulls in nothing.
Does NVIDIA Merlin work without an NVIDIA GPU?
The README describes the project as GPU-accelerated on NVIDIA GPUs and does not document a CPU path for training or for scaling embedding tables beyond GPU and CPU memory. The CPU is mentioned as a memory tier that embedding tables can exceed, not as an execution target.
What is Merlin Systems used for?
It combines recommendation models with other production components such as feature stores, nearest neighbor search and exploration strategies into graphs that can be served with Triton Inference Server. It is the serving layer of the stack.
Which example should I start with in the NVIDIA Merlin repository?
The examples directory includes getting-started-movielens, quick_start, ranking, scaling-criteo, traditional-ml and others. The README does not document commands for running them, so each example directory must be read on its own.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvidia-merlin-merlin)