TorchRec: sharding embedding tables across many GPUs
Pytorch domain library for recommendation systems. TorchRec TorchRec** is a PyTorch domain library built to provide common sparsity and parallelism primitives needed for large-scale recommender systems (RecSys).
At a glance
- What is it?
- TorchRec is a PyTorch domain library for large-scale recommender systems. It supplies the sharding, planning and pipelining primitives that a plain nn.EmbeddingBag does not, and it expects a distributed GPU environment to be useful.
- Who is it for?
- Adopt TorchRec if you are training recommender models whose embedding tables no longer fit on one device and you already run distributed PyTorch. Do not adopt it for a small ranking model that fits in a single GPU's memory, or if you cannot install a matching fbgemm-gpu build, since the optimized kernels come from FBGEMM.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem: embedding tables that outgrow one device
A recommender model is mostly embedding table. The dense part that turns features into a prediction is small; the vocabulary side is where the parameters live, and it grows with users and items rather than with model depth. A single nn.EmbeddingBag is a single tensor, and a single tensor has to sit on a single device. Once the table passes the memory of one GPU, the usual workarounds appear: shrink the embedding dimension, hash the vocabulary, or keep the table on CPU and pay the transfer cost on every step.
TorchRec exists to remove that constraint. The README describes it as a PyTorch domain library providing "common sparsity and parallelism primitives needed for large-scale recommender systems (RecSys)", and states that it powers many production RecSys models at Meta. The intended reader is an engineer who already writes PyTorch models and now needs the embedding side distributed across devices and nodes without hand-rolling the collective communication. If your model fits comfortably on one accelerator, this library is overhead rather than help.
How sharding, planning and pipelining fit together
TorchRec splits the work into three cooperating pieces. Sharders decide how a table is divided. The README lists data-parallel, table-wise, row-wise, table-wise-row-wise, column-wise and table-wise-column-wise sharding. The distinction matters: row-wise sharding cuts the vocabulary so each rank owns a slice of rows, which is what you want when one table is enormous, while table-wise sharding assigns whole tables to ranks, which suits a model with many medium tables and an uneven access pattern.
The planner generates a sharding plan automatically, choosing an assignment rather than making you write one. The README describes it as a planner "that can automatically generate optimized sharding plans for models". In practice this is where most of the tuning happens, because the plan determines how much cross-device traffic each step produces.
The third piece is the pipelined training loop, which overlaps dataloading device transfer (copy to GPU), inter-device communication (input_dist) and computation (forward, backward). Those three stages are named explicitly in the README, and the overlap is the reason the communication cost of sharding does not simply serialize on top of compute. Underneath, the optimized kernels come from FBGEMM, which is why the installation path pulls in fbgemm-gpu before TorchRec itself.
Installing TorchRec and running the installation test
The README points to the Getting Started page in the documentation for recommended setup and states plainly that "Generally, there isn't a need to build from source". The source path is documented for people who want the latest changes. It begins with PyTorch, installed from a CUDA-specific index. This is the CUDA 12.6 case:
pip install torch --index-url https://download.pytorch.org/whl/nightly/cu126FBGEMM comes next, from the matching index, because the optimized RecSys kernels live there:
pip install fbgemm-gpu --index-url https://download.pytorch.org/whl/nightly/cu126Then clone the repository recursively, install the listed requirements, and install TorchRec itself:
git clone --recursive https://github.com/meta-pytorch/torchrec
cd torchrec
pip install -r requirements.txt
python setup.py install developThe recursive clone matters because the build expects submodules. The README then gives a verification step that launches a two-process job through TorchX:
torchx run -s local_cwd dist.ddp -j 1x2 --gpu 2 --script test_installation.pyOn a machine without GPUs, the same script takes a CPU flag:
torchx run -s local_cwd dist.ddp -j 1x2 --script test_installation.py -- --cpu_onlyA successful run means the sharded embedding path initializes and communicates across the two ranks. If it fails, the mismatch is almost always between the torch build, the fbgemm-gpu build and the CUDA version, not in TorchRec's own code. For a fuller example, the README points at the DLRM example in the facebookresearch/dlrm repository.
What the repository does not tell you
The README is a feature list with an installation appendix, and it leaves several things unstated. There is no guidance on choosing a sharder for a given table size, no worked example of reading a planner output, and no discussion of what happens when a sharding plan is wrong. The performance claims are directional: the README says the pipelining is "for increased performance" without publishing numbers, and the external references are papers and integrations rather than benchmarks you can reproduce locally.
The requirements file is also worth reading before you commit. It pins hypothesis==6.70.1 and torchmetrics==1.0.3, which suggests the test suite is sensitive to those versions, and it lists fbgemm-gpu>=1.4.0 as a floor rather than a tested combination. Nothing in the README documents rollback or how to migrate an existing checkpoint into a sharded model, so a team moving from a single-device model should expect to solve checkpoint conversion themselves.
When a plain PyTorch embedding is the better choice
The honest alternative for many teams is not another library, it is nn.EmbeddingBag plus DistributedDataParallel. That setup replicates the full embedding table on every rank and averages gradients. It is simple, it needs no planner, and it works until the table exceeds device memory. TorchRec's model-parallel approach instead gives each rank a slice and routes lookups to the owning rank, which is more machinery but removes the replication ceiling.
The trade-off is real. Data-parallel replication costs memory per rank and scales communication with model size; TorchRec's sharding costs an input_dist step per forward pass and a planning step before training. If your tables fit, replication wins on operational simplicity. If they do not, replication is not an option at all. The other consideration is that TorchRec's optimized kernels come from FBGEMM, so a CPU-only environment gets the CPU path of a library whose feature list is written around multi-GPU training.
Maintenance, versions and the BSD-3-Clause licence
The repository is not archived, and the last push was on 2026-08-13, which is the same date as the v1.8.0 release. Recent releases are v1.6.0 in March 2026, v1.7.0 in June 2026 and v1.8.0 in August 2026, roughly a quarterly cadence. The release tags carry an -rc1 suffix, and setup.py appends the short git SHA to the version unless OFFICIAL_RELEASE is set in the environment, so a source install reports something like a release number plus a commit prefix rather than a clean version string. Pin explicitly if you build from source.
TorchRec is BSD-3-Clause, as stated in the README and the LICENSE file. That is a permissive licence, and it sits alongside a dependency on FBGEMM, which you install separately and which carries its own terms. Nothing here is legal advice; if you redistribute a container that bundles both, read the FBGEMM licence yourself. The upgrade cost is dominated by the CUDA and FBGEMM alignment rather than by TorchRec's API, which means a CUDA upgrade is a coordinated change across three packages.
Editorial conclusion
Adopt TorchRec if you are training recommender models whose embedding tables no longer fit on one device and you already run distributed PyTorch. Do not adopt it for a small ranking model that fits in a single GPU's memory, or if you cannot install a matching fbgemm-gpu build, since the optimized kernels come from FBGEMM. Before committing, verify that a torch, fbgemm-gpu and TorchRec combination resolves for your CUDA version, then run the repository's own test_installation.py through torchx in both GPU and CPU mode to confirm the sharded path works on your hardware.
Frequently asked questions
What is TorchRec?
It is a PyTorch domain library that provides sparsity and parallelism primitives for large-scale recommender systems. The README states that it powers many production RecSys models at Meta and that it allows training and inference of models with large embedding tables sharded across many GPUs.
Is PyTorch a coding language?
No. PyTorch is the framework TorchRec is built on, and TorchRec itself is a Python library whose source is Python. The README's installation steps invoke pip, Python's package installer, and python setup.py.
Does PyTorch use C or C++?
TorchRec's own optimized recommender kernels come from FBGEMM, which the README links as the source of those kernels and which is installed as a separate package. The TorchRec repository itself is Python, with a CMakeLists.txt at the top level.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/meta-pytorch-torchrec)