# TorchRec: sharding embedding tables across many GPUs

> TorchRec is a PyTorch domain library for large-scale recommender systems. It supplies the sharding, planning and pipelining primitives that a plain nn.EmbeddingBag does not, and it expects a distributed GPU environment to be useful.

**meta-pytorch/torchrec** — Pytorch domain library for recommendation systems. TorchRec TorchRec** is a PyTorch domain library built to provide common sparsity and parallelism primitives needed for large-scale recommender systems (RecSys).

- Repository: https://github.com/meta-pytorch/torchrec
- Website: https://pytorch.org/torchrec/
- Stars: 2,607 · Forks: 698
- Language: Python
- License: BSD-3-Clause
- Published: 2026-08-04 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/meta-pytorch-torchrec

## The problem: embedding tables that outgrow one device

A recommender model is mostly embedding table. The dense part that turns features into a prediction is small; the vocabulary side is where the parameters live, and it grows with users and items rather than with model depth. A single nn.EmbeddingBag is a single tensor, and a single tensor has to sit on a single device. Once the table passes the memory of one GPU, the usual workarounds appear: shrink the embedding dimension, hash the vocabulary, or keep the table on CPU and pay the transfer cost on every step.

TorchRec exists to remove that constraint. The README describes it as a PyTorch domain library providing "common sparsity and parallelism primitives needed for large-scale recommender systems (RecSys)", and states that it powers many production RecSys models at Meta. The intended reader is an engineer who already writes PyTorch models and now needs the embedding side distributed across devices and nodes without hand-rolling the collective communication. If your model fits comfortably on one accelerator, this library is overhead rather than help.

## How sharding, planning and pipelining fit together

TorchRec splits the work into three cooperating pieces. Sharders decide how a table is divided. The README lists data-parallel, table-wise, row-wise, table-wise-row-wise, column-wise and table-wise-column-wise sharding. The distinction matters: row-wise sharding cuts the vocabulary so each rank owns a slice of rows, which is what you want when one table is enormous, while table-wise sharding assigns whole tables to ranks, which suits a model with many medium tables and an uneven access pattern.

The planner generates a sharding plan automatically, choosing an assignment rather than making you write one. The README describes it as a planner "that can automatically generate optimized sharding plans for models". In practice this is where most of the tuning happens, because the plan determines how much cross-device traffic each step produces.

The third piece is the pipelined training loop, which overlaps dataloading device transfer (copy to GPU), inter-device communication (input_dist) and computation (forward, backward). Those three stages are named explicitly in the README, and the overlap is the reason the communication cost of sharding does not simply serialize on top of compute. Underneath, the optimized kernels come from FBGEMM, which is why the installation path pulls in fbgemm-gpu before TorchRec itself.

## Installing TorchRec and running the installation test

The README points to the Getting Started page in the documentation for recommended setup and states plainly that "Generally, there isn't a need to build from source". The source path is documented for people who want the latest changes. It begins with PyTorch, installed from a CUDA-specific index. This is the CUDA 12.6 case:

```bash
pip install torch --index-url https://download.pytorch.org/whl/nightly/cu126
```

FBGEMM comes next, from the matching index, because the optimized RecSys kernels live there:

```bash
pip install fbgemm-gpu --index-url https://download.pytorch.org/whl/nightly/cu126
```

Then clone the repository recursively, install the listed requirements, and install TorchRec itself:

```bash
git clone --recursive https://github.com/meta-pytorch/torchrec
cd torchrec
pip install -r requirements.txt
python setup.py install develop
```

The recursive clone matters because the build expects submodules. The README then gives a verification step that launches a two-process job through TorchX:

```bash
torchx run -s local_cwd dist.ddp -j 1x2 --gpu 2 --script test_installation.py
```

On a machine without GPUs, the same script takes a CPU flag:

```bash
torchx run -s local_cwd dist.ddp -j 1x2 --script test_installation.py -- --cpu_only
```

A successful run means the sharded embedding path initializes and communicates across the two ranks. If it fails, the mismatch is almost always between the torch build, the fbgemm-gpu build and the CUDA version, not in TorchRec's own code. For a fuller example, the README points at the DLRM example in the facebookresearch/dlrm repository.

## What the repository does not tell you

The README is a feature list with an installation appendix, and it leaves several things unstated. There is no guidance on choosing a sharder for a given table size, no worked example of reading a planner output, and no discussion of what happens when a sharding plan is wrong. The performance claims are directional: the README says the pipelining is "for increased performance" without publishing numbers, and the external references are papers and integrations rather than benchmarks you can reproduce locally.

The requirements file is also worth reading before you commit. It pins hypothesis==6.70.1 and torchmetrics==1.0.3, which suggests the test suite is sensitive to those versions, and it lists fbgemm-gpu>=1.4.0 as a floor rather than a tested combination. Nothing in the README documents rollback or how to migrate an existing checkpoint into a sharded model, so a team moving from a single-device model should expect to solve checkpoint conversion themselves.

## When a plain PyTorch embedding is the better choice

The honest alternative for many teams is not another library, it is nn.EmbeddingBag plus DistributedDataParallel. That setup replicates the full embedding table on every rank and averages gradients. It is simple, it needs no planner, and it works until the table exceeds device memory. TorchRec's model-parallel approach instead gives each rank a slice and routes lookups to the owning rank, which is more machinery but removes the replication ceiling.

The trade-off is real. Data-parallel replication costs memory per rank and scales communication with model size; TorchRec's sharding costs an input_dist step per forward pass and a planning step before training. If your tables fit, replication wins on operational simplicity. If they do not, replication is not an option at all. The other consideration is that TorchRec's optimized kernels come from FBGEMM, so a CPU-only environment gets the CPU path of a library whose feature list is written around multi-GPU training.

## Maintenance, versions and the BSD-3-Clause licence

The repository is not archived, and the last push was on 2026-08-13, which is the same date as the v1.8.0 release. Recent releases are v1.6.0 in March 2026, v1.7.0 in June 2026 and v1.8.0 in August 2026, roughly a quarterly cadence. The release tags carry an -rc1 suffix, and setup.py appends the short git SHA to the version unless OFFICIAL_RELEASE is set in the environment, so a source install reports something like a release number plus a commit prefix rather than a clean version string. Pin explicitly if you build from source.

TorchRec is BSD-3-Clause, as stated in the README and the LICENSE file. That is a permissive licence, and it sits alongside a dependency on FBGEMM, which you install separately and which carries its own terms. Nothing here is legal advice; if you redistribute a container that bundles both, read the FBGEMM licence yourself. The upgrade cost is dominated by the CUDA and FBGEMM alignment rather than by TorchRec's API, which means a CUDA upgrade is a coordinated change across three packages.

## Conclusion

Adopt TorchRec if you are training recommender models whose embedding tables no longer fit on one device and you already run distributed PyTorch. Do not adopt it for a small ranking model that fits in a single GPU's memory, or if you cannot install a matching fbgemm-gpu build, since the optimized kernels come from FBGEMM. Before committing, verify that a torch, fbgemm-gpu and TorchRec combination resolves for your CUDA version, then run the repository's own test_installation.py through torchx in both GPU and CPU mode to confirm the sharded path works on your hardware.

## FAQ

### What is TorchRec?

It is a PyTorch domain library that provides sparsity and parallelism primitives for large-scale recommender systems. The README states that it powers many production RecSys models at Meta and that it allows training and inference of models with large embedding tables sharded across many GPUs.

### Is PyTorch a coding language?

No. PyTorch is the framework TorchRec is built on, and TorchRec itself is a Python library whose source is Python. The README's installation steps invoke pip, Python's package installer, and python setup.py.

### Does PyTorch use C or C++?

TorchRec's own optimized recommender kernels come from FBGEMM, which the README links as the source of those kernels and which is installed as a separate package. The TorchRec repository itself is Python, with a CMakeLists.txt at the top level.

## Sources

- [Official documentation](https://pytorch.org/torchrec/)
- [Official README](https://github.com/meta-pytorch/torchrec#readme)
- [Project repository](https://github.com/meta-pytorch/torchrec)
- [Release notes](https://github.com/meta-pytorch/torchrec/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/meta-pytorch-torchrec
