polars for the data, PySpark for the models, extras for everything else
A Comprehensive Framework for Building End-to-End Recommendation Systems with State-of-the-Art Models
At a glance
- What is it?
- RePlay is a recommendation framework whose install shape is the first thing you learn: a small core package, PySpark and PyTorch behind extras, and an experimental submodule that only exists in rc-tagged releases. Once it is installed you get splitters, item and feature generators, models from ItemKNN to SASRec and BERT4Rec, metrics wrapped in an Experiment object, and two-level ensembles.
- Who is it for?
- Adopt RePlay if you are comparing recommender architectures on shared datasets and want the comparison to be reproducible, since the numbered notebooks and the MovieLens-1M example exist for exactly that. Do not adopt it expecting the offline-to-online transition to be a solved feature, because the visible documentation states it as a capability without describing the mechanism.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.
Editorial analysis
One package, four install shapes
The quickstart line installs everything:
pip install replay-rec[all]And the plain install installs much less:
pip install replay-recThat gives you the core package without PySpark and without PyTorch, and the README adds the detail that matters most for planning: the experimental submodule is not installed either.
The extras are the two heavy dependencies. [spark] installs the PySpark functionality and [torch] installs PyTorch and Lightning. If you want the experimental submodule you have to name a version with an rc0 suffix, for example pip install replay-rec==XX.YY.ZZrc0, and the extras combine with that form as in pip install replay-rec[spark]==XX.YY.ZZrc0.
So there are four shapes to reason about: a light core, core plus Spark, core plus Torch, and an rc build that adds the experimental code. The split is deliberate rather than incidental, and it means a data scientist prototyping on a laptop does not install a cluster runtime to try a splitter.
Torch on CPU needs its own index URL
One install instruction is specific enough to save an evening. The torch extra can be installed with the CPU-only build of torch by supplying the index URL during installation:
pip install replay-rec[torch] --extra-index-url https://download.pytorch.org/whl/cpuThat matters because the default torch resolution will pull whatever the wheel index offers, and a recommendation experiment on a machine without a GPU is a normal thing to run.
Building from source is a separate path with its own documentation, referenced as the installing-from-the-source section of CONTRIBUTING.md rather than spelled out in the README.
Two CI systems are visible in the tree, a .github/ directory and a .gitlab/ one, alongside a poetry_wrapper.sh, which suggests the project is consumed both from a package index and from an internal packaging pipeline.
Four features need packages you install yourself
The optional features are listed with their install commands rather than as extras, and each one is a real capability rather than a toy.
Hyperparameter search comes from Optuna, pip install optuna. Model compilation via OpenVINO needs openvino, onnx and onnxscript together, which is a three-package install for one feature. Vector database and hierarchical search support needs hnswlib and a fixed-install nmslib, and the unusual package name is the pinned nmslib build rather than the original, which is what you would expect for a library that has historically had no working wheel.
The fourth is LightFM, marked experimental, and it comes with the sharpest caveat in the README: LightFM is not officially supported for Python 3.12 because the library's maintenance has been discontinued, and installing it locally requires a patched fork such as the one the project uses internally.
That last item is worth reading as a policy. A framework that names the fork it depends on, and says plainly that the upstream is unmaintained, is telling you what will happen when you come back to it in a year.
polars in the quickstart, PySpark behind it
The quickstart imports from both worlds, which is the clearest statement of the architecture:
from polars import from_pandas
from rs_datasets import MovieLens
from replay.data import Dataset, FeatureHint, FeatureInfo, FeatureSchema, FeatureType
from replay.metrics import HitRate, NDCG, Experiment
from replay.models import ItemKNN
from replay.splitters import RatioSplitter
ml_1m = MovieLens("1m")
interactions = from_pandas(ml_1m.ratings)Data arrives through rs_datasets and is converted into polars, while the Spark session comes from replay.utils.session_handler.State, with convert2spark in spark_utils for handing data across. Splitters come from replay.splitters, and the metrics are HitRate, NDCG and an Experiment object.
The division is not arbitrary. Notebook 11 is named as a speed comparison of different frameworks, pandas, polars and PySpark, which means the claim that polars-based preprocessing is fast is something the project measures rather than asserts.
The tree explains why Spark is a dependency at all: scala/ and jars/ directories sit beside the Python package, which is what a PySpark deployment with custom UDFs looks like. And the documented hardware support is CPU, GPU, multi-GPU, with PySpark integration for cluster computing.
Fifteen numbered notebooks, then the reinforcement learning ones
The examples directory is the documentation, and it is numbered in a teaching order rather than grouped by feature.
01_replay_basics gets you started. 02_models_comparison is the one that matters for evaluation, a reproducible models comparison on the MovieLens-1M dataset. 03 covers feature preprocessing and LightFM with PySpark doing the preprocessing. 04 is the splitters, 05 the feature generators, 06 item-to-item recommendations, 07 filters, and 08 recommendation for product categories.
The transformer models arrive next: 09 for SASRec, 10 for BERT4Rec, both generating recommendations. 11 is the framework speed comparison across pandas, polars and PySpark.
Past the visible README the numbering continues into less conventional territory, which is the best sign of scope: 12_neural_ts_exp for neural time series, 13_personalized_bandit_comparison, 14_hierarchical_recommender, 15_twotower_example, a cql_compare.py script, features_for_sequential_models.ipynb, and reinforcement learning scripts named train_ddpg.py and train_dt4rec.py alongside an obp_connector directory.
Two models, a hierarchical recommender and a bandit comparison are not the same kind of thing, and having them in one repository is the point of a framework rather than a model collection.
Experiment is an object, not a metric function
Evaluation is part of the framework rather than something you bolt on, and the quickstart shows how: HitRate and NDCG are imported from replay.metrics, alongside an Experiment object.
That object is the part that changes how results get reported. A metric function returns a number for one model on one split. An Experiment is the container, which is what makes a comparison reproducible rather than a set of numbers someone copied into a slide.
The feature list supports the same reading from the other direction. Data preprocessing and splitting are described as streamlining data preparation for efficient processing; the model range is described as spanning state-of-the-art models through commonly-used baselines; and evaluation metrics are listed as a wide set for assessing accuracy and effectiveness.
The baseline matters as much as the headline. ItemKNN appearing in the three-line import list of the quickstart tells you the framework's starting point is a method that predates all of the neural models above it, which is what you want when you are checking whether a new model beats something.
Ensembles are two-level, and the offline-to-online claim is thin
Two features are stated rather than explained, and they deserve different levels of scepticism.
Model Ensemble and Hybridization is concrete: it supports combining predictions from multiple models and creating two-level ensemble models. Two levels means a first layer of models and a second layer over their outputs, which is a specific architecture claim you can check against your own data.
Seamless Mode Transition is not concrete. It is listed as facilitating easy transition from offline experimentation to online production environments, ensuring scalability and flexibility, and the copy of the README available for this review stops before any section describing how that transition is performed.
That asymmetry is worth naming rather than repeating. Ensemble is a data structure you can implement; moving a model from a notebook to serving is an engineering project. If online deployment is your reason for choosing a framework, read the documentation for that path before you install anything, because the one-line feature description is not evidence.
Everything else in the visible feature list, from splitting to metrics to hyperparameter search, is a library-level capability you can evaluate in an afternoon.
Apache-2.0 with a NOTICE, and releases that trail the pushes
The licence is Apache-2.0 with both a LICENSE file and a NOTICE at the root, which is the conventional pairing for a project that vendors anything.
Maintenance is active. The last push was on 2026-09-30, and the recent releases are v0.21.8 on 2026-05-19, v0.21.7 on 2026-04-09 and v0.21.6 on 2026-03-27.
The gap between the newest tag and the last push is about four months, which for a library whose dependency extras include PyTorch and PySpark is unremarkable and worth noting if you pin a version: the extras resolve against whatever those projects publish now, not against what worked when the tag was cut.
The repository layout follows from the scope. The Python package lives in replay/, tests/ and scripts/ cover the code, docs/ and examples/ the documentation, and projects/ holds whatever larger work sits alongside it. The presence of .dockerignore without a Dockerfile in the root listing suggests containers are built elsewhere, which is consistent with a project that expects to run inside someone else's Spark cluster.
Editorial conclusion
Adopt RePlay if you are comparing recommender architectures on shared datasets and want the comparison to be reproducible, since the numbered notebooks and the MovieLens-1M example exist for exactly that. Do not adopt it expecting the offline-to-online transition to be a solved feature, because the visible documentation states it as a capability without describing the mechanism. Verify first which extras you actually need, since Optuna, OpenVINO, hnswlib and LightFM are all installed by hand rather than through extras, and that your Python version is supported by LightFM if you want it.
Frequently asked questions
How do I install RePlay?
pip install replay-rec installs the core package without PySpark and without PyTorch, and without the experimental submodule. Add [spark] or [torch] for those, [all] for everything, and name an rc0 version such as pip install replay-rec==XX.YY.ZZrc0 when you need the experimental code. Building from source is documented in CONTRIBUTING.md.
How do I install RePlay without a GPU build of torch?
Install the torch extra against PyTorch's CPU index with pip install replay-rec[torch] --extra-index-url https://download.pytorch.org/whl/cpu. Separately, the README notes LightFM is not officially supported on Python 3.12 because that library's maintenance was discontinued, so a patched fork is needed.
Which RePlay features need extra packages installed by hand?
Hyperparameter search via Optuna, model compilation via OpenVINO together with onnx and onnxscript, vector database and hierarchical search support via hnswlib and a fixed-install nmslib, and experimental LightFM support. None of these come through a pip extra, so each is installed manually.
What do the RePlay examples cover?
Basics, a reproducible models comparison on MovieLens-1M, feature preprocessing with LightFM and PySpark, splitters, feature generators, item-to-item recommendations, filters, category recommendation, SASRec, BERT4Rec and a pandas versus polars versus PySpark speed comparison, then neural time series, personalized bandits, a hierarchical recommender, a two-tower model and reinforcement learning scripts.
Does RePlay move a model from offline experiments to online serving?
Seamless Mode Transition is listed as a feature, covering the transition from offline experimentation to online production with scalability and flexibility. The visible README does not describe the mechanism behind that transition, so read the documentation for that path rather than relying on the one-line description.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/sb-ai-lab-replay)