Model or dataset
datawhalechina/fun-rec avatar
datawhalechina/fun-rec

datawhalechina/fun-rec: A Chinese-Language Recommendation Systems Course from Cascading Architectures to Generative Paradigms

推荐系统入门教程,在线阅读地址:https://datawhalechina.github.io/fun-rec/

7,381 stars1,027 forksPythonLicense varies

At a glance

What is it?
fun-rec (深度推荐算法实践, the wheat book) is a Datawhale open textbook that walks from collaborative filtering and vector recall to HSTU, OneRec and diffusion-based recommendation. It is a reading and reference project, not a library you import.
Who is it for?
Adopt fun-rec if you already know machine learning basics and want a structured map of recommendation techniques from collaborative filtering through HSTU, OneRec and diffusion models, with runnable code in src/ and a production chapter at the end. Do not adopt it as a library: the pyproject.toml version is 0.0.1, the README states the project is still under development and is not accepting pull requests, and the package is not published as a stable API.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 94 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What fun-rec actually is, and who the README is written for

fun-rec is a book, not a framework. The README describes it as a systematic treatment of how recommendation systems evolved from traditional cascading architectures to generative paradigms, and splits the content into two halves: an upper part on candidate retrieval, ranking and re-ranking, and a lower part on generative recommendation. The repository carries a src/ directory, a docs/ directory, a web_project/ directory and a pyproject.toml that names the package funrec at version 0.0.1, so there is installable code, but the README's own framing is pedagogical. The audience it names is readers who already have machine learning fundamentals and want both the core principles and the engineering practice of recommendation algorithms.

The scope is unusually wide for a single repository. Chapter 2 covers item-based and user-based collaborative filtering plus matrix factorization, then vector recall (I2I and U2I) and sequence recall with streaming indexes. Chapter 3 moves into feature crossing, sequence modeling with attention, multi-task learning and multi-scenario modeling. Chapter 4 handles re-ranking diversity through greedy methods such as maximum marginal relevance and determinantal point processes, plus personalized re-ranking with Transformer and permutation-based models. The lower half is the part that distinguishes fun-rec from the older generation of recommendation tutorials: HSTU and scaling-law architecture exploration, OneRec V1 and V2, OneSug and OneSearch, the OneRec-Think reasoning framework, and diffusion-based approaches including DiffuASR, Diff-MSR, AsymDiffRec and DMSG.

If you are looking for a scikit-learn-style API that gives you a recommender in ten lines, this is the wrong repository. If you want to understand why industrial recommenders moved from a retrieval-ranking-re-ranking cascade toward end-to-end generative modeling, the chapter list is a reasonable syllabus.

How the repository is organized: docs, src and a production project

The top level contains .env.example, .gitignore, README.md, README_en.md, docs/, imgs/, pyproject.toml, requirements.txt, src/ and web_project/. That layout tells you the intended workflow. The docs/ tree holds the book text, which is also published online through the GitHub Pages site linked in the README. The src/ tree holds the package code, with setuptools configured to find packages under src and to include YAML and Python files from a config directory inside the funrec package. The web_project/ directory corresponds to Chapter 10, which the README describes as a production-grade recommendation system build covering system architecture, offline pipelines, online flows, frontend interaction, and deployment and operations.

Data paths are externalized. The .env.example file defines two variables:

bash
FUNREC_RAW_DATA_PATH=<YOUR_RAW_DATA_PATH>
FUNREC_PROCESSED_DATA_PATH=<YOUR_PROCESSED_DATA_PATH>

python-dotenv is a pinned dependency, which is consistent with the code reading those variables at runtime rather than hardcoding dataset locations. The README does not document the expected directory structure under those paths, so you will need to read the chapter code to learn what filenames it expects.

The dependency list is opinionated in a way worth noting. tensorflow==2.13.0 is pinned exactly, not with a range, and pandas==2.0.3 is pinned exactly as well. The requirements.txt adds seaborn==0.13.2, faiss-cpu==1.7.4 and lightgbm==4.6.0 under a comment saying they are needed for the news recommendation chapter. faiss-cpu rather than faiss-gpu is a deliberate default that keeps installation simple but caps you at CPU vector search, which matters if you intend to run the recall chapters on a real dataset.

Installing funrec and running a first chapter

The package declares requires-python >=3.8 and depends on tensorflow==2.13.0. Because that TensorFlow version is pinned exactly, create an isolated environment rather than installing into an existing one. From the repository root:

bash
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
pip install -e .

The first command creates the environment, the second activates it on Linux or macOS, the third installs the pinned dependency set including the news-recommendation extras, and the fourth installs the funrec package itself in editable mode using the setuptools configuration in pyproject.toml. On Windows the activation line is .venv\Scripts\activate instead.

Before running any chapter code, copy the environment template and fill in your own paths:

bash
cp .env.example .env

Then edit .env so that FUNREC_RAW_DATA_PATH points at your raw dataset directory and FUNREC_PROCESSED_DATA_PATH points at a writable directory for intermediate files. The README does not state which datasets the chapters expect, so read the chapter you intend to run before choosing directories.

What you should see after a successful install is an importable funrec package and a resolved dependency tree with TensorFlow 2.13. If pip resolves a different TensorFlow version, the pin has been overridden by another requirement in your environment and you should rebuild the virtualenv cleanly. Beyond this, the README gives no quickstart command and no sample invocation of the package API, so the honest answer is that the first real use is running a chapter's scripts from src/ rather than calling a documented entry point.

Where fun-rec stops being useful

The README carries a warning near the top: the project is still under development, code and content are updated frequently, and pull requests are not accepted at this time. Feedback is directed to the issue tracker. For a reader this is mostly harmless. For anyone planning to build on the code, it is a real constraint, because the interfaces you depend on can change without a deprecation cycle and you cannot upstream a fix.

The exact TensorFlow pin is the second constraint. tensorflow==2.13.0 does not coexist easily with projects that require a newer TensorFlow, and it will not run on Python versions that 2.13 does not support. If your team has standardized on a recent TensorFlow or on PyTorch, fun-rec is a reading resource rather than something to add to your dependency graph.

Language is the third. The README is bilingual, with README.md in Chinese and README_en.md in English, but the book content under docs/ is the Chinese-language material described in the README's table of contents, and the online version is served from the Chinese GitHub Pages site. An English reader gets the repository description and the chapter titles, not the chapters themselves. That is a genuine barrier, not a footnote.

Finally, the licence is CC BY-NC-SA 4.0. The NonCommercial clause is the one that catches people. A commercial team can read the material, but redistributing adapted content inside a commercial training program is a different question, and the licence text, not this article, governs it.

How fun-rec differs from other recommendation learning resources

The obvious comparison is RecBole, the Python recommendation library from a different research group. RecBole's premise is that you pick a model, point it at a dataset in a supported format, and run training and evaluation through a unified pipeline with dozens of implemented algorithms. fun-rec's premise is the opposite: each chapter explains a technique and the accompanying code in src/ demonstrates it, so the value is in the explanation rather than in a stable abstraction over models. If your goal is to benchmark five recall models on your own data this week, RecBole is the tool. If your goal is to understand what problem each of those models was invented to solve, fun-rec is the better fit, and the two are complementary rather than competing.

A second comparison is the older Datawhale recommendation material and the many blog-series tutorials on the same topic. What separates fun-rec is the lower half of the table of contents. Most Chinese-language recommendation tutorials stop at multi-task ranking models. fun-rec continues into HSTU's scaling-law exploration, OneRec's end-to-end generative architecture, tokenizer design for recommendation, chain-of-thought reasoning frameworks such as OneRec-Think and PLUM, and diffusion-based augmentation. That is current research territory, and covering it in a structured syllabus is the repository's main claim to attention.

The cost of that ambition is depth per topic. Ten chapters covering everything from matrix factorization to diffusion models cannot go as deep on any single model as a dedicated paper walkthrough would. Treat the chapters as guided entry points with references, not as the final word.

Maintenance, licence and the cost of following along

The repository is not archived, and the last push was on 2026-06-27. There are no retrieved releases, which is consistent with a version number of 0.0.1 and with a project that ships content rather than versioned artifacts. Upgrading means pulling the master branch and re-reading whatever changed, since there is no changelog in the repository and no release notes to consult.

The practical upgrade cost sits in the dependency pins. tensorflow==2.13.0, pandas==2.0.3, gensim==4.3.3, networkx==3.1 and faiss-cpu==1.7.4 are all fixed versions. If you keep a long-lived environment for the book's code, expect to rebuild it when you move to a newer Python, because the pinned TensorFlow will hold you back. The alternative is to treat the code as illustrative and reimplement the ideas against your own stack, which is what most readers of a textbook end up doing anyway.

On licensing: pyproject.toml declares license = { text = "CC BY-NC-SA 4.0" } and the README links to the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International licence. The NonCommercial term restricts commercial use of the material, and ShareAlike means derivative works carry the same licence. This is a content licence applied to a repository that also contains code, which is an unusual combination and worth checking against your organization's policy before you copy chapter code into an internal project. Nothing here is legal advice; read the licence text.

Editorial conclusion

Adopt fun-rec if you already know machine learning basics and want a structured map of recommendation techniques from collaborative filtering through HSTU, OneRec and diffusion models, with runnable code in src/ and a production chapter at the end. Do not adopt it as a library: the pyproject.toml version is 0.0.1, the README states the project is still under development and is not accepting pull requests, and the package is not published as a stable API. Before committing time, verify three things: that your Python environment can hold tensorflow==2.13.0 alongside your existing stack, that you can read Chinese or are willing to work through the Chinese docs/ tree, and that the chapter you care about is actually written, since the README lists a full table of contents but says content is updated frequently.

Frequently asked questions

What is fun-rec?

fun-rec is a Datawhale open textbook on recommendation systems, titled 深度推荐算法实践 and subtitled from cascading architectures to generative paradigms. It covers candidate recall, ranking, re-ranking, and then generative approaches such as HSTU, OneRec and diffusion-based recommendation, with a production system chapter at the end.

Is fun-rec a library I can install and call?

It is primarily a book. The repository does contain a src/ package named funrec at version 0.0.1 with a pyproject.toml, and requirements.txt lists the dependencies, but the README frames the project as a tutorial and does not document a stable public API.

What Python and TensorFlow versions does fun-rec need?

pyproject.toml declares requires-python >=3.8 and pins tensorflow==2.13.0 and pandas==2.0.3 exactly. requirements.txt adds seaborn==0.13.2, faiss-cpu==1.7.4 and lightgbm==4.6.0 for the news recommendation chapter.

Is fun-rec still being updated?

The repository is not archived and the last push was on 2026-06-27. The README states the project is still under development, that code and content are updated frequently, and that pull requests are not accepted at this time, with feedback directed to the issue tracker.

What licence does fun-rec use?

pyproject.toml declares CC BY-NC-SA 4.0 and the README links to the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International licence. The NonCommercial and ShareAlike terms both apply to reuse of the material.

Official sources

  1. datawhalechina/fun-rec on GitHub
  2. Issues
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/datawhalechina-fun-rec.svg)](https://hysenlabs.com/projects/datawhalechina-fun-rec)