Open-source project
yandexdataschool/Practical_RL avatar
yandexdataschool/Practical_RL

Practical_RL: the YSDA reinforcement learning course you run yourself

A course in reinforcement learning in the wild

6,580 stars1,808 forksJupyter NotebookUnlicense

At a glance

What is it?
Practical_RL is an open reinforcement learning course from YSDA and HSE, distributed as Jupyter notebooks organised into weekly folders. It is a strong fit for engineers who already write Python and want labs, not for anyone looking for a maintained library or a pip-installable toolkit.
Who is it for?
Adopt Practical_RL if you already write Python and want a structured sequence of notebooks that takes you from crossentropy methods to PPO, with homework you can actually run. Do not adopt it if you need a supported library, a stable API, or a single install that pins every dependency for you: this is course material, and the README points at an issue thread for local setup rather than a setup script.
Can I use it commercially?
Yes. Unlicense is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What Practical_RL actually is, and the gap it fills

Most reinforcement learning material falls into two buckets. Textbooks give you the derivations and leave the debugging to you. Single-file reference implementations give you a working agent and no explanation of why the pieces are arranged that way. Practical_RL sits between them: it is a taught course, released as a repository of Jupyter notebooks, with a lecture and a seminar for each week and a homework description in each week's README.

The README states the course is taught on-campus at HSE and YSDA and maintained to be friendly to online students in both English and Russian. The audience is therefore someone who wants the on-campus experience without the campus: a working programmer who has read about Q-learning and wants to implement it against a real environment, then move on to policy gradients and trust regions in the same structured way.

The manifesto is unusually explicit about the editorial stance. It says the course optimises for the curious, that materials not covered in detail link out to Sutton, Silver and blogs, and that assignments carry bonus sections. It also says practicality comes first, and that heuristics and tricks are not shunned. That last point matters. A course that admits heuristics exist is more useful to a practitioner than one that pretends every method follows cleanly from theory.

The week-by-week structure and how the notebooks fit together

The syllabus is organised as top-level directories, one per topic, and the repository layout confirms them: week01_intro, week02_value_based, week03_model_free, week04_[recap]_deep_learning, week04_approx_rl, week05_explore, week06_policy_based, week07_seq2seq, week08_pomdp, week09_policy_II, week10_planning, and yet_another_week. There is also a week07_[recap]_rnn directory, so two different weeks use the number seven in their folder names. If you script anything against these paths, read the actual directory listing rather than assuming a clean sequence.

The ordering is deliberate. Week one covers decision processes and the crossentropy method, comparing parameter-space search against action-space search. Week two moves to value iteration and policy iteration on discounted reward MDPs. Week three drops the model and covers Q-learning, SARSA, off-policy against on-policy, n-step methods and TD(lambda). Only after that does the course introduce function approximation in week04_approx_rl, with experience replay, target networks and the DQN variants. Exploration gets its own week, covering contextual bandits, Thompson sampling and UCB, including UCB for MCTS. Policy gradients arrive in week six, planning and model-based methods in week ten.

The README notes the syllabus is approximate, and that lectures may occur in a slightly different order or a topic may take two weeks. Treat the folder order as the intended reading order, not as a contract. The README also says the course massively refers to CS294 and uses pictures from the Berkeley AI course, so some slides will look familiar if you have watched those lectures.

Installing Practical_RL and running your first notebook

There is no package to install. The README gives two routes: a virtual environment on Google Colab, or a local install. For Colab it says to set open, then github, then yandexdataschool/practical_rl, then pick a branch and select any notebook. Note that the README spells the path with a typo, practical_rl rather than Practical_RL, so if the Colab picker does not resolve, correct the casing by hand.

For a local machine the README does not provide a requirements file or an install command. It links to issue thread 1, titled as the technical issues thread, which is where dependency installation is discussed. That is the honest state of the project: setup guidance lives in an issue, not in a script. The repository does include a docker directory, so a container definition exists in the tree, but the README does not document how to build or run it.

The repository also ships setup_colab.sh at the top level, which is the script the Colab notebooks use to prepare the environment. If you are working locally, that file is the closest thing to a dependency manifest, and it is worth reading before you start installing things by hand.

Once an environment exists, the workflow is the same as any notebook course. Open the week's notebook, run the cells in order, and read the week's README for the homework description. The README says homework descriptions live in files like week1/README.md, so the assignment text and the notebook are separate artefacts and you need both.

Frameworks, environments and the dependency risk

The repository topics list both pytorch and tensorflow alongside keras, and the README credits several TensorFlow assignments to a contributor. That means the course is not framework-neutral in practice: different weeks were written against different libraries, and some weeks may exist in more than one version. The README's manifesto explicitly invites pull requests for an alternative-framework version of the code, which tells you the maintainers expect framework drift and treat community ports as the answer.

The environments come from OpenAI Gym, which the README names directly in the week one seminar description: tabular CEM for Taxi-v0 and deep CEM for box2d environments. Gym has changed its API and its distribution since this course was written, and the README does not address that. Nothing in the project documentation states which Gym version the notebooks assume.

The practical consequence is that the hardest part of this course is not the reinforcement learning. It is getting a notebook from an older dependency era to execute on a current machine. Budget for that. If you are evaluating the course for a team, the honest test is to pick one mid-course notebook, say the approximate Q-learning seminar with experience replay, and try to run it end to end before you plan anything around the rest.

Where the course stops being the right tool

Practical_RL is teaching material, and it behaves like teaching material. There is no versioned API, no changelog for the notebooks, and no support commitment. The releases listed are spring20 and spring19, dated 2020-08-04 and 2019-06-17, so the release history is sparse and old. The last push to the repository was on 2026-03-31, which is recent enough that the project is not abandoned, but a recent push does not mean the notebooks were re-verified against current library versions.

Two specific failure modes are worth naming. First, if you want to train a model that does something useful, this is the wrong repository. The labs target toy environments: Taxi-v0, CartPole, box2d, simple robot control. That is by design, and it is the right design for learning, but it will not give you a production agent.

Second, if you need reproducibility, the course does not offer it in the sense a research codebase does. Homework descriptions are separate from notebooks, the syllabus is described as approximate, and the dependency story runs through an issue thread. You can still use it, but you will be assembling your own pinned environment, and the course will not tell you which versions to pin.

A third case is more subtle. If you already know deep learning and want to skip to modern methods, the recap weeks will feel slow. The README includes a deep learning recap week and an RNN recap week, which is appropriate for the on-campus audience and redundant for someone who trains networks daily.

How it compares with CleanRL and single-file implementations

CleanRL takes the opposite approach on almost every axis. It is a set of self-contained training scripts, one file per algorithm, designed so you can read the whole algorithm in one sitting and run it without a course structure around it. Practical_RL gives you explanation, sequencing and homework, and asks you to work through notebooks over weeks. CleanRL gives you code and expects you to bring the theory.

The trade-off is real in both directions. With CleanRL you get something you can copy into a project today, but you get no guidance on what to learn next or why one algorithm beats another on a given problem. With Practical_RL you get the guidance and the progression, but the code is written to be understood in a seminar, not to be dropped into a training pipeline. The notebooks are also tied to a specific era of Gym and to a mix of PyTorch and TensorFlow, which CleanRL avoids by being narrower and more current.

There is a middle path worth mentioning: the course's own reading list. The README links an RL reading group wiki and says the materials point to Sutton, Silver and related work wherever detail is missing. If you use Practical_RL, use those links. The course is explicit that it is a starting point rather than a complete treatment, and the yet_another_week directory is labelled as all the cool RL material you will not learn from this course.

Licence, maintenance and what an upgrade costs you

The repository is released under the Unlicense. That is a public-domain dedication rather than a permissive licence with conditions, which means the usual attribution and copyleft questions do not arise from the licence text itself. This is not legal advice, and if you plan to reuse course material in a commercial setting you should read LICENSE.md and the third-party attributions yourself. The README's contributions section credits pictures from the Berkeley AI course and heavy reference to CS294, so some assets inside the repository carry their own provenance that the Unlicense does not override.

On maintenance: the last push was on 2026-03-31, and the most recent release is spring20 from 2020-08-04. That combination describes a course that receives occasional fixes and contributions but whose packaged releases stopped six years ago. The README's git-course manifesto invites pull requests for typos, formulas, links and readability, so small improvements do flow in. Do not read that as a support commitment.

The upgrade cost is therefore not a version bump. It is the cost of reconstructing an environment that matches notebooks written across several years, with a mixed PyTorch and TensorFlow surface and an undocumented Gym version. If you are adopting this for a study group, that reconstruction is a one-time task you can share. If you are adopting it alone, it is the first thing you will do and possibly the thing that stops you.

Editorial conclusion

Adopt Practical_RL if you already write Python and want a structured sequence of notebooks that takes you from crossentropy methods to PPO, with homework you can actually run. Do not adopt it if you need a supported library, a stable API, or a single install that pins every dependency for you: this is course material, and the README points at an issue thread for local setup rather than a setup script. Before committing time, open the week you care about, check that its notebook matches the framework you use (the repository carries both PyTorch and TensorFlow assignments), and confirm the README's local install path still resolves on your machine.

Frequently asked questions

What exactly is reinforcement learning, and how does Practical_RL introduce it?

The course opens with week01_intro, whose lecture covers RL problems, decision processes, stochastic optimization and the crossentropy method, comparing parameter-space search with action-space search. The seminar then moves into OpenAI Gym with tabular CEM for Taxi-v0 and deep CEM for box2d environments.

What are some practical examples of reinforcement learning covered by Practical_RL?

The labs target concrete environments rather than abstract descriptions: Taxi-v0 and box2d in week one, CartPole and Atari in the approximate RL week, a character-level RNN language model in the sequence week, and approximate TRPO for simple robot control in the advanced policy week.

How do I install Practical_RL and run it locally?

There is no install command in the README. It recommends installing dependencies on your local machine and links to the technical issues thread (issue 1) for that, or alternatively running the notebooks on Google Colab by selecting the yandexdataschool/practical_rl repository and a branch. The repository also contains a docker directory and a setup_colab.sh script.

Which frameworks does Practical_RL use, PyTorch or TensorFlow?

Both appear. The repository topics list pytorch, tensorflow and keras, and the README credits several TensorFlow assignments to a contributor. The manifesto invites pull requests that port code to alternative frameworks, so the framework surface is mixed rather than uniform.

Is Practical_RL actively maintained?

The last push was on 2026-03-31, so the repository still receives changes, but the most recent release is spring20 from 2020-08-04. The README describes the project as a git-course that welcomes pull requests for fixes and improvements, which is not the same as a support commitment.

Official sources

  1. Issues
  2. License: Unlicense
  3. README
  4. Releases
  5. yandexdataschool/Practical_RL on GitHub
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/yandexdataschool-practical-rl.svg)](https://hysenlabs.com/projects/yandexdataschool-practical-rl)
Community notes

Community notes