Model or dataset
udacity/deep-reinforcement-learning avatar
udacity/deep-reinforcement-learning

udacity/deep-reinforcement-learning: A PyTorch Course Repository, Not a Library

Repo for the Deep Reinforcement Learning Nanodegree program

5,173 stars2,369 forksJupyter NotebookMIT

At a glance

What is it?
The Udacity Deep Reinforcement Learning Nanodegree repository is a set of Jupyter notebooks that walks through dynamic programming, DQN, DDPG and policy gradients in PyTorch. It is a teaching artifact with course-shaped boundaries, and it should be judged as one.
Who is it for?
Adopt this repository if you are working through the Deep Reinforcement Learning Nanodegree or want the algorithm notebooks as a reference implementation to read alongside a textbook. Do not adopt it as a dependency: there is no package to install, the notebooks pin PyTorch v0.4 and Python 3, and the README marks PPO and the CarRacing DQN as coming soon.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 84 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What the Udacity Deep Reinforcement Learning Repository Actually Is

This is the companion repository for Udacity's Deep Reinforcement Learning Nanodegree, published under the MIT license and written almost entirely in Jupyter Notebook. It is not a framework, not a package on PyPI, and not something you import into a production training pipeline. The README describes it as material related to the program, and the directory listing confirms that reading: dynamic-programming, monte-carlo, temporal-difference, discretization, tile-coding, dqn, hill-climbing, cross-entropy, reinforce, ddpg-pendulum, ddpg-bipedal, finance, lab-taxi, p1_navigation, p2_continuous-control and p3_collab-compet.

The audience is narrow and explicit. The tutorials section says the code is in PyTorch v0.4 and Python 3, and the projects section says the labs and projects use simulation environments from Unity ML-Agents. Someone looking for a maintained RL library with a stable API will not find it here. Someone working through the course, or reading reinforcement learning theory and wanting runnable notebooks that match the chapters, is exactly who this was built for.

How the Notebooks and Unity Environments Fit Together

The repository splits into two mechanisms. The tutorials are self-contained notebooks that implement an algorithm and run it against an OpenAI Gym environment. The README's benchmark list makes that pairing concrete: Cartpole-v0 with Hill Climbing, solved in 13 episodes; Cartpole-v0 with REINFORCE, solved in 691 episodes; MountainCarContinuous-v0 with the Cross-Entropy Method, solved in 47 iterations; LunarLander-v2 with DQN, solved in 1504 episodes; MountainCar-v0 with uniform-grid discretization and Q-Learning, solved in under 50000 episodes. Each line points at a specific notebook file, so the algorithm and the environment are not separated by an abstraction layer.

The projects work differently. They depend on Unity ML-Agents environments rather than Gym, which is why p1_navigation, p2_continuous-control and p3_collab-compet each carry their own setup burden. The README also notes that in the Nanodegree program you receive a review of your project, with personalized feedback. That review loop is the actual product; the repository is the substrate. Reading the code without the review is possible, but you lose the part of the design that assumes a human will read your submission.

One structural detail worth noting: the Robotics tutorial is an external link to dusty-nv/jetson-reinforcement, a C++ API, not a notebook in this tree. The README marks it as an external link, so the repository is honest about the boundary of what it contains.

Setting Up the Python Environment and Opening a Notebook

The README's Dependencies section begins by telling you to create and activate a new environment with Python, and the text available stops there. It does not spell out the exact conda or pip commands, and the repository has a python/ directory at the top level where the environment files live. Because the README does not give a runnable install command, this section stays at the level the repository documents: create and activate an environment with Python, install the dependencies found in python/, then launch Jupyter from the repository root so the relative paths inside each notebook resolve.

The README does give one concrete command, in step 1 of the Dependencies section: create and activate a new environment with Python. The repository's python/ directory is where the dependency files sit, so check that directory before installing anything. The README does not name the file, so no install command can be quoted here without inventing one.

Once the environment is ready, start the notebook server from the repository root. Then open a tutorial notebook, for example the Hill Climbing notebook that the README lists against Cartpole-v0, and run the cells in order. You should see episode-by-episode output as the agent trains, and the README states that this particular configuration solves Cartpole-v0 in 13 episodes. If you see import errors instead, the environment is the problem, not the notebook.

The PyTorch v0.4 Pin Is the Real Constraint

The README states plainly that all tutorial code is in PyTorch v0.4 and Python 3. That is the single most consequential fact about using this repository today. PyTorch v0.4 predates a large amount of API surface that current PyTorch code assumes, and a notebook written against it will not necessarily run unchanged on a modern install. The repository has been pushed to recently, but a recent push does not rewrite the notebooks, and the README still carries the v0.4 statement.

This is not a defect in a teaching repository. Course material is written against the version the course teaches, and rewriting notebooks every time PyTorch ships a release would break the correspondence between the lessons and the code. But it means the repository is the wrong tool if what you want is a drop-in reference for current PyTorch idioms. You will spend time translating, not training.

The same applies to the Gym environments. The benchmark list names Acrobot-v1, Cartpole-v0, MountainCarContinuous-v0, MountainCar-v0, Pendulum-v0, BipedalWalker-v2, CarRacing-v0, LunarLander-v2, FrozenLake-v0, Blackjack-v0 and CliffWalking-v0. Several of those carry version suffixes that later Gym releases changed. The README is a snapshot of what the course targeted, and it should be read as one.

Where the Repository Is Incomplete by Design

The README lists Proximal Policy Optimization with the note coming soon, and separately lists CarRacing-v0 with Deep Q-Networks with the same note. Neither is present as a finished notebook. If PPO is what you came for, this repository does not have it, and the README says so rather than leaving you to discover it.

The finance tutorial is listed as training an agent to discover optimal trading strategies, with no benchmark entry in the OpenAI Gym section and no environment named. That is thin compared to the Cartpole and LunarLander entries, which give episode counts. The same asymmetry shows up in the DDPG entries: Pendulum-v0 and BipedalWalker-v2 are listed without solved-in numbers, unlike the classic control and Box2d DQN results. Where the README gives a number, it is telling you the configuration was run to completion. Where it does not, you should not assume the same.

A second limitation is that the projects depend on Unity ML-Agents environments that the repository does not ship. The README points at Unity's own repository. Setting those up is a separate task with its own version constraints, and the README does not walk through it.

Compared with Stable-Baselines3 and CleanRL

The obvious alternative for someone who wants to train an agent rather than study one is Stable-Baselines3, which ships as an installable Python package with a consistent API across algorithms. The difference in approach is fundamental. Stable-Baselines3 hides the training loop behind a model object; this repository puts the loop in a notebook cell so you can read every line and change it. If you want to run DQN on a new environment this afternoon, the library wins. If you want to understand why the target network exists, the notebook wins.

CleanRL takes a third position: single-file implementations, no shared abstraction, closer to research code than to a library. That is nearer in spirit to this repository than Stable-Baselines3 is, but CleanRL is maintained as a set of reference implementations rather than as course material tied to a specific PyTorch version and a specific set of Gym environments. The trade-off is that CleanRL assumes more background from the reader. This repository assumes less, which is the point of a nanodegree.

Licence, Maintenance and Upgrade Cost

The repository is MIT licensed, which permits reuse, modification and redistribution with the licence and copyright notice retained. That is permissive enough for the notebooks to be adapted into internal training material or used as a starting point for your own implementations. Nothing in the README suggests any additional restriction, and this is a description of the licence text rather than legal advice; if you plan to redistribute modified versions commercially, read the LICENSE file at the repository root.

The last push was on 2026-07-08, and the repository is not archived. That tells you the tree is still being touched, but it does not tell you the notebooks have been modernized, because the README still states PyTorch v0.4. The upgrade cost of adopting this code is therefore front-loaded: you pay it once, when you decide whether to run the notebooks as-is in an older environment or port them forward. Porting is a per-notebook task, not a single migration, since each notebook imports PyTorch directly and there is no shared compatibility layer to fix in one place. Budget for that before you start.

Editorial conclusion

Adopt this repository if you are working through the Deep Reinforcement Learning Nanodegree or want the algorithm notebooks as a reference implementation to read alongside a textbook. Do not adopt it as a dependency: there is no package to install, the notebooks pin PyTorch v0.4 and Python 3, and the README marks PPO and the CarRacing DQN as coming soon. Before relying on any notebook, check whether the environment it targets still installs on your machine, and read the Dependencies section of the README in full rather than assuming the setup is current.

Frequently asked questions

What is the udacity/deep-reinforcement-learning repository?

It is the material repository for Udacity's Deep Reinforcement Learning Nanodegree, containing PyTorch tutorial notebooks and Unity ML-Agents project folders. The README describes it as material related to the program rather than a standalone library.

How do I run a deep reinforcement learning example from this repository?

The README's Dependencies section says to create and activate a new Python environment, and the repository has a python/ directory for the dependency files. After that you install the requirements and launch Jupyter from the repository root so relative paths resolve.

Is DQN deep reinforcement learning in udacity/deep-reinforcement-learning?

Yes. The repository has a dqn directory, and the README lists LunarLander-v2 with Deep Q-Networks as solved in 1504 episodes. CarRacing-v0 with DQN is marked coming soon.

Is PPO deep reinforcement learning in udacity/deep-reinforcement-learning?

No. The README lists Proximal Policy Optimization in the tutorials table with the note coming soon, so it is not present as a finished notebook.

What is deep reinforcement learning in AI, and does this repository cover it?

The repository covers it through tutorials that pair an algorithm with an environment, such as DQN on LunarLander-v2 and the Cross-Entropy Method on MountainCarContinuous-v0. The README states the tutorial code is in PyTorch v0.4 and Python 3.

What is the difference between deep reinforcement learning and reinforcement learning in this repository?

The repository covers both. The dynamic-programming, Monte Carlo, temporal-difference, discretization and tile-coding tutorials work with tabular or discretized methods, while the dqn, ddpg-pendulum, ddpg-bipedal and reinforce notebooks use neural networks.

Official sources

  1. Issues
  2. License: MIT
  3. Project website
  4. README
  5. udacity/deep-reinforcement-learning on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/udacity-deep-reinforcement-learning.svg)](https://hysenlabs.com/projects/udacity-deep-reinforcement-learning)