keon/deep-q-learning: DQN and Double DQN in Under 100 Lines of Keras
Minimal Deep Q Learning (DQN & DDQN) implementations in Keras
At a glance
- What is it?
- A minimal reference implementation of Deep Q Learning and Double Deep Q Learning on Gymnasium, written in Keras and small enough to read in one sitting. The repository is a teaching artifact, not a training framework, and the unstable training in dqn.py is the point the Double DQN file exists to make.
- Who is it for?
- Adopt keon/deep-q-learning if you need to read a complete DQN loop end to end, or if you want a baseline to diff against a PyTorch reimplementation. Do not adopt it for production training, for environments with continuous action spaces, or for any pipeline that needs checkpointing beyond weights and epsilon.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 136 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What keon/deep-q-learning actually solves
Most DQN repositories are frameworks. They ship a replay buffer abstraction, a trainer class, a config system and a logging layer, and by the time you find the Bellman update it is wrapped in four levels of indirection. keon/deep-q-learning takes the opposite position. The README says the implementation is "Minimal and Simple Deep Q Learning Implemenation in Keras and Gym. Under 100 lines of code!" and the file list backs that up: dqn.py, ddqn.py, dqn_batch.py, requirements.txt. There is no package, no setup.py, no CLI. You clone it and run a file.
The audience is someone who has read the Deep Q Learning paper and now wants to see the pieces in working Python. Experience replay, a target network, epsilon-greedy action selection, a Huber or MSE loss on the TD error: all of it is visible in a single scroll. The repository is also the companion code for a blog article at keon.kim, which the README links directly for the explanation of dqn.py. That pairing matters, because the code is deliberately under-commented in places where the article carries the reasoning.
It is not for someone who wants to train an agent on a custom environment and ship it. There is no evaluation harness, no vectorized environment support, and no experiment tracking. The three files are three different answers to the same question, and you pick one.
How the agent loop and replay memory are wired
The mechanism is the standard DQN loop, kept flat. The agent holds an online network that maps observations to Q-values for each discrete action, and a target network that produces the bootstrap value in the TD target. Actions come from an epsilon-greedy policy over the online network's output. Transitions of the form (state, action, reward, next state, done) go into a replay buffer, and training samples random minibatches from that buffer rather than learning from the most recent step.
Two implementation choices are worth calling out because the README names them explicitly. First, the memory is a deque rather than a list, which the README explains is "in order to limit the maximum number of elements in the memory." That single change gives the buffer a fixed capacity and makes old transitions fall off the end automatically, instead of growing without bound until the process runs out of RAM. Second, the repository adds save() and load() helpers, which the README describes as minor tweaks made for convenience. The 2026-05 changelog notes that these now persist epsilon alongside the weights, so a resumed run keeps its exploration schedule rather than restarting at the initial epsilon.
The 2026-05 changelog also records the migration from gym to gymnasium with the modern reset/step API, and the move to Keras 3 idioms: Adam(learning_rate=), an explicit Input layer, and keras.losses.Huber. The same changelog states that ddqn.py now implements actual Double DQN, where the online network selects the action and the target network evaluates it. If you have seen Double DQN implementations that only swap in a second network without splitting selection from evaluation, this is the distinction being corrected.
Installing it and running a first training loop
There is no package on an index and no console script. The repository gives a requirements.txt listing four dependencies: numpy, keras, gymnasium and tensorflow. Install them, then run one of the agent files from the repository root. The requirements file pins nothing, so you are installing whatever the current releases are; the 2026-05 changelog assumes Keras 3, which is the version that introduced the Input layer and the learning_rate keyword used in the code.
pip install -r requirements.txtAfter that, the entry point is the script itself. Running dqn.py executes the training loop against CartPole, which is the environment the README names when it discusses hyperparameter tuning. Expect a stream of per-episode output; the README does not document the exact print format, so treat what appears as whatever the script emits rather than a fixed contract.
python dqn.pyFor the Double DQN variant, run the other file. The README states that training might be unstable for dqn.py and that this problem is mitigated in ddqn.py, so if your first run of dqn.py fails to improve, that is a documented outcome rather than a bug you introduced. The 2026-05 changelog also notes that CartPole-specific reward shaping was removed from ddqn.py so it is environment-agnostic, which means the file is the better starting point if you intend to swap in a different Gymnasium environment.
python ddqn.pyThere is a third file, dqn_batch.py, which the 2018-08 changelog describes as splitting batched-update training into its own file. The README does not explain when to prefer it over dqn.py, so read the source before assuming it is a drop-in replacement.
Where this implementation breaks down
The README is candid about the main failure mode: "The training might be unstable for dqn.py." That is not a caveat bolted on at the end, it is the reason ddqn.py exists, and it tells you the plain DQN file is a demonstration of the algorithm rather than a reliable trainer. Overestimation bias in the max operator is the textbook explanation, and Double DQN is the textbook fix, but the practical consequence here is that you should not build a benchmark on dqn.py and then conclude DQN does not work.
The harder limits are structural. The action space must be discrete, because the network emits one Q-value per action and the policy takes an argmax over them; nothing in the repository addresses continuous control. There is no prioritized replay, no n-step returns, no dueling architecture, no noisy networks. The save() and load() helpers persist weights and epsilon, and the README does not document rollback, versioned checkpoints or optimizer state, so resuming a long run is not equivalent to continuing it. And the requirements.txt pins no versions at all, which means a fresh install months from now may resolve to a TensorFlow or Keras release that the code was not written against.
If your goal is to reproduce a published result, this is the wrong tool. If your goal is to understand why the target network exists, it is a good one.
Double DQN here versus a PyTorch rewrite
The most common alternative people reach for is a PyTorch DQN, and the difference is not just the framework. PyTorch implementations of DQN typically arrive as a training script with a config dictionary, an evaluation callback and a checkpoint directory, often derived from the CleanRL or Stable-Baselines3 lineage. They give you logging, seeding and vectorized environments out of the box, and they are what you want if the question is "how well does this agent do" rather than "how does this agent work."
keon/deep-q-learning makes the opposite trade. There is no config layer, so every hyperparameter is a literal in the file and changing one means editing Python. There is no evaluation loop separate from training. There is no seeding discipline documented in the README, so run-to-run variance is something you will observe rather than control. What you get in exchange is that the entire agent fits in one screen, and the Double DQN correction is expressed as a handful of lines you can point at.
A second alternative is to skip the implementation entirely and use a maintained library's DQN. That is the right call when DQN is a means to an end. The reason to prefer this repository is that it is a reference you can hold in your head, and the 2026-05 changelog shows it is being kept current with the modern Gymnasium and Keras 3 APIs rather than frozen at its 2017 state.
Maintenance, licensing and upgrade cost
The repository is not archived and the last push was on 2026-05-18. The changelog shows a real modernization pass at that date rather than a cosmetic commit: gymnasium migration, Keras 3 API updates, a corrected Double DQN, removal of CartPole-specific reward shaping, and epsilon persistence in save() and load(). Before that, the changelog jumps back to 2018-08 for dqn_batch.py and to 2017 for the Keras 2 migration and the hyperparameter work. The pattern is long quiet periods punctuated by a maintenance burst when an upstream API breaks, which is exactly what you would expect from a teaching repository.
The practical upgrade cost sits in requirements.txt. It lists numpy, keras, gymnasium and tensorflow with no version constraints, so an environment built today and an environment built a year from now can differ in ways the code does not guard against. Pinning those four lines yourself is the cheapest insurance, and it is the one change this repository most obviously invites.
The licence is MIT, per the LICENSE file at the repository root. That permits reuse and modification with the copyright notice retained; it says nothing about the licence terms of the dependencies you install alongside it, which carry their own. This is a description of the licence text, not legal advice.
Editorial conclusion
Adopt keon/deep-q-learning if you need to read a complete DQN loop end to end, or if you want a baseline to diff against a PyTorch reimplementation. Do not adopt it for production training, for environments with continuous action spaces, or for any pipeline that needs checkpointing beyond weights and epsilon. Before running anything, confirm that gymnasium, keras and tensorflow resolve to versions compatible with the Keras 3 API used in the 2026-05 changelog, and read ddqn.py rather than dqn.py if you want a run that the README describes as stable.
Frequently asked questions
What is deep Q-learning, and how does keon/deep-q-learning implement it?
Deep Q Learning replaces the Q-table of tabular Q-learning with a neural network that predicts a Q-value for each discrete action. This repository implements that loop in Keras and Gymnasium, with a replay memory stored as a deque, a target network for the bootstrap value, and save/load helpers for weights and epsilon.
What is the difference between dqn.py and ddqn.py in this repository?
The README states that training might be unstable for dqn.py and that this problem is mitigated in ddqn.py. The 2026-05 changelog records that ddqn.py now implements actual Double DQN, where the online network selects the action and the target network evaluates it.
What is the difference between Q-learning and deep Q-learning?
Q-learning stores action values in a table indexed by state and action, which does not scale past small discrete state spaces. Deep Q Learning approximates those values with a neural network, which is what keon/deep-q-learning does using Keras, so the memory holds transitions rather than a table.
What is deep Q-learning in reinforcement learning?
It is a value-based reinforcement learning method that learns a Q-function from experience replay using a neural network approximator. In keon/deep-q-learning the agent stores transitions in a deque-bounded memory, samples minibatches from it, and updates an online network against a target network.
What is the deep Q-learning algorithm?
The loop is: act epsilon-greedily on the online network's Q-values, store the transition, sample a minibatch from replay memory, and regress the online network toward the reward plus the discounted target network value of the next state. The 2026-05 changelog notes that ddqn.py splits selection and evaluation between the two networks.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/keon-deep-q-learning)