lucidrains/q-transformer: Autoregressive Q-Learning in Python
Implementation of Q-Transformer, Scalable Offline Reinforcement Learning via Autoregressive Q-Functions, out of Google Deepmind
At a glance
- What is it?
- A Python implementation of the Q-Transformer paper from Google Deepmind, with an attention model, a Q-learner and a replay-memory pipeline. The install is one pip command; the environment is yours to write.
- Who is it for?
- Adopt it if you already have a robot or simulator environment and want to compare a single-action Q-learner against the autoregressive multi-action formulation without reimplementing the paper. Do not adopt it if you need a working policy out of the box: the README states you must supply your own environment by overriding BaseEnvironment, and the bundled MockEnvironment returns random tensors, not a task.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 49 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What the Q-Transformer implementation is for
This repository packages the method from Q-Transformer: Scalable Offline Reinforcement Learning via Autoregressive Q-Functions, the Google Deepmind paper cited in the README. The paper's problem is that a Q-function over a multi-dimensional action space grows combinatorially: a robot arm with several joints and a discretized bin per joint has an action space that a single regression head handles badly. The repository's answer is to factorize the Q-function autoregressively, generating one action dimension at a time and conditioning each on the ones already chosen, with the reward attached only after the last action. The README's todo list marks that main proposal as done.
The intended reader is someone doing offline reinforcement learning research who wants the architecture and the training loop in readable Python rather than a full training stack. The README is explicit that the single-action Q-learning logic is kept alongside the autoregressive version, both for comparison and, in the author's words, to serve as education for himself and the public. That framing matters: this is a reference implementation with a training entry point, not a robotics product.
How the autoregressive Q-function is wired
Three pieces do the work. QRoboticTransformer is the attention model. It takes a vision-transformer configuration for the image tower, a num_actions count, an action_bins count for the discretization, and transformer parameters for the action-decoding side. The README example passes depth = 1, heads = 8, dim_head = 64, cond_drop_prob = 0.2 and dueling = True, so a deep dueling head is available as a flag rather than a separate class.
Agent drives the model through an environment to produce experience. Its constructor takes the model, an environment, num_episodes and max_num_steps_per_episode, and calling it runs the loop. The README defines the environment contract in comments: init() returns instructions and an initial state as a tuple of a string and a tensor of shape state_shape, and calling the environment with actions returns rewards, the next state and a done flag. QLearner then trains on a ReplayMemoryDataset with num_train_steps, learning_rate, batch_size and grad_accum_every.
Inference is a single call. model.get_optimal_actions(video, instructions) takes a video tensor shaped (batch, channels, frames, height, width) and a list of instruction strings, and returns actions. The README's example uses a tensor of shape (2, 3, 6, 224, 224) with two natural-language instructions, which shows the intended input is a short video clip plus language, not a single frame.
Installing q-transformer and running the README example
The README gives one install line. It pulls the package from PyPI under the name q-transformer, while the import name is q_transformer.
pip install q-transformerAfter that, the README's usage example is the fastest way to see the shapes the model expects. The following block is adapted from it: it builds the model, attaches a mock environment, collects episodes and then trains. The mock environment is imported from q_transformer.mocks and exists so the API can be exercised without a simulator.
import torch
from q_transformer import (
QRoboticTransformer,
QLearner,
Agent,
ReplayMemoryDataset
)
model = QRoboticTransformer(
vit = dict(
num_classes = 1000,
dim_conv_stem = 64,
dim = 64,
dim_head = 64,
depth = (2, 2, 5, 2),
window_size = 7,
mbconv_expansion_rate = 4,
mbconv_shrinkage_rate = 0.25,
dropout = 0.1
),
num_actions = 8,
action_bins = 256,
depth = 1,
heads = 8,
dim_head = 64,
cond_drop_prob = 0.2,
dueling = True
)The next block wires the environment and the two loops. The README notes that you need to supply your own environment by overriding BaseEnvironment; MockEnvironment is only a stand-in with a state shape of (3, 6, 224, 224) and a text embedding shape of (768,).
from q_transformer.mocks import MockEnvironment
env = MockEnvironment(
state_shape = (3, 6, 224, 224),
text_embed_shape = (768,)
)
agent = Agent(
model,
environment = env,
num_episodes = 1000,
max_num_steps_per_episode = 100,
)
agent()
q_learner = QLearner(
model,
dataset = ReplayMemoryDataset(),
num_train_steps = 10000,
learning_rate = 3e-4,
batch_size = 4,
grad_accum_every = 16,
)
q_learner()What you should see is the agent loop filling a replay memory and the learner stepping the model. The README adds, after this point, that your robot should be better at selecting optimal actions, which is a statement about the intended outcome rather than a measured result. Finally, action selection at inference is one call:
video = torch.randn(2, 3, 6, 224, 224)
instructions = [
'bring me that apple sitting on the table',
'please pass the butter'
]
actions = model.get_optimal_actions(video, instructions)The environment contract is where most adopters will stall
The repository ships train_cartpole.py at the top level, which suggests a small classical-control example is the intended first target. Beyond that, the burden of connecting a real simulator falls entirely on the user. The README's comment block is the whole specification: init() returns instructions and the initial state, and the call operator returns rewards, next state and done. There is no adapter for Gym, no dataset format for an existing offline corpus, and no description of what a well-formed instruction string looks like beyond the two examples.
The MockEnvironment shapes are a second constraint worth reading carefully. state_shape (3, 6, 224, 224) means three channels, six frames and 224 by 224 pixels, and text_embed_shape (768,) means the language side expects precomputed embeddings of width 768 rather than raw text. If your observations are low-dimensional vectors, you are not using this model as designed; the vision transformer tower in the configuration above is doing the perception work.
There is also a stated open problem. The todo list asks for consultation with RL experts on delusional bias, citing a 2020 paper on the subject, and lists randomized action ordering and a beam search function for optimal actions as unfinished. The README does not document rollback, checkpointing or resuming a training run, so plan for that yourself.
lucidrains/q-transformer against Decision Transformer
The related searches for this project surface Decision Transformer, and the contrast is real. Decision Transformer casts offline RL as sequence modeling over returns, states and actions, and predicts actions conditioned on a target return. There is no Q-function and no Bellman backup; the objective is supervised next-token prediction. Q-Transformer keeps the Q-learning objective and makes the Q-function itself autoregressive over action dimensions, with the reward applied only at the last action, as the README's todo entry describes.
The practical difference is what you need to supply. Decision Transformer style code typically consumes a fixed offline dataset of trajectories. This repository's pipeline is built around an Agent that interacts with an environment to create that dataset, then a QLearner that trains on it. If you have a large static dataset and no simulator, the Decision Transformer framing matches your situation more directly. If you have a simulator and care about how the Q-function scales across action dimensions, this is the relevant code.
Maintenance, licence and upgrade cost
The last push to the default branch was on 2026-08-11, which is recent enough that the code is not stale. That is a statement about timing, not a guarantee of support. The release history shows 0.4.5 on 2025-05-20, preceded by 0.4.3 and 0.4.2 earlier that year, so releases are infrequent and the version numbers move in small increments.
The licence is MIT, which is permissive and places few obligations on how you redistribute or modify the code. This is not legal advice; read the LICENSE file in the repository if your organisation has specific requirements. The upgrade cost is dominated by the model constructor, not by dependency churn: the QRoboticTransformer signature carries nested configuration dictionaries for the vision tower, so a change to a key name inside vit or to the dueling flag is the kind of edit that will break a pinned call site. Pinning the version and reading the diff of the constructor between releases is cheaper than tracking the default branch.
Editorial conclusion
Adopt it if you already have a robot or simulator environment and want to compare a single-action Q-learner against the autoregressive multi-action formulation without reimplementing the paper. Do not adopt it if you need a working policy out of the box: the README states you must supply your own environment by overriding BaseEnvironment, and the bundled MockEnvironment returns random tensors, not a task. Before committing, check that your environment's state and text-embedding shapes match what you pass to MockEnvironment, because the model's vision tower expects a fixed image shape, and check the tests directory for which components have coverage.
Frequently asked questions
Can you give me an example of Q-learning?
The README's usage example is one: build a QRoboticTransformer, wrap it in a QLearner with a ReplayMemoryDataset, and call the learner to run training steps against the replay memory. The repository also keeps single-action Q-learning logic alongside the autoregressive version for comparison.
What is the definition of the q-function?
In this repository the Q-function is the model that QRoboticTransformer implements, and in the paper it cites the function is factorized autoregressively over action dimensions with the reward applied only on the last action. The README's todo list records that formulation as completed.
What are the disadvantages of Q-learning?
The repository points at one directly: the todo list asks for consultation with RL experts on delusional bias, citing a 2020 paper. The README also links to a blog post titled Q Learning is Not Yet Scalable, which is a claim about the method rather than about this implementation.
Who proposed Q-learning?
The README does not attribute the original Q-learning algorithm. It attributes the autoregressive Q-function formulation to the Q-Transformer paper out of Google Deepmind, cited in the repository.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lucidrains-q-transformer)