Open-source project
nv-tlabs/DriveGAN_code avatar
nv-tlabs/DriveGAN_code

nv-tlabs/DriveGAN_code: two training stages, one Drive folder, four borrowed codebases

Code release for DriveGAN (CVPR 2021)

100 stars16 forksCSSNOASSERTION

At a glance

What is it?
The PyTorch release for DriveGAN, a learned driving simulator from a CVPR 2021 oral paper, trained on unannotated video and action pairs rather than on a hand-built engine. The pipeline is a VAE-GAN stage feeding a dynamics engine, with a browser front end on top, and the README is organised around the checkpoints you can skip downloading.
Who is it for?
Use this code if you want to study or extend a learned driving simulator from that paper, and if you can source the checkpoints rather than train from scratch, since the README makes every stage downloadable. Three things to weigh first.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Probably not. The repository last received commits 59 months ago, on November 11, 2021.
What is it written in?
Mainly CSS, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A VAE-GAN stage feeds a dynamics engine

Training is split into two stages and the split is not arbitrary. Stage one is a VAE-GAN that learns a latent representation of driving frames, and once its validation loss converges you encode the dataset with the learned model. Stage two is the dynamics engine, trained on that encoded data rather than on pixels. That ordering is what makes the eventual simulator differentable enough to re-simulate a sequence, since the dynamics engine works in the latent space stage one produced. Each stage has a shell script, one for training and one for encoding, and both take a comma-separated list of data directories rather than a single path. The README also makes the shortcut explicit: if you did not train stage one, go to the stage two section and download the encoded data instead.

Your data directories have to be renamed first

Stage one expects a naming pattern that does not match what the archives extract to, and this is the step most likely to waste your afternoon. The README tells you to download the numbered archives, extract them, and notes that the extracted directories have names starting with 6405, which you must change to data1 through data6.

bash
cd DriveGAN_code/latent_decoder_model
mkdir img_data && cd img_data
tar -xvzf {0-5}.tar.gz
mv 6405x data{1-6}

The rename uses a glob, so every directory with that prefix is renamed in order. The training invocation then lists all six explicitly, comma separated, so the same convention appears twice: once in the filesystem and once in the command line.

bash
./scripts/train.sh ./img_data/data1,./img_data/data2,./img_data/data3,./img_data/data4,./img_data/data5,./img_data/data6

Progress is monitored with tensorboard against the log directory named inside the training script, which means the script itself is where you configure the output location.

Every artefact comes from one shared Drive folder

There is a single file sharing link in the README, and it serves every artefact you might want. The numbered image archives for stage one come from it, and so do the encoded data archive and the VAE-GAN checkpoint for people skipping stage one, and the dynamics engine checkpoint and the simulated checkpoint for people skipping training altogether. That arrangement is generous, and it comes with a caveat worth stating plainly: nothing in the repository publishes a checksum for those files, so you are trusting a shared folder rather than verifying an artefact. The file names are also meaningful, with iteration numbers embedded in them, which tells you these are specific training states rather than the latest of anything.

The dataset is non-commercial and the code is not

The licensing is split across artefacts, and the split matters if you are doing anything commercial. The codebase and the trained models are distributed under the NVIDIA source code licence. The dataset is a separate matter entirely: it is derived from the Carla simulator and is distributed under a Creative Commons Attribution-NonCommercial 4.0 International licence. So a non-commercial licence governs the data you would train on or evaluate with, while the code and the model weights sit under a different, proprietary-adjacent source licence. Nobody in this repository has stated that the trained models inherit the data's non-commercial terms, and if that matters to your use, it is a question to raise rather than to assume either way.

PyTorch 1.7.1 and Python 3.6.9 are what was tested

The environment section names one machine configuration and hedges it. The codebase is tested with Ubuntu 18.04 and Python 3.6.9, and the README says it most likely works with other close Python 3 versions. The requirements file is mostly unpinned, with things like termcolor, opencv-python, moviepy, scikit-image, IPython, tornado, simplejson and tensorboard, and then two exact pins.

code
torch==1.7.1
torchvision==0.8.2

Those two lines are the real constraint. A PyTorch release from that era will not install against a current Python on many systems without a specific older build, and the README offers no container or environment file to sidestep it. So the first thing to budget for in this repository is not compute, it is getting an interpreter that the pinned torch will accept.

Controllability is what the disentanglement buys

The paper's contribution, as the abstract states it, is learning to simulate a dynamic environment directly in pixel space by watching unannotated sequences of frames and their associated action pairs, with no supervision. The payoff is controllability obtained by disentangling components without labels. Beyond steering, that means controls for sampling features of a scene, the abstract names weather and the location of non-player objects as examples. The other property worth noting is differentiability: because the simulator is fully differentiable it supports re-simulation of a given video sequence, so an agent can drive through a recorded scene again while taking different actions. That is the distinction from a video generator, which produces frames without being steerable in the middle of one. The model is trained on multiple datasets including 160 hours of real-world driving data, on top of the Carla-derived set, and the paper was presented as an oral at CVPR in 2021 with pages 5820 to 5829 of the proceedings.

The browser front end is four keys and a refresh

Playing with a trained model means starting a server script with three positional arguments and opening a browser.

bash
./scripts/play/server.sh ${path to saved dynamics engine} ${port e.g. 8888} ${path to saved vae-gan model}

You then navigate to the local port, and the README says it was tested on Chrome. The controls are listed in one line: w speeds up, s slows down, a steers left, d steers right. There are additional buttons for changing scene contents, and sampling a completely new scene is done by refreshing the page, which follows from the model generating the scene rather than playing back a recording. There is also a Colab notebook contributed by someone outside the team, linked from the notebooks directory, for people who would rather not set the stack up at all.

Four borrowed codebases, four licence files

The repository carries four licence files at its root, and each maps to a piece of code that came from somewhere else. The VAE-GAN implementation is adapted from a PyTorch port of StyleGAN2, with its own licence file. The perceptual similarity code is imported from another project, with a second file. The custom StyleGAN operations are imported from NVIDIA's own repository, with a third. And the interactive interface uses a semantic user interface framework, licensed separately again. Alongside those sit the two model directories that mirror the two training stages, a parallel training launcher named for what it does, a distributed training module, a trainer, a configuration file, a data directory, and the Carla assets with a pickle describing how the dataset was split. There is no environment file, no container definition and no lock file in the tree, which is consistent with the requirements file pinning two packages exactly and leaving the rest loose.

Editorial conclusion

Use this code if you want to study or extend a learned driving simulator from that paper, and if you can source the checkpoints rather than train from scratch, since the README makes every stage downloadable. Three things to weigh first. The stack is pinned to a 2020-era PyTorch on Python 3.6, so expect friction on a current machine. The dataset is under a non-commercial licence while the code carries NVIDIA's source licence, so the terms differ by artefact and a commercial use of either needs reading. And the training interface wants your data directories renamed into a fixed pattern before the scripts will accept them.

Frequently asked questions

What is DriveGAN?

A high quality neural simulator that learns to simulate a dynamic environment directly in pixel space from unannotated sequences of frames and their action pairs. It achieves controllability by disentangling components without supervision, and because it is differentiable it supports re-simulation of a video sequence with different actions.

What are the two training stages in DriveGAN_code?

Stage one is a VAE-GAN that learns a latent representation of driving frames, after which the dataset is encoded with that model. Stage two trains the dynamics engine on the encoded data rather than on pixels. Either stage can be skipped by downloading the corresponding checkpoint instead.

Can I skip training and just run DriveGAN_code?

Yes. The README says to download the simulator checkpoint and the VAE-GAN checkpoint, then run the play server script with the dynamics engine path, a port and the VAE-GAN model path, and open that port in a browser.

What are the controls in the DriveGAN web interface?

The keyboard controls are w to speed up, s to slow down, a to steer left and d to steer right. There are additional buttons for changing scene contents, and refreshing the page samples a new scene. The interface was tested on Chrome.

What licence applies to the DriveGAN code and data?

The codebase and trained models are distributed under the NVIDIA source code licence, while the dataset, which is derived from the Carla simulator, is distributed under a Creative Commons Attribution-NonCommercial 4.0 International licence. Four further licence files at the repository root cover code borrowed from other projects.

Official sources

  1. Official README
  2. Project repository