Model or dataset
commaai/research avatar
commaai/research

commaai/research: the 2016 comma.ai driving dataset and its two training scripts

GitHub describes it as dataset and code for 2016 paper "Learning a Driving Simulator". The repository metadata lists Python as its primary language. The metadata lists the BSD-3-Clause license. This article stays within the project description and details documented in the GitHub repository README.

4,125 stars1,158 forksPythonBSD-3-Clause

At a glance

What is it?
commaai/research is the dataset and code release for the 2016 paper Learning a Driving Simulator. It ships 7.25 hours of highway video with aligned sensor logs, plus two Keras experiments, and it pins you to TensorFlow 0.9 and Keras 1.0.6.
Who is it for?
Adopt commaai/research if you want a small, fully documented driving dataset with a known sensor layout and you are willing to rebuild a TensorFlow 0.9 and Keras 1.0.6 environment for the training scripts, or to write your own loader against the HDF5 files. Do not adopt it if you need city driving, night or rain footage, current deep learning frameworks, or any commercial use: the dataset is CC BY-NC-SA 3.0 while the repository code is BSD-3-Clause.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Probably not. The repository last received commits 50 months ago, on August 16, 2022.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What the comma.ai driving dataset actually contains

The repository exists to support one paper, Learning a Driving Simulator, and one artefact: 7 and a quarter hours of largely highway driving recorded on an Acura ILX 2016. A camera sat on the windshield and captured video at 20 Hz. In parallel the car logged speed, acceleration, steering angle, GPS coordinates and gyroscope angles. Those measurements were resampled onto a uniform 100 Hz time base, so the video and the sensor channels do not share a clock in the raw files and have to be joined. The README lists 11 clips under driver names (dog, emily, frodo) with per-clip sizes from 2.7 GB to 13 GB, 45 GB compressed and 80 GB uncompressed.

The intended audience is narrow. This is not a perception benchmark with bounding boxes or a planning benchmark with route goals. It is a steering-angle regression corpus plus a generative image modelling corpus, which is exactly what the two experiment write-ups cover. If your question is "can a network learn to predict steering from a windshield image", the data answers it. If your question is "can a network handle a roundabout in the rain", the data does not contain the question.

How the HDF5 files and cam1_ptr alignment work

Everything is stored in HDF5, and files are named after the time they were recorded. The camera datasets have shape number_frames x 3 x 160 x 320 and dtype uint8, so frames arrive as small RGB images, not full-resolution video. The log datasets hold the 100 Hz measurements. The README singles out one log dataset, cam1_ptr, which addresses the alignment between camera frames and the other measurements.

That single pointer is the whole data model, and it is worth being blunt about the consequences. The camera runs at 20 Hz and the logs at 100 Hz, so five log samples fall between consecutive frames. Any training script has to decide what to do with the four it discards, and the repository does not document that decision beyond providing the pointer. The repository layout shows dask_generator.py, which suggests the training scripts stream batches from disk rather than loading an 8 GB clip into memory, and train_steering_model.py, train_generative_model.py, view_steering_model.py and view_generative_model.py, which map onto the two experiments in SelfSteering.md and DriveSim.md. There is also a ros/ directory, and the README does not explain what it is for. If you need ROS bag output, read the directory before assuming it is a maintained conversion path.

Installing the requirements and running the steering experiment

The README points at Anaconda for the Python environment and lists four requirements: tensorflow-0.9, keras-1.0.6, cv2 (linked to the menpo opencv3 conda channel), and Anaconda itself. Those versions are from 2016. TensorFlow 0.9 predates the 1.0 API and Keras 1.0.6 predates the Keras 2 API, so the training scripts will not run against a modern pip install. Expect to build an isolated environment, or to port the model definitions yourself.

The dataset download is a shell script at the repository root. It fetches the clips described in the README, or you can pull the same data from the archive.org comma dataset page.

bash
./get_data.sh

After it finishes you should have a dataset directory containing camera and log subdirectories, each holding folders named after recording timestamps. The README shows the shape of that tree explicitly:

bash
+-- dataset
|   +-- camera
|   |   +-- 2016-04-21--14-48-08
|   |   ...
|   +-- log
|   |   +-- 2016-04-21--14-48-08
|   |   ...

With the data in place, the steering experiment runs through train_steering_model.py, and SelfSteering.md is the write-up to read alongside it. The generative experiment runs through train_generative_model.py with DriveSim.md as its companion. The two view_* scripts are for inspecting what a trained model produces. The README does not document command-line flags for any of these scripts, so read the files themselves before running them, and check whether they expect a GPU or will fall back to CPU.

Where this dataset stops being the right tool

The licence is the first hard boundary. The README states that the dataset is copyrighted by comma.ai and published under Creative Commons Attribution-NonCommercial-ShareAlike 3.0, which means attribution is required, commercial use is not permitted, and derivative works must carry the same licence. The repository code is BSD-3-Clause. That split is easy to miss: you can read and reuse the training scripts under a permissive licence while the data they consume forbids commercial deployment. Anyone building a product on this corpus has a licensing problem, not a technical one.

The second boundary is scope. The README describes the footage as largely highway driving. There is no city traffic, no pedestrians, no construction zones, and nothing said about weather or night conditions. A model trained here learns lane-following on open road. Evaluating it on anything else tells you about the distribution shift, not about the model.

The third is maintenance. The repository is not archived, but no release has been cut and the environment pins are from 2016. Treat it as a frozen artefact tied to a paper, not as a library that will track upstream TensorFlow or Keras. The README ends with a hiring note inviting people to show off work on the dataset, which tells you how it was positioned: a recruiting and research release, not a supported product.

commaai/research against a modern driving dataset

The obvious alternative is a current autonomous driving dataset such as nuScenes or the Berkeley DeepDrive set, which ship multi-camera rigs, lidar, and annotations for detection and tracking. The difference is not just size. Those datasets are built for perception tasks with labelled objects and a held-out evaluation server. commaai/research is built for end-to-end imitation: one forward camera, one steering signal, and a generative model that predicts the next frame. There is no annotation layer to evaluate against, so there is no leaderboard and no standard metric.

That makes the comparison a question of intent rather than quality. If you want to train a detector, the annotated sets are the right input and this repository is the wrong one. If you want the smallest possible end-to-end driving setup that fits on a single machine and comes with a paper describing the experiment, this is close to the minimum viable version of that. The 160 x 320 frame size is a deliberate part of that: it keeps a full training run within reach of one GPU from 2016.

Licence, upgrade cost and what you are signing up for

Two licences apply. The repository code is BSD-3-Clause, which permits commercial use with attribution and without copyleft. The dataset is CC BY-NC-SA 3.0, which does not permit commercial use and requires share-alike on derivatives. Mixing the two in one project means your code can be proprietary while your trained weights, if they are considered a derivative of the data, may not be. That is a question for a lawyer, not for this article, but it is the first thing to resolve before any product work.

The upgrade cost is real and it is front-loaded. tensorflow-0.9 and keras-1.0.6 are not installable from a current package index without pinning an old Python, and the menpo opencv3 conda channel referenced in the README is a 2016-era channel. The practical paths are to containerise an old environment, or to port the model definitions to a current framework and keep only the data loading logic. The second path is more work but leaves you with something you can maintain. The data itself does not rot: HDF5 files from 2016 still open today, and the format is documented in the README.

Editorial conclusion

Adopt commaai/research if you want a small, fully documented driving dataset with a known sensor layout and you are willing to rebuild a TensorFlow 0.9 and Keras 1.0.6 environment for the training scripts, or to write your own loader against the HDF5 files. Do not adopt it if you need city driving, night or rain footage, current deep learning frameworks, or any commercial use: the dataset is CC BY-NC-SA 3.0 while the repository code is BSD-3-Clause. Before committing, run ./get_data.sh on a machine with 80 GB free, open camera/2016-04-21--14-48-08 and log/2016-04-21--14-48-08 with an HDF5 reader, and confirm that cam1_ptr lines the frames up with the 100 Hz measurements the way your pipeline expects.

Frequently asked questions

What is the size of the commaai/research driving dataset?

The README states it is 45 GB compressed and 80 GB uncompressed, split across 11 clips of variable size, from 2.7 GB to 13 GB each.

What licence covers the commaai/research dataset?

The README states the dataset is copyrighted by comma.ai and published under Creative Commons Attribution-NonCommercial-ShareAlike 3.0, which requires attribution, forbids commercial use and requires derivative works to carry the same licence. The repository code is BSD-3-Clause.

Which frameworks does commaai/research require?

The README lists anaconda, tensorflow-0.9, keras-1.0.6 and cv2 from the menpo opencv3 conda channel. Those are 2016-era versions and predate the current TensorFlow and Keras APIs.

Official sources

  1. Official README
  2. Project repository