Library / SDK
kevinhughes27/TensorKart avatar
kevinhughes27/TensorKart

TensorKart: training a TensorFlow agent to drive Mario Kart 64

self-driving MarioKart with TensorFlow

1,574 stars249 forksPythonMIT

At a glance

What is it?
TensorKart records joystick input and emulator screenshots, trains a convolutional model on the pairs, and replays the result through gym-mupen64plus. It is a small, readable imitation-learning project, not a turnkey racing bot.
Who is it for?
TensorKart fits engineers who want a compact, end-to-end imitation-learning example they can read in an afternoon and adapt to their own capture setup. It does not fit anyone expecting a packaged racing agent: the README gives no install instructions for gym-mupen64plus, no sample data, and no rollback path, and the pinned requirements (tensorflow==2.5.0, numpy==1.19.5, matplotlib==1.5.2) will need attention on a current Python.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 59 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What TensorKart actually solves, and for whom

Supervised imitation learning is the cheapest way to get a driving agent into a game, and TensorKart is a compact demonstration of that pipeline. You drive Mario Kart 64 yourself, the project captures what the screen looked like and what the joystick was doing at the same moment, and a TensorFlow model learns the mapping from image to control. The README frames it as "self-driving MarioKart with TensorFlow", and the whole thing lives in four Python scripts at the repository root: record.py, utils.py, train.py and play.py.

The intended user is someone with a working Nintendo 64 emulator setup who wants a readable end-to-end example rather than a framework. There is no package to install from PyPI, no CLI, and no configuration file. The README's training note is modest: the model was trained with 4 races on Luigi Raceway, 2 races on Kalimari Desert and 2 races on Mario Raceway, and it sometimes generalizes to a track it has never seen, such as Royal Raceway. That generalization claim is the interesting part, and it is also the part the README leaves unquantified. If you need a benchmark, this project does not offer one.

The record, prepare, train, play data flow

The mechanism is a four-stage loop over one data directory per recording session. record.py captures the emulator window from the top left corner of the screen and writes screenshots alongside a data.csv of joystick readings. utils.py then does two different jobs depending on the subcommand: `viewer` displays a sample directory, and `prepare` walks one or more sample directories and builds an X matrix of images and a y matrix of joystick outputs.

The shape of y is documented explicitly in the README: index 0 is the joystick x axis, 1 is the joystick y axis, 2 is button a, 3 is button b, and 4 is button rb. Five numbers, one per frame. train.py fits a model to those pairs with TensorFlow and cuDNN for GPU acceleration, and the README states that the best model across all epochs is the one written to disk. play.py inverts the flow: the gym-mupen64plus environment supplies screenshots, the model returns a joystick command, and the command goes back to the emulator. Holding the LB button on the controller overrides the model, which is the only manual escape hatch described.

Two capture details matter more than they look. The GUI stops updating while recording, per the README, to avoid slowdowns, so you cannot watch the preview to confirm framing mid-run. And the README warns that sometimes the captured screenshot is the desktop instead of the game, which means you have to inspect the data and delete the offending rows from data.csv by hand. Data quality here is a manual, eyes-on process.

Installing TensorKart and recording your first race

There is no packaging step. The README says to install Python and pip, install the Python dependencies from requirements.txt, and install mupen64plus via apt-get. The pinned versions are old (tensorflow==2.5.0, numpy==1.19.5, scikit-image==0.12.3, matplotlib==1.5.2), so a modern Python interpreter is likely to reject some of them; expect to resolve that before anything runs.

bash
pip install -r requirements.txt
sudo apt-get install mupen64plus

With those in place, start the emulator, load Mario Kart 64, and confirm that mupen64plus is using the sdl input plugin and that your joystick is connected. Then run the recorder.

bash
python record.py

The README's sequence is specific: make sure the graph responds to joystick input, position the emulator window so the image is captured by the program at the top left corner, press record, and play through a level. Afterwards you can trim images off the front and back of the run by deleting lines from data.csv. Once you have a directory of samples, you can look at them before committing to a training run.

bash
python utils.py viewer samples/luigi_raceway

When the samples look right, build the training matrices. The README gives `samples/*` as the argument, and notes that zsh expands the glob or that you can pass it directly.

bash
python utils.py prepare samples/*

That command should leave you with an X array of images and a y array of joystick outputs. Training is then a single script, and the README puts the runtime at roughly an hour depending on data volume and hardware.

Where TensorKart breaks down

The failure modes are mostly about data and about the gap between the two halves of the pipeline. Recording is brittle by design: the capture region is fixed to the top left corner of the screen, so any window movement, notification, or resolution change silently poisons the dataset, and the README's own advice is to check for desktop screenshots and delete those rows from data.csv. Nothing in the described workflow validates that an image actually contains game pixels.

The generalization story is honest but thin. The README says the model is "sometimes able" to handle a new track after training on three others, and gives no metric, no per-track breakdown, and no evaluation script. There is no held-out test split described, so you cannot tell from the repository alone whether a given training run is better or worse than the last one except by watching play.py drive.

play.py also depends on gym-mupen64plus, which is not in requirements.txt. The README links to the bzier/gym-mupen64plus repository but does not describe installing it. That is the single largest gap between the documented workflow and a working setup, and it is worth confirming before you invest an hour in recording. Finally, the reward signal mentioned under future work is `-1` per time-step, described as a baseline for reinforcement learning that has not been added. If your goal is an agent that improves itself, this project is not there yet; it is pure imitation.

TensorKart against Xbox Game AI and Donkey Gym

The README lists its own neighbours, and the differences are structural rather than cosmetic. Xbox Game AI uses PYXInput for direct control of any Xbox or PC game, which sidesteps screenshot capture entirely: it reads and writes game input rather than watching pixels. TensorKart is the opposite approach, learning from images, which makes it more general in principle and far more sensitive to capture quality in practice.

Donkey Gym, the OpenAI Gym environments for the Donkey Car, is the closest analogue in spirit. Both wrap a driving task as a Gym-style environment, but Donkey Car targets a physical RC car with its own hardware and data-collection tooling, while TensorKart targets an emulator window on your desktop. If you want a simulator with a documented environment API and a reward signal, AirSim is a different order of magnitude: an Unreal Engine simulator for autonomous vehicles, not a script that watches mupen64plus. SerpentAI sits at another level again, a general Game Agent Framework for building agents across games. TensorKart's advantage over all of them is size. Four scripts, one requirements file, and a data format you can inspect in a spreadsheet.

Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-08-03. The licence is MIT, declared in LICENSE.txt at the repository root. MIT is permissive: it allows use, modification and redistribution with the copyright notice and permission notice retained. That is the extent of what the repository states; questions about your own distribution obligations belong with a lawyer, not with this article.

The upgrade cost is concentrated in requirements.txt. The pins date from the TensorFlow 2.5 era, and several of them (matplotlib==1.5.2, Pillow==4.1.0, scikit-image==0.12.3, h5py==2.7.0) are far behind current releases. There are no releases in the repository, so there is no versioned upgrade path and no changelog to read: you track master. The README does not document rollback, does not describe a test suite, and does not mention CI. In practice, adopting TensorKart means adopting the dependency set as a starting point and doing your own environment work, which is a real cost even though the code itself is small enough to read completely.

Editorial conclusion

TensorKart fits engineers who want a compact, end-to-end imitation-learning example they can read in an afternoon and adapt to their own capture setup. It does not fit anyone expecting a packaged racing agent: the README gives no install instructions for gym-mupen64plus, no sample data, and no rollback path, and the pinned requirements (tensorflow==2.5.0, numpy==1.19.5, matplotlib==1.5.2) will need attention on a current Python. Before committing, verify that you can get mupen64plus running with the sdl input plugin, that record.py produces a data.csv whose rows match real screenshots rather than your desktop, and that utils.py prepare emits the X and y arrays train.py expects.

Frequently asked questions

What is TensorKart?

TensorKart is a Python project that trains a TensorFlow model to drive Mario Kart 64 by imitating recorded human play. It captures emulator screenshots and joystick input with record.py, builds training matrices with utils.py prepare, trains with train.py, and drives with play.py through the gym-mupen64plus environment.

How do I install TensorKart?

The README says to install Python and pip, run `pip install -r requirements.txt`, and install mupen64plus via apt-get. There is no package on PyPI and no installer; play.py additionally needs gym-mupen64plus, which the README links to but does not include in requirements.txt.

Why does TensorKart capture my desktop instead of the game?

The recorder captures a fixed region starting at the top left corner of the screen, so the emulator window has to be positioned there and stay there. The README notes that sometimes the screenshot is the desktop instead of the game and says to remove the corresponding lines from data.csv.

How long does training a TensorKart model take?

The README states that training can take around an hour depending on how much data you are training with and your system specs. train.py saves the best model from all epochs when it finishes.

Official sources

  1. Issues
  2. kevinhughes27/TensorKart on GitHub
  3. License: MIT
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/kevinhughes27-tensorkart.svg)](https://hysenlabs.com/projects/kevinhughes27-tensorkart)