Kaggle docker-python: three v171 images, three commits, one undocumented variant
Kaggle Python docker image
At a glance
- What is it?
- The repository that builds the CPU and GPU images behind Kaggle Notebooks. Its release tags embed a build number and a full commit hash per variant, which makes them precise and, at the same build number, mutually inconsistent. A TPU image is published and the page never mentions TPU.
- Who is it for?
- Adopt the Kaggle docker-python repository if you need to reproduce the environment Kaggle Notebooks actually run in, or if you need a package added to that image and are willing to write a test for it.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Three v171 images built from three different commits
The release tags are the most informative thing in the repository. There are three of them published on 2026-10-01, and each name has three parts: the build number, the variant, and a full forty-character commit hash. `v171-CPU-9154f8c8...`, `v171-GPU-22be155a...` and `v171-TPU-a48bbe96...`.
That scheme is genuinely good practice for images. Because the commit is in the tag, an image tag is a reproducible reference rather than a label, and you can tell exactly which tree a notebook was running.
Read the three hashes together and the picture changes. The three images carrying the same build number 171 were built from three different commits. The GPU image was published at 16:56, the CPU at 16:44 and the TPU at 16:37, so the ordering of publication does not line up with a single rebuild either. And the last push to the branch was on 2026-09-05, four weeks before any of them were published, so all three lag the branch head by the same margin.
For a notebook environment, where a CPU and a GPU kernel are supposed to be the same code with different hardware, that divergence is the thing to raise before pinning.
Dockerfile.tmpl, renderizer/ and four undocumented executables
The README says the repository includes the Dockerfile for building the images. The file on disk is `Dockerfile.tmpl`, a template, and there is a `renderizer/` directory beside it whose job is not stated anywhere in the page. `patches/`, `tools/` and `clean-layer.sh` are likewise present and undocumented.
Four executables sit at the repository root. Two are documented: `build` and `test`. The other two, `diff` and `push`, appear in the file listing and in no section of the README. A `Jenkinsfile` and a `.github/` directory are both present, so CI is configured in two systems.
There is also a `dev.Dockerfile` next to `Dockerfile.tmpl`, which implies a development image distinct from the build image without saying which one `build` produces by default.
The practical consequence is that this repository documents its entry points and not its machinery. If you want to change what goes into the image rather than which packages are in it, you are reading `renderizer/` and `patches/` with no documentation to orient you.
kaggle_requirements.txt is the whole contribution surface
The pull request procedure is five steps and only one of them edits anything. Change `kaggle_requirements.txt`, build the image, add a test for the new package, test the image, open the pull request.
The test step names a concrete example, `tests/test_fastai.py`, which tells you the expected shape of a package contribution: a file per package under `tests/`, which the testing section says is where the suite lives.
That is a small, well-defined contribution surface, and it is worth understanding what it means. Anyone changing the image contents changes the environment of every notebook that runs on it, and the mechanism the project offers for reviewing that change is one line of requirements plus one test file. There is no documented dependency review step, no size or vulnerability check, and no mention of what happens to a package that is later found to conflict with one already in the base image.
For a package that installs cleanly in a notebook on demand, none of that matters. For one that shadows an existing dependency, the pull request is the only gate.
The first step of every package request is install it yourself
The Requesting new packages section opens by telling you not to make a request. It says to evaluate first whether installing the package yourself in your own notebooks suits your needs, and links a guide for missing packages.
That is the right default for a hosted notebook environment, where every user paying for storage and compute shares one image, and where a per-notebook install is cheap and reversible. It also explains why the contribution surface is so thin: the project is trying to keep dependency requests down rather than process them quickly.
The escape route is an issue or a pull request when your own install does not work, which in practice means a package that needs a system library, a compiled extension against the image's exact Python and CUDA versions, or something that must be present before the notebook kernel starts.
The page does not say what the guide contains, so the boundary between the two cases is documented somewhere other than here.
./test takes -p, -i and --gpu
Building is one command with two flags:
./build`--gpu` selects a GPU image and `--use-cache` is described as making iterative builds faster. There is no TPU flag, which the release tags say should exist.
Testing is the same shape:
./testThree flags. `--gpu` to test the GPU image, `--pattern test_keras.py` or `-p test_keras.py` to run a single test, and `--image gcr.io/kaggle-images/python:ci-pretest` or `-i` with the same value to test against a specific image instead of the one you just built.
That last flag is the useful one for anyone reproducing a reported problem: point the suite at a published tag and you find out whether the failure belongs to the image or to your change. Note the example uses the `:ci-pretest` tag, a channel the README otherwise never mentions, and that the test file naming has to match `--pattern` exactly.
Locally it is kaggle/python-build, in the registry it is gcr.io
Running the CPU image is two commands, one for what you just built and one for what is published:
# Run the image built locally:
docker run --rm -it kaggle/python-build /bin/bash
# Run the pre-built image from gcr.io
docker run --rm -it gcr.io/kaggle-images/python /bin/bashThe GPU pair adds `--runtime nvidia` and changes both names, to `kaggle/python-gpu-build` and `gcr.io/kaggle-gpu-images/python`:
docker run --runtime nvidia --rm -it kaggle/python-gpu-build /bin/bash
docker run --runtime nvidia --rm -it gcr.io/kaggle-gpu-images/python /bin/bashTwo details are easy to trip over. The local images are built under the `kaggle/` Docker Hub namespace while the published ones live on Google Container Registry under `kaggle-images` and `kaggle-gpu-images`, so `docker pull` will not find what `build` produced and `docker run` will not find what is published. And GPU access is not just the flag: the page sends you to a pinned comment on issue 361 for what your container needs before the GPU is visible.
Only these two registries are named, CPU and GPU. There is no third one.
A TPU image ships, and the page never mentions TPU
There is a `tpu/` directory at the repository root and a published `v171-TPU-...` release, and the words TPU, tpu and accelerator appear nowhere in the README.
Everything about the TPU variant is missing from the documentation. There is no `gcr.io` registry for it, no `--tpu` flag on either `./build` or `./test`, no `docker run` example, and no statement of which accelerator library it carries or whether it is the same base as the GPU image with a different runtime. The release tag exists and the directory exists; the instructions do not.
For the CPU and GPU variants the page is complete enough to run from: build, test, pull, run. For TPU, an adopter has to read `tpu/` and infer the rest.
That gap is also the reason to read the release list rather than the README. The tags are where the variants are enumerated, and the tags list three while the documentation covers two.
Editorial conclusion
Adopt the Kaggle docker-python repository if you need to reproduce the environment Kaggle Notebooks actually run in, or if you need a package added to that image and are willing to write a test for it. Do not adopt it as a general Python base image: the published tags encode a build number and a commit hash per variant, the CPU, GPU and TPU images at build 171 come from three different commits, and a TPU image ships with no TPU flag, no TPU registry and no TPU run example anywhere in the page. Verify first which variant you need and whether the two documented registries cover it, then read kaggle_requirements.txt before opening a pull request, because one file is the whole contribution surface.
Frequently asked questions
What does the Kaggle docker-python repository contain?
It contains the Dockerfile used to build the CPU-only and GPU images that run Python Notebooks on Kaggle, stored on disk as Dockerfile.tmpl. The built images are published at gcr.io/kaggle-images/python for CPU and gcr.io/kaggle-gpu-images/python for GPU.
How do you build a Kaggle docker-python image?
Run `./build` from the repository root. The documented flags are `--gpu` to build a GPU image and `--use-cache` for faster iterative builds. No TPU flag is documented, although TPU images are published.
How do you test a Kaggle docker-python image?
Run `./test`. Use `--gpu` for the GPU image, `--pattern test_keras.py` or `-p test_keras.py` to run one test, and `--image gcr.io/kaggle-images/python:ci-pretest` or `-i` with that value to test against a specific published image instead of the one just built.
How do you request a new package for Kaggle notebooks?
The page says to evaluate first whether installing the package yourself in your own notebooks suits your needs, following its missing-packages guide. If it does not, edit kaggle_requirements.txt, build the image, add a test for the package, and open a pull request.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/kaggle-docker-python)