docker-whisperX: 175 images a week, tag names that are not versions, and one model that needs alignment switched off
Dockerfile for WhisperX: Automatic Speech Recognition with Word-Level Timestamps and Speaker Diarization (Dockerfile, CI image build and test)
At a glance
- What is it?
- A community image set for WhisperX, rebuilt every week against upstream main, keyed by model and language. Convenient tag names, and no immutable build anywhere in the scheme.
- Who is it for?
- docker-whisperX is a good answer if you want WhisperX running on a GPU tonight and do not need to prove which commit you ran, because the tags are readable, the GPU prerequisites are written down, and the alignment-model cache can be shared across containers.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly Dockerfile, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Tags name a model and a language, never a build
The tag scheme is `WHISPER_MODEL`-`LANG`, so `tiny-en`, `base-de` and `large-v3-zh` are all valid, plus a `no_model` tag that is also called `latest`. What no tag carries is a build identity. The Dockerfile opens with `ARG VERSION=EDGE` and `ARG RELEASE=0`, and the scheduling note says a workflow runs weekly against the main branch of the upstream project and rebuilds every image, with the WhisperX code base inside each image aligned to the git submodule commit hash. The repository publishes no GitHub releases at all. Put together, that means `ghcr.io/jim60105/whisperx:base-en` is a moving name: the same string can resolve to a different WhisperX commit and a different set of Python packages on different days. If a result has to be repeatable, build your own image from a submodule commit you pin, and treat the published tags as a convenience for exploration. The run commands themselves are short, with everything after the second `--` handed to WhisperX:
docker run --gpus all -it -v ".:/app" ghcr.io/jim60105/whisperx:base-en -- --output_format srt audio.mp3
docker run --gpus all -it -v ".:/app" ghcr.io/jim60105/whisperx:large-v3-ja -- --output_format srt audio.mp3
docker run --gpus all -it -v ".:/app" ghcr.io/jim60105/whisperx:no_model -- --model tiny --language en --output_format srt audio.mp3The third line supplies the model and the language at run time instead, which is what the `no_model` tag is for. Every one of them asks for all GPUs rather than a device you choose.
Excluded model names and published tags collide on the same suffix
Three model families are left out on purpose: the `*.en` variants, `large-v1` and `large-v2`, excluded because the author judged them not frequently used, with instructions to build them yourself. That exclusion sits next to published tags like `base-en`, `tiny-en` and `large-v3-en`, because the tag format is model then language joined by a hyphen. Whisper's English-only `base.en` and this project's `base-en` differ by one character and mean different things: one is an English-only model, the other is the multilingual base model paired with the English alignment model. With the dotted forms excluded and the hyphenated forms published, a tag list is genuinely ambiguous to read at a glance. Check the build matrix at `.github/workflows/04-build-matrix-images.yml` and the tag listing on ghcr.io before you assume a particular model is present.
Language carries two different meanings in this project
At build time, `LANG` is a build argument defaulting to `en`, and a space separated list such as `"LANG=pl fr en"` embeds several alignment models in one image. At run time, you are told you do not have to pass the language at all, because WhisperX still detects it or falls back to en and uses that to choose the alignment model, and alignment models are language specific. The README is explicit that the multi-language build argument exists purely for embedding models, and equally explicit that WhisperX does not do well with multiple languages inside a single audio file. So the list widens what an image can align, not what a recording can contain. If your audio switches languages mid-file, neither the build list nor the runtime flag addresses it.
A Hokkien model that emits Mandarin characters and cannot be aligned
The `breeze-asr-26-zh` image wraps a Taiwanese Hokkien recogniser published by MediaTek Research and re-packaged for the faster-whisper runtime by another party. It is fine-tuned from Whisper on roughly 10,000 hours of synthetic Taigi speech, including Taigi and Mandarin code-switching, and it transcribes spoken Taigi into Mandarin Chinese characters, which is why it ships under the `zh` pairing rather than a Taigi one. The cost is stated directly: on genuine Taigi audio, phoneme-level alignment will not work, because the bundled `zh` wav2vec2 alignment model is trained on Mandarin phonology and cannot align Taigi pronunciations against Mandarin-character output. The workaround is a flag, `--no_align`, passed after the second `--` in the run command. Word-level timestamps are the headline of WhisperX, and this is the one image in the set where you have to turn that headline off.
The diarization download ignores TORCH_HOME, so CACHE_HOME is pinned
Four arguments near the top of the Dockerfile exist to work around a specific bug, and the comment cites the issue number:
ARG CACHE_HOME=/.cache
ARG CONFIG_HOME=/.config
ARG TORCH_HOME=${CACHE_HOME}/torch
ARG HF_HOME=${CACHE_HOME}/huggingfaceThe note says the diarization model download, when it uses an auth token, does not respect `TORCH_HOME`, so `CACHE_HOME` has to be set to exactly the same path as the default. Two consequences follow. Relocate that cache and the workaround stops applying. And the README's caching instructions mount precisely that path, which is why the shared-volume example is:
docker run --gpus all -it -v ".:/app" -v whisper_cache:/.cache ghcr.io/jim60105/whisperx:latest -- --model large-v3 --language en --output_format srt audio.mp3That guidance applies to the `no_model` and `latest` tags, since the pre-built images already carry their weights and the volume is there for the alignment models.
Two base stages, one per architecture, and a package outside Debian main
The file opens with `prepare_base_amd64` and `prepare_base_arm64`, both `FROM docker.io/library/python:3.13-slim`, each declaring `TARGETARCH` and `TARGETVARIANT`. Two environment lines are baked into the image rather than passed at run time, `NVIDIA_VISIBLE_DEVICES=all` and `NVIDIA_DRIVER_CAPABILITIES=compute,utility`, so the image always asks for every GPU the runtime offers. The NVIDIA libraries are not in Debian's main component, so the build rewrites the sources list to add `non-free` and `non-free-firmware` and installs `libnppicc12` with `--no-install-recommends`. Each architecture gets its own apt caches, keyed by architecture and variant, with `sharing=locked` on the package lists so parallel builds of the same target cannot write that cache at once. The comment above the amd64 mount points at a buildx issue about multi-arch cache mounts, which is what that locking is answering. Two details sit above the stages as well. The build declares `UID=1001`, and `LOAD_WHISPER_STAGE=load_whisper` and `NO_MODEL_STAGE=no_model` are marked as arguments for caching stage builds in CI that should be left alone when building locally, which is a hint that the stage names are load-bearing for the matrix jobs rather than decoration. The `# syntax=docker/dockerfile:1` line at the very top is required for the mount syntax used below it.
175 images of 10GB, rebuilt every week, from a moving submodule
The stated purpose of the repository is managing the continuous integration build on a GitHub Free runner on a weekly schedule: 175 images built in parallel, each 10GB. That is about 1.75TB of images pushed to ghcr.io every week, which is the reason the model list is pruned at all, and it explains why layer reuse and cache read and write ordering are the stated strategy rather than build speed. The recurring inputs are also the fragile part. WhisperX arrives as a git submodule, so a plain clone is not enough and the build instructions ask for `git clone --recursive`, and the weekly job points at the upstream main branch rather than a tag. The build command section itself is cut off after naming the `en` language and the beginning of the word `large`, so the exact invocation is not on the page. The two build arguments, `LANG` and `WHISPER_MODEL`, are described above it, with `WHISPER_MODEL` defaulting to `base` and `LANG` to `en`, and the default that ships in the Dockerfile is the same pair.
Editorial conclusion
docker-whisperX is a good answer if you want WhisperX running on a GPU tonight and do not need to prove which commit you ran, because the tags are readable, the GPU prerequisites are written down, and the alignment-model cache can be shared across containers. It is the wrong answer if reproducibility is the point: there are no releases, the build defaults to VERSION=EDGE, and a weekly job rebuilds every tag against upstream main, so a tag you recorded last month is not the image you get this month. Build from a submodule commit you control, and if you transcribe Taiwanese Hokkien, plan on passing --no_align from the first run.
Frequently asked questions
How are image tags named in jim60105/docker-whisperX?
As `WHISPER_MODEL`-`LANG`, for example `tiny-en`, `base-de` or `large-v3-zh`. A `no_model` tag with no pre-downloaded weights is also published under the name `latest`, and the actual build matrix lives in `.github/workflows/04-build-matrix-images.yml`.
Can I pin a specific version of jim60105/docker-whisperX?
Not from a tag. The repository publishes no GitHub releases, the Dockerfile defaults `VERSION` to `EDGE` and `RELEASE` to `0`, and a scheduled workflow rebuilds every image weekly against the main branch of the upstream project. The WhisperX code inside each image aligns to the git submodule commit hash, so record that hash yourself if you need reproducibility.
Why does the breeze-asr-26-zh image need the --no_align flag?
It is a Taiwanese Hokkien model that transcribes spoken Taigi into Mandarin Chinese characters, and the bundled `zh` wav2vec2 alignment model is trained on Mandarin phonology. On genuine Taigi audio it cannot align Taigi pronunciations against Mandarin-character output, so the alignment pass has to be skipped with `--no_align`.
How do I share downloaded alignment models between containers in jim60105/docker-whisperX?
Mount the cache path as a volume and use the `no_model` or `latest` tag, for example `-v whisper_cache:/.cache` against `ghcr.io/jim60105/whisperx:latest`. The path matters: the Dockerfile pins `CACHE_HOME` to `/.cache` to work around a diarization download that ignores `TORCH_HOME`.
Which Whisper models does jim60105/docker-whisperX not build?
The `*.en` variants along with `large-v1` and `large-v2`, excluded because the author judged them not frequently used. Those need to be built locally. Note that `base-en` and `tiny-en` are published tags and mean the multilingual model paired with an English alignment model, not the excluded dotted model names.
What does jim60105/docker-whisperX require before it can use a GPU?
On Windows, Docker Desktop, the CUDA Toolkit, the NVIDIA Windows Driver and Docker running under WSL2. On Linux and macOS, an NVIDIA GPU driver and the NVIDIA Container Toolkit. The images also set `NVIDIA_VISIBLE_DEVICES=all` and `NVIDIA_DRIVER_CAPABILITIES=compute,utility` in the image itself.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/jim60105-docker-whisperx)