subsai wraps eight speech backends behind one CLI, and installs all eight by default
🎞️ Subtitles generation tool (Web-UI + CLI + Python package) powered by OpenAI's Whisper and its variants 🎞️
At a glance
- What is it?
- Subs AI is a GPL-3.0 Python tool that turns audio or video into subtitles using Whisper and seven other backends, with a Streamlit web interface, a batch CLI and six output formats. Its dependency file treats backend specific packages as required, so a plain install pulls in every engine it lists.
- Who is it for?
- Subs AI is worth its place when one job dominates: a folder of recordings that all need the same treatment, in a format a specific tool expects, on a machine where the models already exist. Two things decide whether it fits.
- Can I use it commercially?
- Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository last received commits 166 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
requirements.txt installs the backends it says are specific
The dependency file is organised in commented groups, and one of the groups is not commented. Under a heading that reads Backend specific dependencies, the lines that follow are active requirements: openai 1.60.1, ffmpeg-python, joblib, whisper-timestamped pulled from a git URL, pywhispercpp 1.3.1, faster_whisper with no version at all, stable-ts 2.18.2, whisperx pulled from a git URL pinned to commit 8c58c54635cd6ee2d9d8665a3cf789863f6ed700, transformers 4.48.1 and openai-whisper 20240930. So a plain pip install of the project installs every engine the README lists, including two that compile or fetch at install time. Above that group sit the always needed lines: ffsubsync, pysubs2, dl_translate 0.3.0 and pydub, then a group explicitly there to avoid dependency version problems, numpy<2, torch==2.2.0 and torchaudio==2.2.0. The web interface adds pandas 1.5.3, streamlit around 1.20, streamlit_player and streamlit-aggrid. Choosing a backend is therefore a file edit, not a command-line switch.
Installed from a git URL, allowed on Python 3.8, recommended on 3.10
There is no PyPI instruction on the page. Installation is one line:
pip install git+https://github.com/absadiki/subsaiBefore it, ffmpeg has to be present, and the commands for that are quoted from the Whisper installation instructions, one per package manager, with a note that Rust may be needed if tokenizers ships no wheel for the platform, and that a missing setuptools_rust module means running pip install setuptools-rust. The version story is where the packaging and the page disagree. requires-python reads >=3.8, while the note under the install command recommends Python 3.10 or 3.11 and says 3.12 or later may have compatibility issues. The manifest also sets license to the literal text LICENSE rather than to an SPDX identifier, and it requires setuptools-scm in the build system while carrying a static version of 1.6.2. Two console scripts are registered: subsai pointing at subsai.cli:main and subsai-webui pointing at subsai.webui:run. A second note covers the most common install failure, a GPU that torch cannot see because pip brought in the CPU wheel, with an issue number given for it.
One positional argument that takes a file, a list or a manifest
The CLI is a single argparse surface with one positional argument and a set of short options:
usage: subsai [-h] [--version] [-m MODEL] [-mc MODEL_CONFIGS] [-f FORMAT] [-df DESTINATION_FOLDER] [-tm TRANSLATION_MODEL]
[-tc TRANSLATION_CONFIGS] [-tsl TRANSLATION_SOURCE_LANG] [-ttl TRANSLATION_TARGET_LANG]
media_file [media_file ...]Each option changes one decision. The model option selects the backend and the model configuration option supplies settings for it, so a backend that needs extra arguments gets them without a second command. The format option picks the subtitle container and the destination folder decides where files land. The four translation options are a set rather than a flag: a translation model, translation configurations, a source language and a target language, which means the same run can transcribe and translate in one pass. The positional argument accepts the path of a media file, a list of files, or a text file holding them, which is the batch path, and the bracketed repetition after it confirms that several files can be given directly. Audio and video are both accepted as input. The web interface, by contrast, is one command, subsai-webui, which opens a page in the default browser.
A Streamlit interface with the file watcher and usage stats turned off
The web interface is built on Streamlit, pinned to roughly 1.20 with the player and aggrid add-ons, and the container starts it in a deliberately quiet mode: the entrypoint passes --server.fileWatcherType none and --browser.gatherUsageStats false, so the running server neither rescans files nor reports usage. The interface is stated to be fully offline with no third-party services, to work on Linux, Mac and Windows, and to allow editing the subtitles it produces. Three tools are integrated rather than bolted on. Translation goes through dl-translate with four translation models selectable, facebook/nllb-200-distilled-600M, facebook/m2m100_418M, facebook/m2m100_1.2B and facebook/mbart-large-50-many-to-many-mmt. Auto-sync uses ffsubsync, and the interface can merge the subtitles back into the video file. Output format is not the model's business but the container library's: six subtitle formats are supported through pysubs2, namely SubRip, WebVTT, substation alpha, MicroDVD, MPL2 and TMP.
Compose reserves a GPU the Compose schema does not read, and publishes no ports
The compose file defines two services built from the same Dockerfile. One is subsai-webui, which carries a GPU reservation written as a deploy stanza:
deploy:
resources:
reservations:
devices:
- driver: nvidia
capabilities: [ gpu ]The other, subsai-webui-cpu, is the same image with no reservation at all. The reservation uses the key Compose inherits from Swarm rather than the runtime or gpus key, so it describes intent for a Swarm deployment and is not what a plain compose up reads when deciding whether to give the container a GPU. Neither service declares a ports mapping, so the 8501 the Dockerfile exposes is not published by this file; a reader who expects compose to make the interface reachable has to add the mapping or run the image directly. The file also opens with version 3.8, which is the legacy schema key that current Compose releases warn about, and both services inherit whatever device access the default configuration grants.
The image ships torch 2.0.1 and the requirements ask for 2.2.0
The container starts from a PyTorch runtime image, pytorch/pytorch:2.0.1-cuda11.7-cudnn8-runtime, with a python:3.10.6 line left commented out above it, which shows the base was swapped rather than replaced. It installs git, gcc and mono-mcs from apt, with an apt-get upgrade in the same layer, and mono-mcs is the Mono C# compiler, which no file in the repository gives a visible consumer. It then installs requirements with no cache, copies the project, and pip installs the package itself, so the pinned torch==2.2.0 and torchaudio==2.2.0 lines are resolved over the 2.0.1 already present in the base image rather than reusing it. The CUDA 11.7 and cuDNN 8 runtime in the image name is therefore the toolkit the older torch was built against. Work happens in /subsai, the port exposed is 8501, and the entrypoint runs the web interface as a script path with the two Streamlit flags rather than through the subsai-webui console script the manifest registers.
Version 1.6.2, no releases, and two different homepages
The manifest records version 1.6.2, and the repository has no GitHub releases, so nothing on the project page gives you a downloadable artifact or a changelog entry to pin against. The last push to main was 2026-04-20. Two addresses compete for the word homepage: the project's URL table points Homepage at the GitHub repository itself, while the repository's declared homepage is the documentation site at absadiki.github.io/subsai/, built from the mkdocs.yml file and the docs directory sitting at the root. That site is where the deeper documentation lives, and the README is the short version of it, carrying a generated table of contents and a feature checklist that fills most of the page with quoted descriptions of each upstream project. Two notebooks in examples/ go further than the page does: one covers translation and one covers voice activity detection, which is the preprocessing step Whisper backends use to decide what is speech. tests/ sits at the root as well, so the packaging, the documentation and the examples are all in one tree.
Editorial conclusion
Subs AI is worth its place when one job dominates: a folder of recordings that all need the same treatment, in a format a specific tool expects, on a machine where the models already exist. Two things decide whether it fits. First, the dependency file, because the packages grouped under backend specific dependencies are not optional markers but active requirements, so an install brings in a pinned torch 2.2.0, openai-whisper, stable-ts, pywhispercpp, faster-whisper, transformers 4.48.1 and two git dependencies, and you choose your backend by editing that file rather than by passing a flag. Second, the Python version, since the metadata allows 3.8 and above while the page recommends 3.10 or 3.11 and warns about 3.12 and later, and the container image starts from a torch 2.0.1 build that the requirements then overwrite. The web interface runs fully offline with no third-party services, the CLI handles batch work, and the last push to main was 2026-04-20 with no GitHub releases to pin.
Frequently asked questions
Which speech recognition backends can subsai use?
Eight are listed: openai/whisper, whisper-timestamped for word-level timestamps and confidence, whisper.cpp through the absadiki/pywhispercpp binding, faster-whisper, whisperX with diarization, stable-ts for more reliable timestamps, any Hugging Face Transformers speech recognition model, and the OpenAI speech-to-text API or any OpenAI-compatible endpoint such as speaches.ai.
How do I install subsai?
ffmpeg has to be installed first, then the package is installed straight from git with pip install git+https://github.com/absadiki/subsai. The page recommends Python 3.10 or 3.11 and warns that 3.12 or later may have compatibility issues, even though the packaging metadata allows 3.8 and above.
Which subtitle formats does subsai write?
Six, through the pysubs2 library: SubRip, WebVTT, substation alpha, MicroDVD, MPL2 and TMP. The format is selected on the command line with the -f option, and the web interface can also merge the finished subtitles back into the video file.
Does the subsai web interface send my audio anywhere?
The page describes the web interface as fully offline with no third-party services, and it can be run locally with the subsai-webui command or through the container. Choosing the OpenAI API backend is the exception, since that backend is a hosted service by definition.
What does subsai add on top of running Whisper directly?
A web interface with subtitle editing, translation through four selectable translation models, auto-sync using ffsubsync, merging into the video, six output formats, a command line that accepts a file, a list or a manifest for batch work, and the project as an importable Python package.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/absadiki-subsai)