pyVideoTrans: Open Source Video Translation with AI Dubbing and Subtitle Generation
Translate the video from one language to another and embed dubbing & subtitles.
At a glance
- What is it?
- pyVideoTrans is a Python tool that translates video from one language to another through a four-stage pipeline: speech recognition, subtitle translation, AI voice synthesis, and video remuxing. It supports local offline models alongside a wide range of cloud APIs, and ships a GUI, a CLI, a WebUI, and a Docker image.
- Who is it for?
- pyVideoTrans is the right tool for individuals and small teams who need to translate video content between languages and want control over which ASR, translation, and TTS models are used, including fully local offline operation. It is not suited for production workflows that require guaranteed uptime or formal SLAs: it is an open-source Python application, not a managed service.
- Can I use it commercially?
- Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What pyVideoTrans Does and Who It Is For
pyVideoTrans solves the problem of converting a video in one spoken language into a watchable version in another language, including the voice track. The four-stage pipeline is: speech recognition (ASR) to produce a subtitle file, translation of those subtitles through a chosen model or API, voice synthesis (TTS) to generate a new audio track from the translated text, and remuxing to merge the new audio and subtitles back into the video file. Each stage can be paused for manual review and correction before continuing.
The README positions this as a one-click fully automatic workflow, but also makes the interactive editing available at every stage. This makes it practical for content creators who want a first pass quickly but need to correct names, technical terms, or unusual phrasing before the final export.
The tool is aimed at content creators, educators, and developers who need to localize video content and either want to avoid per-minute API costs by running local models, or want the flexibility to mix and match the best available models for each stage of the pipeline. Windows users get a pre-packaged `.exe` that requires no Python environment. macOS, Linux, and developer users work from source with the uv package manager.
The Four-Stage Pipeline and Supported Models
Each stage of the pipeline has multiple options. For speech recognition, the README lists Faster-Whisper (local, recommended for speed and accuracy), WhisperX and Parakeet (which support speaker diarization and timestamp alignment), and cloud APIs from Alibaba Qwen, ByteDance Volcano, Azure, and Google. For translation, the supported providers include DeepSeek, ChatGPT, Claude, Gemini, MiniMax, Ollama (local), and Alibaba Bailian. For voice synthesis, Edge-TTS is listed as free, with OpenAI, Azure, and Minimaxi as cloud alternatives, and ChatTTS and ChatterBox as additional options.
Voice cloning is a separate capability. The README lists F5-TTS, CosyVoice, and GPT-SoVITS as zero-shot voice cloning models that can replicate a target speaker's voice from a short sample.
Speaker diarization, which assigns different translation voices to different speakers in the original audio, is supported through WhisperX and Parakeet. The README calls this multi-role AI dubbing and describes assigning different AI voices to different speakers.
The utility toolkit listed in the README includes vocal separation, video and subtitle merging, audio-video alignment, and transcript matching. These are separate tools that can be used independently of the full translation pipeline.
Installing and Running pyVideoTrans from Source
The project requires Python 3.10 (the pyproject.toml specifies `>=3.10, <3.11`) and FFmpeg. On macOS, the README provides specific Homebrew commands for FFmpeg and libsndfile:
brew install libsndfile git [email protected]On Ubuntu or Debian:
sudo apt-get install ffmpeg libsndfile1-devInstall the uv package manager, clone the repository, and install dependencies:
git clone https://github.com/jianchang512/pyvideotrans.git
cd pyvideotrans
uv syncLaunch the GUI with:
uv run sp.pyThe CLI mode uses `cli.py` with task-specific flags. A video translation task looks like:
uv run cli.py --task vtv --name "./video.mp4" --source_language_code zh-cn --target_language_code en --voice_role "en-US-GuyNeural"The WebUI, which supports remote or internal network access, requires installing the optional webui extra:
uv sync --extra webui
uv run webui.pyDocker Deployment and GPU Acceleration
A Dockerfile is included in the repository. The CPU build and GPU build are controlled by the `USE_CUDA` build argument. The Dockerfile uses a two-stage base selection: the CPU path uses `python:3.10-slim` and the GPU path uses `nvidia/cuda:12.8.0-cudnn-runtime-ubuntu22.04`.
Build and run the WebUI container:
docker build -t pyvideotrans-webui .
docker run -d -p 7860:7860 --name pyvideotrans pyvideotrans-webuiTo persist configuration and output across container restarts:
docker run -d -p 7860:7860 \
-v ./data/output:/app/output \
-v ./data/config:/app/videotrans \
--name pyvideotrans pyvideotrans-webuiFor NVIDIA GPU acceleration outside of Docker, the README instructs removing the CPU-only PyTorch version and reinstalling the CUDA variant:
uv remove torch torchaudio
uv add torch==2.7 torchaudio==2.7 --index-url https://download.pytorch.org/whl/cu128GPU acceleration requires NVIDIA hardware with CUDA 12.8 and cuDNN 9.11 installed. AMD GPU acceleration through Whisper.NET is documented separately in `docs/whisper_net_setup.md`.
Real Limitations and Cases Where It Is the Wrong Tool
pyVideoTrans requires Python exactly 3.10. The pyproject.toml `requires-python = ">=3.10, <3.11"` rules out Python 3.11 and later. Teams that standardize on a newer Python version must either maintain a separate virtual environment or use the Docker image.
The dependency list in pyproject.toml is extremely long: dozens of pinned packages covering audio processing, cloud SDK clients, machine learning frameworks, and UI libraries. A fresh installation downloads several gigabytes of packages. Model weights for local ASR models like Faster-Whisper are downloaded separately at runtime and can be several gigabytes more. This makes pyVideoTrans unsuitable as a lightweight background utility.
The translation quality depends entirely on the external model or API chosen for the translation stage. pyVideoTrans handles the orchestration; it does not provide its own translation engine. For languages with complex morphology or for highly technical content, the output will reflect the limitations of the underlying translation provider.
The GPL-3.0 license has practical implications for commercial deployment. Any application that incorporates pyVideoTrans as a component and is distributed to users must be released under GPL-3.0 as well. Internal tools that are not distributed are not affected by this requirement.
CLI and WebUI for Headless and Remote Use
The CLI mode, documented in `docs/cli.md`, supports batch processing and server deployment without a display. The README lists four task types: `vtv` for full video translation, `stt` for audio or video to subtitle, `sts` for subtitle translation, and `tts` for text-to-speech. Each task accepts command-line flags for the source and target language codes, the model name, and the output voice role.
An audio-to-subtitle task using the large-v3 Whisper model looks like:
uv run cli.py --task stt --name "./audio.wav" --model_name large-v3The WebUI mode exposes the same functionality through a Gradio-based web interface accessible in a browser. The Dockerfile sets `GRADIO_SERVER_NAME=0.0.0.0` and `GRADIO_SERVER_PORT=7860`, so the WebUI is available on port 7860 after the container starts. This makes it practical to deploy on a server and access it from another machine on the same network without forwarding X11 or VNC.
Development Activity and License
pyVideoTrans is under active development. Three releases appeared in September 2026 alone: v4.12 on 2026-09-06, v4.13 on 2026-09-20, and v4.14 on 2026-09-26. The last push to the repository was on 2026-09-27. The project is not archived.
The license is GPL-3.0. Forks or applications that distribute pyVideoTrans as a component must release their source code under the same license. The repository also includes a `law.txt` file in the top level, which the project name suggests addresses legal or compliance notices relevant to distribution.
The homepage at `pyvideotrans.com` and a community forum at `bbs.pyvideotrans.com` are referenced in the README. Full CLI documentation is in `docs/cli.md`, WebUI documentation in `docs/webui.md`, and AMD GPU setup in `docs/whisper_net_setup.md`.
Editorial conclusion
pyVideoTrans is the right tool for individuals and small teams who need to translate video content between languages and want control over which ASR, translation, and TTS models are used, including fully local offline operation. It is not suited for production workflows that require guaranteed uptime or formal SLAs: it is an open-source Python application, not a managed service. Before adopting it, verify that your system has Python exactly 3.10 (the pyproject.toml specifies `>=3.10, <3.11`), FFmpeg installed and on the system PATH, and enough disk space for the Faster-Whisper model weights if local ASR is the goal. The license is GPL-3.0, which requires derivative works that are distributed to users to be released under the same license.
Frequently asked questions
How do I use pyVideoTrans to translate a video?
Install dependencies with uv sync after cloning the repository, then launch the GUI with uv run sp.py. Select the source and target languages, choose your ASR and translation providers, and load the video file. Each stage (recognition, translation, dubbing) can be paused for manual review. The CLI equivalent for a full translation is uv run cli.py --task vtv with the source file, language codes, and voice role as flags.
Does pyVideoTrans support offline operation without an internet connection?
Yes for the ASR and translation stages: Faster-Whisper runs locally for speech recognition, and Ollama provides local LLM translation. Edge-TTS requires an internet connection, but the README lists other TTS options. Model weights for local models are downloaded on first use and then cached.
What Python version and system dependencies does pyVideoTrans require?
The pyproject.toml specifies Python >=3.10, <3.11, meaning exactly Python 3.10 is required. FFmpeg must be installed and available on the system PATH. On macOS, libsndfile is also required; on Ubuntu and Debian, the package is libsndfile1-dev.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/jianchang512-pyvideotrans)