# pyVideoTrans: Open Source Video Translation with AI Dubbing and Subtitle Generation

> pyVideoTrans is a Python tool that translates video from one language to another through a four-stage pipeline: speech recognition, subtitle translation, AI voice synthesis, and video remuxing. It supports local offline models alongside a wide range of cloud APIs, and ships a GUI, a CLI, a WebUI, and a Docker image.

**jianchang512/pyvideotrans** — Translate the video from one language to another and embed dubbing & subtitles.

- Repository: https://github.com/jianchang512/pyvideotrans
- Website: https://pyvideotrans.com
- Stars: 19,179 · Forks: 2,372
- Language: Python
- License: GPL-3.0
- Published: 2026-09-16 · Updated: 2026-09-16 · Language: en
- Canonical page: https://hysenlabs.com/projects/jianchang512-pyvideotrans

## What pyVideoTrans Does and Who It Is For

pyVideoTrans solves the problem of converting a video in one spoken language into a watchable version in another language, including the voice track. The four-stage pipeline is: speech recognition (ASR) to produce a subtitle file, translation of those subtitles through a chosen model or API, voice synthesis (TTS) to generate a new audio track from the translated text, and remuxing to merge the new audio and subtitles back into the video file. Each stage can be paused for manual review and correction before continuing.

The README positions this as a one-click fully automatic workflow, but also makes the interactive editing available at every stage. This makes it practical for content creators who want a first pass quickly but need to correct names, technical terms, or unusual phrasing before the final export.

The tool is aimed at content creators, educators, and developers who need to localize video content and either want to avoid per-minute API costs by running local models, or want the flexibility to mix and match the best available models for each stage of the pipeline. Windows users get a pre-packaged `.exe` that requires no Python environment. macOS, Linux, and developer users work from source with the uv package manager.

## The Four-Stage Pipeline and Supported Models

Each stage of the pipeline has multiple options. For speech recognition, the README lists Faster-Whisper (local, recommended for speed and accuracy), WhisperX and Parakeet (which support speaker diarization and timestamp alignment), and cloud APIs from Alibaba Qwen, ByteDance Volcano, Azure, and Google. For translation, the supported providers include DeepSeek, ChatGPT, Claude, Gemini, MiniMax, Ollama (local), and Alibaba Bailian. For voice synthesis, Edge-TTS is listed as free, with OpenAI, Azure, and Minimaxi as cloud alternatives, and ChatTTS and ChatterBox as additional options.

Voice cloning is a separate capability. The README lists F5-TTS, CosyVoice, and GPT-SoVITS as zero-shot voice cloning models that can replicate a target speaker's voice from a short sample.

Speaker diarization, which assigns different translation voices to different speakers in the original audio, is supported through WhisperX and Parakeet. The README calls this multi-role AI dubbing and describes assigning different AI voices to different speakers.

The utility toolkit listed in the README includes vocal separation, video and subtitle merging, audio-video alignment, and transcript matching. These are separate tools that can be used independently of the full translation pipeline.

## Installing and Running pyVideoTrans from Source

The project requires Python 3.10 (the pyproject.toml specifies `>=3.10, <3.11`) and FFmpeg. On macOS, the README provides specific Homebrew commands for FFmpeg and libsndfile:

```bash
brew install libsndfile  git  python@3.10
```

On Ubuntu or Debian:

```bash
sudo apt-get install ffmpeg libsndfile1-dev
```

Install the uv package manager, clone the repository, and install dependencies:

```bash
git clone https://github.com/jianchang512/pyvideotrans.git
cd pyvideotrans
uv sync
```

Launch the GUI with:

```bash
uv run sp.py
```

The CLI mode uses `cli.py` with task-specific flags. A video translation task looks like:

```bash
uv run cli.py --task vtv --name "./video.mp4" --source_language_code zh-cn --target_language_code en --voice_role "en-US-GuyNeural"
```

The WebUI, which supports remote or internal network access, requires installing the optional webui extra:

```bash
uv sync --extra webui
uv run webui.py
```

## Docker Deployment and GPU Acceleration

A Dockerfile is included in the repository. The CPU build and GPU build are controlled by the `USE_CUDA` build argument. The Dockerfile uses a two-stage base selection: the CPU path uses `python:3.10-slim` and the GPU path uses `nvidia/cuda:12.8.0-cudnn-runtime-ubuntu22.04`.

Build and run the WebUI container:

```bash
docker build -t pyvideotrans-webui .
docker run -d -p 7860:7860 --name pyvideotrans pyvideotrans-webui
```

To persist configuration and output across container restarts:

```bash
docker run -d -p 7860:7860 \
  -v ./data/output:/app/output \
  -v ./data/config:/app/videotrans \
  --name pyvideotrans pyvideotrans-webui
```

For NVIDIA GPU acceleration outside of Docker, the README instructs removing the CPU-only PyTorch version and reinstalling the CUDA variant:

```bash
uv remove torch torchaudio
uv add torch==2.7 torchaudio==2.7 --index-url https://download.pytorch.org/whl/cu128
```

GPU acceleration requires NVIDIA hardware with CUDA 12.8 and cuDNN 9.11 installed. AMD GPU acceleration through Whisper.NET is documented separately in `docs/whisper_net_setup.md`.

## Real Limitations and Cases Where It Is the Wrong Tool

pyVideoTrans requires Python exactly 3.10. The pyproject.toml `requires-python = ">=3.10, <3.11"` rules out Python 3.11 and later. Teams that standardize on a newer Python version must either maintain a separate virtual environment or use the Docker image.

The dependency list in pyproject.toml is extremely long: dozens of pinned packages covering audio processing, cloud SDK clients, machine learning frameworks, and UI libraries. A fresh installation downloads several gigabytes of packages. Model weights for local ASR models like Faster-Whisper are downloaded separately at runtime and can be several gigabytes more. This makes pyVideoTrans unsuitable as a lightweight background utility.

The translation quality depends entirely on the external model or API chosen for the translation stage. pyVideoTrans handles the orchestration; it does not provide its own translation engine. For languages with complex morphology or for highly technical content, the output will reflect the limitations of the underlying translation provider.

The GPL-3.0 license has practical implications for commercial deployment. Any application that incorporates pyVideoTrans as a component and is distributed to users must be released under GPL-3.0 as well. Internal tools that are not distributed are not affected by this requirement.

## CLI and WebUI for Headless and Remote Use

The CLI mode, documented in `docs/cli.md`, supports batch processing and server deployment without a display. The README lists four task types: `vtv` for full video translation, `stt` for audio or video to subtitle, `sts` for subtitle translation, and `tts` for text-to-speech. Each task accepts command-line flags for the source and target language codes, the model name, and the output voice role.

An audio-to-subtitle task using the large-v3 Whisper model looks like:

```bash
uv run cli.py --task stt --name "./audio.wav" --model_name large-v3
```

The WebUI mode exposes the same functionality through a Gradio-based web interface accessible in a browser. The Dockerfile sets `GRADIO_SERVER_NAME=0.0.0.0` and `GRADIO_SERVER_PORT=7860`, so the WebUI is available on port 7860 after the container starts. This makes it practical to deploy on a server and access it from another machine on the same network without forwarding X11 or VNC.

## Development Activity and License

pyVideoTrans is under active development. Three releases appeared in September 2026 alone: v4.12 on 2026-09-06, v4.13 on 2026-09-20, and v4.14 on 2026-09-26. The last push to the repository was on 2026-09-27. The project is not archived.

The license is GPL-3.0. Forks or applications that distribute pyVideoTrans as a component must release their source code under the same license. The repository also includes a `law.txt` file in the top level, which the project name suggests addresses legal or compliance notices relevant to distribution.

The homepage at `pyvideotrans.com` and a community forum at `bbs.pyvideotrans.com` are referenced in the README. Full CLI documentation is in `docs/cli.md`, WebUI documentation in `docs/webui.md`, and AMD GPU setup in `docs/whisper_net_setup.md`.

## Conclusion

pyVideoTrans is the right tool for individuals and small teams who need to translate video content between languages and want control over which ASR, translation, and TTS models are used, including fully local offline operation. It is not suited for production workflows that require guaranteed uptime or formal SLAs: it is an open-source Python application, not a managed service. Before adopting it, verify that your system has Python exactly 3.10 (the pyproject.toml specifies `>=3.10, <3.11`), FFmpeg installed and on the system PATH, and enough disk space for the Faster-Whisper model weights if local ASR is the goal. The license is GPL-3.0, which requires derivative works that are distributed to users to be released under the same license.

## FAQ

### How do I use pyVideoTrans to translate a video?

Install dependencies with uv sync after cloning the repository, then launch the GUI with uv run sp.py. Select the source and target languages, choose your ASR and translation providers, and load the video file. Each stage (recognition, translation, dubbing) can be paused for manual review. The CLI equivalent for a full translation is uv run cli.py --task vtv with the source file, language codes, and voice role as flags.

### Does pyVideoTrans support offline operation without an internet connection?

Yes for the ASR and translation stages: Faster-Whisper runs locally for speech recognition, and Ollama provides local LLM translation. Edge-TTS requires an internet connection, but the README lists other TTS options. Model weights for local models are downloaded on first use and then cached.

### What Python version and system dependencies does pyVideoTrans require?

The pyproject.toml specifies Python >=3.10, <3.11, meaning exactly Python 3.10 is required. FFmpeg must be installed and available on the system PATH. On macOS, libsndfile is also required; on Ubuntu and Debian, the package is libsndfile1-dev.

## Sources

- [jianchang512/pyvideotrans on GitHub](https://github.com/jianchang512/pyvideotrans)
- [License: GPL-3.0](https://github.com/jianchang512/pyvideotrans/blob/main/LICENSE)
- [Project website](https://pyvideotrans.com)
- [README](https://github.com/jianchang512/pyvideotrans/blob/main/README.md)
- [Releases](https://github.com/jianchang512/pyvideotrans/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/jianchang512-pyvideotrans
