# VideoLingo review: Netflix-style subtitle cutting, translation and dubbing in one Streamlit app

> VideoLingo is an Apache-2.0 Python pipeline that chains yt-dlp, WhisperX, an LLM and several TTS backends into one Streamlit UI. The interesting part is not the stack, it is the subtitle segmentation step, and the constraints are just as specific.

**Huanshere/VideoLingo** — Netflix-level subtitle cutting, translation, alignment, and even dubbing - one-click fully automated AI video subtitle team | Netflix级字幕切割、翻译、对齐、甚至加上配音，一键全自动视频搬运AI字幕组

- Repository: https://github.com/Huanshere/VideoLingo
- Website: https://docs.videolingo.io
- Stars: 18,508 · Forks: 2,036
- Language: Python
- License: Apache-2.0
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/huanshere-videolingo

## The problem VideoLingo targets: subtitles that read like subtitles

Most automatic video translation tools fail at the same place. Speech recognition returns word timestamps, a machine translation model returns a sentence, and the two get glued together with the original line breaks. The result is a subtitle track that is technically correct and unpleasant to watch, because the segmentation follows the audio's pauses rather than the sentence structure of the target language.

VideoLingo's README describes its core feature as NLP and AI-powered subtitle segmentation, followed by custom plus AI-generated terminology for coherent translation. That ordering is the point. The project treats segmentation as a distinct stage between recognition and translation, and it carries a terminology list (the repository ships custom_terms.xlsx) so that names and domain words stay consistent across a video.

The audience is a single operator, not a studio. Everything runs behind a Streamlit interface on localhost, and the README frames the whole flow as one-click startup and processing. If you are a solo creator moving a channel between languages, or an engineer who wants to see each stage rather than trust a black box, the shape of the tool fits. If you need a review queue with roles and audit trails, it does not.

## How the pipeline is wired: yt-dlp, WhisperX, an LLM, then TTS

The data flow is linear and the components are visible in the repository. yt-dlp pulls the source video. WhisperX performs word-level speech recognition and alignment. From there the text goes to an OpenAI-compatible Chat Completions provider, which handles segmentation, translation, and an optional reflection and rewriting pass. Finally a TTS backend synthesizes the dubbed track, and FFmpeg handles the media work.

The README lists the TTS options explicitly: GPT-SoVITS, Azure, OpenAI, Fish TTS, SiliconFlow Fish/CosyVoice2, Edge TTS, F5-TTS, and a custom adapter at core/tts_backend/custom_tts.py. That custom adapter is the escape hatch for a service the project does not ship.

Speech recognition has two paths. You either run WhisperX locally, which pulls in the CUDA side, or you call the ElevenLabs API. The local path is where most of the installation weight sits, and it is also where the version pinning comes from: requirements.txt pins torch 2.8.0, torchaudio 2.8.0, torchvision 0.23.0 and torchcodec 0.7, with a comment noting that WhisperX 3.8 uses that matched family and that Transformers 5 would require huggingface-hub >= 1, which conflicts with WhisperX's hub<1 requirement. That is a real dependency knot, not a stylistic choice.

Two operational details matter more than they look. The README advertises pause, resume, or stop at any step, and detailed logging with progress resumption. In practice that means a failed run does not force a full re-transcription, which on a long video is the difference between a retry and an evening.

## Installing VideoLingo with uv and running the first video

The README asks for Git, uv and FFmpeg before anything else, and says to reopen the terminal and verify all three. That is not boilerplate: the host installer chooses the PyTorch wheel family by inspecting the machine. If nvidia-smi reports CUDA 12.8 or higher it selects cu128, otherwise cu126, and without an NVIDIA GPU it selects CPU packages. This selects Python packages only, not a system CUDA Toolkit.

The FFmpeg requirement is stricter than most projects. The README states that the application needs FFmpeg 7 shared libraries for the pinned TorchCodec 0.7, and that FFmpeg 8 or 9 alone is not compatible. On Windows it asks for a shared-library build with its bin directory on PATH; on macOS and Linux the documented commands are:

```bash
brew install ffmpeg
```

```bash
sudo apt install ffmpeg
```

With the prerequisites in place, clone the repository and let setup_env.py build an isolated environment. uv downloads Python 3.13 itself, so no preinstalled Python is needed for this command, though the application supports 3.10 through 3.13.

```bash
git clone https://github.com/Huanshere/VideoLingo.git
cd VideoLingo
uv run --no-project --python 3.13 setup_env.py
```

Start the app with the interpreter from the generated environment. The README gives both variants:

```bash
.venv\Scripts\streamlit run st.py        # Windows
.venv/bin/streamlit run st.py            # macOS / Linux
```

On Windows there is also OneKeyStart.bat, which prefers ~/.venvs/videolingo when present and falls back to the project .venv. Open http://localhost:8501 and fill in the API URL, key and model in the sidebar. The sidebar model field supports search and filtering across the provider's full model list, which is a small feature that saves a lot of copy-pasting.

For a Linux NVIDIA container the README documents a Docker path. The image defaults to CUDA 12.8.1 and the Dockerfile accepts either 12.8.1 or 12.6.3, mapping them to cu128 and cu126 respectively and failing the build on anything else.

```bash
docker build -t videolingo .
docker run -d -p 8501:8501 --gpus all videolingo
```

The Dockerfile exposes port 8501 and starts Streamlit with --server.address=0.0.0.0 and --server.headless=true. The README points to the Docker docs for the matched CUDA 12.6 alternative and for persistence settings, which is worth reading before you treat a container as disposable, since output lives inside it by default.

## Where VideoLingo breaks: JSON contracts, mixed languages and dubbing timing

The README's Current Limitations section is unusually candid, and three items deserve attention before you plan around this tool.

First, the LLM must return JSON in the structure the workflow expects. When it does not, the README directs you to output/gpt_log/error.json. It also warns against deleting all output as a first troubleshooting step, noting that successful response caches and completed outputs can be reused on retry, and that changing the model alone does not regenerate every completed stage. Anyone who has debugged a pipeline by wiping its cache will recognize this as a warning earned the hard way. The practical consequence is that your choice of LLM is constrained by structured output reliability, not by translation quality alone.

Second, mixed-language speech is not handled well. Local WhisperX uses one recognition and alignment language per segment, so a video that switches languages mid-sentence is not guaranteed to keep accurate text and timing in every language. If your source material is code-switching, this is the wrong tool for the alignment stage.

Third, dubbing quality and timing depend on the translation, the TTS service and the speech rate, and the README states plainly that speed adjustment does not guarantee natural delivery or perfect synchronization. Dubbing is offered as a feature, not as a solved problem.

There is a fourth constraint in the recognition section: background noise and language-specific alignment models affect word timestamps, vocal separation may help, and numbers and symbols may lack reliable word timings. The README's advice is to inspect the resulting subtitles. For Chinese input specifically, it notes that you must explicitly select Chinese to get the punctuation-enhanced Belle Whisper model, which is easy to miss in a dropdown.

## VideoLingo compared with pyvideotrans, KrillinAI and SoniTranslate

The alternatives people search for alongside VideoLingo differ in where they put the human. Jianchang512 pyvideotrans is the closest in spirit: a desktop-oriented translation tool that also covers recognition, translation and dubbing. The difference that matters is not the feature list but the interface model. VideoLingo is a Streamlit web app you start yourself and reach at localhost:8501, and its pipeline is built around a terminology file and an LLM-driven segmentation and rewriting stage. If you want a desktop application rather than a browser tab, that is a reason to look elsewhere.

KrillinAI is another project in the same search space. Without running either, the honest statement is that both sit in the same category of automated video localization, and the deciding factor for a given team will be which backends and which languages they already have credentials for. VideoLingo's documented TTS list is broad, and its custom adapter at core/tts_backend/custom_tts.py means a service that is not on the list can be added without forking the project.

SoniTranslate and Linly-Dubbing appear in the same searches. The general trade-off is consistent: tools that hide the stages are faster to start and harder to correct, while tools that expose the stages ask you to configure an LLM endpoint, an ASR path and a TTS provider. VideoLingo is firmly in the second group. It asks for an API URL, key and model before it will do anything, and it writes logs and intermediate output you are expected to read.

## Licence, maintenance and what an upgrade actually costs

VideoLingo is licensed under Apache-2.0, and the LICENSE file sits at the top level of the repository. Apache-2.0 is a permissive licence with an explicit patent grant and a requirement to preserve notices; the practical implication for most users is that shipping the tool inside a product is permitted, but the usual obligations around attribution and notice retention still apply. Nothing here is legal advice, and the terms that matter for your situation are in the LICENSE file itself, not in a review.

The repository is not archived, and the last push was on 2026-09-15, which is days before this article. Release v3.0.4 landed the same day, following v3.0.1 in February 2026 and v3.0.0 in May 2025. That cadence suggests the project is being worked on, but the gap between v3.0.0 and v3.0.1 is worth noting if you depend on a feature that arrives in a minor release.

Upgrade cost is dominated by the dependency pins, not by the application code. requirements.txt holds torch at exactly 2.8.0, torchaudio at 2.8.0, torchvision at 0.23.0 and torchcodec between 0.7 and 0.8, with comments explaining that these move together with WhisperX 3.8 and that Transformers 5 conflicts with WhisperX's huggingface-hub constraint. The FFmpeg 7 shared-library requirement is tied to the same TorchCodec pin. In other words, you cannot casually bump one package, and a future WhisperX release that relaxes the hub constraint is the event that unblocks the rest. The Dockerfile's refusal to build on any CUDA version other than 12.8.1 or 12.6.3 is the same philosophy applied to the container: fail loudly rather than produce a mismatched wheel set.

## Conclusion

Adopt VideoLingo if you already run a WhisperX-class GPU box and want subtitle segmentation and terminology handling inside one interface, and if you are willing to keep the pinned Torch 2.8 / TorchCodec 0.7 / FFmpeg 7 combination intact. Do not adopt it if you need guaranteed timing on mixed-language audio, or if you want a hosted product that does not ask you for an API key. Before committing, check three things: that your FFmpeg build ships 7.x shared libraries, that your LLM can return the structured JSON the workflow expects, and that you have read the note about not deleting all output as a first troubleshooting step.

## FAQ

### Is there a way to auto translate a video with VideoLingo?

Yes. The README describes a one-click flow in the Streamlit interface that combines yt-dlp download, WhisperX recognition, LLM translation and optional dubbing, and it produces subtitle files plus optionally subtitled or dubbed videos. You still have to enter an API URL, key and model in the sidebar before it runs.

### Is VideoLingo an app that can translate audio from a video?

It is a self-hosted Python application, not a hosted service. It runs as a Streamlit app on localhost:8501, and its dubbing stage uses one of several TTS backends such as Azure, OpenAI, Edge TTS, GPT-SoVITS or Fish TTS to produce translated audio.

### What does VideoLingo cost to run?

The software itself is Apache-2.0 and free to install. Costs come from the services you connect: an OpenAI-compatible LLM provider, optionally the ElevenLabs speech recognition API, and whichever TTS backend you choose. The README does not state a price for any of these.

## Sources

- [Huanshere/VideoLingo on GitHub](https://github.com/Huanshere/VideoLingo)
- [License: Apache-2.0](https://github.com/Huanshere/VideoLingo/blob/main/LICENSE)
- [Project website](https://docs.videolingo.io)
- [README](https://github.com/Huanshere/VideoLingo/blob/main/README.md)
- [Releases](https://github.com/Huanshere/VideoLingo/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/huanshere-videolingo
