Self-hosted service
YaoFANGUK/video-subtitle-remover avatar
YaoFANGUK/video-subtitle-remover

video-subtitle-remover: inpainting hard subtitles out of video, locally

基于AI的图片/视频硬字幕去除、文本水印去除,无损分辨率生成去字幕、去水印后的图片/视频文件。无需申请第三方API,本地实现。AI-based tool for removing hard-coded subtitles and text-like watermarks from videos or Pictures.

13,097 stars1,656 forksPythonApache-2.0

At a glance

What is it?
YaoFANGUK/video-subtitle-remover is an Apache-2.0 tool that removes hard-coded subtitles and text watermarks from video and images by AI inpainting, at the original resolution, with no third-party API. It runs locally, which means you supply the GPU and install PaddlePaddle, PyTorch and ONNX Runtime yourself.
Who is it for?
Use video-subtitle-remover if you need hard-coded subtitles or text watermarks removed at full resolution and the video cannot be uploaded to someone else's service, since it runs locally under Apache-2.0 with no third-party API. It suits owners of a CUDA-capable NVIDIA card, or anyone willing to use one of the published Docker images, and it handles images in batches as well.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 92 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What it actually does

Video-subtitle-remover, or VSR, is an AI tool that removes hard subtitles from video, and text watermarks from images. The README describes it as generating a subtitle-free file at lossless resolution, meaning the output keeps the original dimensions rather than cropping away the subtitle band.

That distinction is the whole value. Cropping a video to cut off burned-in subtitles throws away part of the frame, and blurring leaves an obvious rectangle. Inpainting reconstructs the pixels behind the text, so the frame keeps its size.

The README says the fill is done by a strong AI model rather than by adjacent pixel filling or mosaicing, which is the difference between plausible reconstruction and a smudge.

Two modes are offered: you can pass a subtitle area and only that region is processed, or pass no area and all text in the whole video is removed. Images can be handled in batches for watermark removal, and there is a companion project, video-subtitle-extractor, for extracting the original subtitle text before removing it.

The inpainting modes are the real choice

The command line exposes the algorithm as a parameter, which is the most informative thing in the README:

bash
--inpaint-mode {sttn-auto,sttn-det,lama,propainter,opencv}

The default is sttn-auto. Five modes across two families is a lot of choice, and the naming suggests the trade: STTN appeared in release 1.1.0 as a new algorithm that skips subtitle detection, described in the release notes as greatly increasing removal speed, while the det variants presumably run detection first.

The other names are separate inpainting models, LaMa and ProPainter, plus a plain OpenCV fallback. Someone processing a long video will care about the speed difference; someone processing a single difficult frame will care that LaMa and ProPainter reconstruct differently.

The README does not document what each mode is best at, so the practical approach is to run a short clip through two modes and compare before committing to a long batch.

Subtitle areas are given as coordinates, ymin ymax xmin xmax, and the flag can be repeated for multiple regions, which is what lets different tasks use different areas.

Installing it from source

The README asks for Python 3.12 or newer and recommends a virtual environment:

shell
python -m venv videoEnv

After activating it, dependencies come from the requirements file:

shell
pip install -r requirements.txt

The heavy part is the frameworks underneath, and the README gives per-platform instructions for them. PaddlePaddle and PyTorch are installed separately with CUDA-matched wheels, and Linux additionally needs onnxruntime-gpu, with different versions for CUDA 11.x and 12.x.

Four runtime modes are supported: CUDA for NVIDIA cards, CPU with no GPU, DirectML for AMD and Intel graphics on Windows, and Apple Silicon on macOS. Getting the CUDA version right matters, since the README ties specific PyTorch and PaddlePaddle builds to specific CUDA releases.

This is not a pip install and go project. Expect to spend an evening on the environment, which is why the README leads with prebuilt packages instead.

The prebuilt and Docker routes

The README points first at prebuilt Windows archives on the releases page, and publishes a comparison table for them. The packages ship Python 3.12 with PaddlePaddle 3.0.0, and differ in PyTorch and target: vsr-windows-cpu.7z with Torch 2.7.0, vsr-windows-directml.7z with Torch 2.4.1 for non-NVIDIA Windows GPUs, and three CUDA builds, 11.8, 12.6 and 12.8, each with a stated compute capability range.

That table is worth reading before downloading, because the CUDA build you pick has to match your card's compute capability.

The Docker route avoids the environment work entirely. The README gives a command per GPU generation:

shell
docker run -it --name vsr --gpus all eritpchy/video-subtitle-remover:1.4.0-cuda12.8 python backend/main.py -i test/test.mp4 -o test/test_no_sub.mp4

Images are published for cuda11.8, cuda12.6, cuda12.8, directml and cpu. Results are copied back out of the container:

shell
docker cp vsr:/vsr/test/test_no_sub.mp4 ./

For anyone who just wants to try it, the Docker image is the shortest path. There is also a GUI, run through gui.py, with a ui/ directory in the repository.

Language, platform and community

The project is Chinese-language first. The README is in Simplified Chinese with a separate README_en.md, the release notes are in Chinese, and the support channel listed is a set of QQ groups, three of which the README marks as full.

That is worth knowing before you file an issue. English-language users can read the translated README and the code, but discussion happens in a Chinese-language community, and an English question may not get an answer.

Platform support is broad: Windows, macOS including Apple Silicon, and Linux, with the prebuilt packages currently Windows-only and the Docker images covering the rest.

Release 1.4.0, published on 2026-04-09, added support for different subtitle areas per task, updated the PP-OCRv5 model, added a configurable output path and batch image watermark removal, fixed subtitle selection box anomalies in some environments, and added Python 3.13 and macOS Apple Silicon support. The last push to the repository was on 2026-06-30.

Where it will disappoint

Inpainting is reconstruction, and reconstruction fails where there is nothing to reconstruct from. A subtitle sitting over a plain sky or a wall is the easy case. Text over a face, over moving detail, or over a scene that changes rapidly frame to frame is the hard case, and no model recovers what was never visible.

The README does not publish failure cases or quality metrics, so this is a property of the method rather than a documented limitation. Test on a representative clip rather than a demo clip.

Cost is the second constraint. This runs locally, and locally means you own the compute. A feature-length video on CPU is slow, and the README's own table of CUDA builds tells you the project expects a GPU.

There is also an ethical boundary the README does not address. Removing a watermark is fine on content you own and not fine on someone else's, and the tool does not distinguish. That is your responsibility rather than the software's.

Online services as the alternative

The alternative most people consider is a commercial online subtitle or watermark removal service, and the difference is who owns the compute and where the video goes.

An online service needs your video uploaded, processes it on their hardware and returns a file, usually metered by minutes or by credit. You get no installation, no CUDA matching and no GPU purchase, and you give up confidentiality: a private or unreleased video is now on someone else's storage.

VSR runs on your machine under Apache-2.0, so nothing leaves and there is no per-minute cost. You pay in setup time and in hardware.

The deciding question is usually the content. For a public video you downloaded and want cleaned up, an online service is less hassle. For anything unreleased, client-owned or under embargo, local processing is the only defensible option, and this is one of the few open tools that does it at full resolution.

Licence and upkeep

The project is Apache-2.0, which is permissive, allows commercial use and carries an explicit patent grant. That is the right licence for a tool that might end up inside a video pipeline at a company.

Maintenance is steady but slow. Release 1.4.0 came on 2026-04-09, 1.1.1 on 2025-04-24 and 1.1.0 on 2023-12-28, with the last push on 2026-06-30. The gap between 1.1.0 and 1.1.1 is over a year, so releases are occasional rather than scheduled.

The real upkeep burden is the dependency stack. PaddlePaddle, PyTorch and ONNX Runtime all move, and each CUDA generation brings a new matrix of compatible wheels. The project's answer is prebuilt packages and Docker images, which is the right answer, and it means adopting this mostly means tracking which image tag matches your GPU.

Pin the image tag rather than using latest, and keep the container workflow in your notes, since rebuilding the environment from scratch is the expensive part.

Editorial conclusion

Use video-subtitle-remover if you need hard-coded subtitles or text watermarks removed at full resolution and the video cannot be uploaded to someone else's service, since it runs locally under Apache-2.0 with no third-party API. It suits owners of a CUDA-capable NVIDIA card, or anyone willing to use one of the published Docker images, and it handles images in batches as well. Do not expect a quick install from source: PaddlePaddle, PyTorch and ONNX Runtime all need CUDA-matched wheels, and the prebuilt packages are Windows-only. Test on a clip that resembles your real material, because inpainting reconstructs rather than reveals and the README publishes no failure cases. Start with the Docker image for your GPU generation, and if text sits over moving detail, compare two inpaint modes before running the whole file.

Frequently asked questions

What is video-subtitle-remover (VSR)?

It is an AI tool that removes hard-coded subtitles from video and text watermarks from images, generating output at the original resolution, with the fill done by an inpainting model rather than by adjacent pixel filling or mosaicing.

Does video-subtitle-remover need a GPU?

It supports CUDA for NVIDIA cards, CPU with no GPU, DirectML for AMD and Intel graphics on Windows, and Apple Silicon on macOS. CPU works but is slow for long video, and the project publishes CUDA-matched packages and Docker images.

Can I remove subtitles from only part of the video?

Yes. The command line takes --subtitle-area-coords with ymin ymax xmin xmax, and the flag can be given multiple times for several areas. With no area specified, all text in the video is removed.

Is there a Docker image for video-subtitle-remover?

Yes. Images are published as eritpchy/video-subtitle-remover with tags for cuda11.8, cuda12.6, cuda12.8, directml and cpu, and results are copied out with docker cp.

Which inpainting mode should I use?

The default is sttn-auto, and the options are sttn-auto, sttn-det, lama, propainter and opencv. The README does not say which suits which material, so comparing two modes on a short clip is the practical approach.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. Releases
  5. YaoFANGUK/video-subtitle-remover on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/yaofanguk-video-subtitle-remover.svg)](https://hysenlabs.com/projects/yaofanguk-video-subtitle-remover)