# Video-subtitle-extractor: local OCR for hard-coded subtitles

> VSE detects and reads burned-in subtitles on your own machine and writes srt files, with no third-party OCR API. The trade-off is a heavyweight PaddlePaddle install and a subtitle detection engine that is slow in its most accurate mode.

**YaoFANGUK/video-subtitle-extractor** — 视频硬字幕提取，生成srt文件。无需申请第三方API，本地实现文本识别。基于深度学习的视频字幕提取框架，包含字幕区域检测、字幕内容提取。A GUI tool for extracting hard-coded subtitle (hardsub) from videos and generating srt files. 

- Repository: https://github.com/YaoFANGUK/video-subtitle-extractor
- Stars: 9,541 · Forks: 972
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/yaofanguk-video-subtitle-extractor

## Hard-coded subtitles are pixels, and VSE reads them that way

A soft subtitle lives in its own stream and can be copied out with a demuxer. A hard subtitle is drawn into the video frames, so recovering it means finding the text region in each frame and running optical character recognition on it. Video-subtitle-extractor (VSE) is built for that second case. The README describes the pipeline as key frame extraction, text region detection, text recognition, filtering of non-subtitle text, removal of watermarks and logos, de-duplication of repeated lines, and finally an srt or txt file. The audience is people archiving or translating material whose only subtitle is baked into the picture, and who want the OCR to happen locally instead of through Baidu or Alibaba cloud OCR. The project states that no third-party API is required, which is the main reason to pick it over a cloud OCR wrapper. It is a GUI program first: the README's usage section says to click Open, choose the video, adjust the subtitle region, and click Run.

## Detection, recognition and the three modes that decide your runtime

The repository splits into backend/, ui/, gui.py and a test/ directory, with requirements.txt at the top level. Two engines appear in the README's mode table: VideoSubFinder, used for subtitle detection in fast and automatic modes on Windows, Linux and macOS, and an engine the table calls VSE, used only in precise mode. The OCR model size also changes by mode. Fast uses the mini model. Automatic picks the mini model on CPU and the large model on GPU. Precise uses the large model and, according to the table, detects frame by frame. The README is direct about the consequences: fast may drop a few lines and produce a few wrong characters, automatic may drop a few lines but has almost no wrong characters, and precise drops nothing and has almost no wrong characters but is very slow. Its own advice is to start with fast or automatic and only move to precise when lines are being lost. That is a sensible default, and it also tells you the project's accuracy problem is recall, not character quality. Non-subtitle text is handled by a typo map at backend/configs/typoMap.json, a JSON object that maps a wrong string to a corrected one and maps a string to an empty value to delete it. The README's example maps "性感荷官在线发牌" to an empty string, which removes every occurrence.

## Installing VSE from source and running the GUI

The README recommends downloading a release archive and running it directly, and only falling back to a conda or virtual environment install from source if that fails. Source installs need Python 3.12 or newer. Create and activate a virtual environment, then move into the source directory. This is the CPU path, which needs no GPU and installs the CPU build of PaddlePaddle 3.3.1 from the project's own package index before the pinned requirements:

```bash
python -m venv videoEnv
source videoEnv/bin/activate
pip install paddlepaddle==3.3.1 -i https://www.paddlepaddle.org.cn/packages/stable/cpu/
pip install -r requirements.txt
python gui.py
```

On Windows the activation line is videoEnv\Scripts\activate instead. For an NVIDIA card the README recommends CUDA 11.8 with cuDNN 8.6.0 and a different PaddlePaddle index, then the same requirements file:

```bash
pip install paddlepaddle-gpu==3.3.1 -i https://www.paddlepaddle.org.cn/packages/stable/cu118/
pip install -r requirements.txt
```

For AMD or Intel GPUs on Windows there is a DirectML path that adds a second requirements file, and an ONNX path for macOS and AMD ROCm that the README marks as untested and asks you not to file issues about. The CLI entry point is python ./backend/main.py. After launch, open one video for a single extraction or several for batch extraction, adjust the subtitle region, and press Run. Batch mode requires every video to share the same resolution and subtitle region, which is the constraint that decides whether batching saves you time or silently misreads frames.

## The install surface is the real cost, not the OCR

requirements.txt pins paddlepaddle to 3.3 and paddleocr to 3.4.0 alongside opencv-python, pyside6, pyside6-fluent-widgets, numpy 2.2 and smaller libraries such as pysrt, Levenshtein, wordsegment and pyclipper. That is a large dependency set for one task, and the PaddlePaddle version is the part most likely to fight your machine. The README notes that NVIDIA 50-series cards need CUDA 12.8.0 or newer, that Paddle 3.3.1 does not yet support it, and that the DirectML build is the suggested workaround. It also says the MacOS path does not support CUDA at all. The README's first troubleshooting entry is simply that CUDA and cuDNN problems are solved by installing the versions matching your card and driver. Two further constraints are stated as hard requirements rather than suggestions: video and program paths must not contain Chinese characters or spaces, with D:\下载\vse\运行程序.exe and E:\study\kaoyan\sanshang youya.mp4 given as failing examples, and the 7z extraction error is fixed by updating 7-Zip. If your library lives under a directory with a space in its name, plan to move it before you start.

## Where VSE is the wrong tool

The clearest wrong case is a video whose subtitles are a separate stream. VSE reads pixels; it does not demux, so running it on a soft-subbed file converts clean text into OCR output with the error rate that implies. The second case is timing precision. Every mode relies on detecting subtitle regions across frames, and the README's own mode table admits fast and automatic can drop lines. Precise mode drops nothing but is described as very slow, so if your source has faint, animated or rapidly changing subtitles, you are choosing between an incomplete srt and a long render. Third, VSE is not a subtitle remover. The README points to a separate project, video-subtitle-remover, for erasing watermarks, station logos or the original hard subtitle, and only claims that VSE can filter such text out of the recognized output through typoMap.json. Fourth, batch extraction assumes uniform resolution and subtitle position, so a mixed-resolution folder breaks the premise. Finally, the README does not document rollback, an upgrade path between releases, or a headless server mode, so anyone planning unattended pipelines should treat the GUI as the intended interface.

## How it compares with VideoSubFinder and cloud OCR

VideoSubFinder is the closest reference point, and it is not a rival bolted on from outside: the README's mode table lists it as the subtitle detection engine for fast and automatic modes in every supported OS, with the project's own VSE engine reserved for precise mode. So the meaningful comparison is between VSE's mode selection and running VideoSubFinder alone. VideoSubFinder is a dedicated detector that produces subtitle images; VSE wraps detection with PaddleOCR recognition, de-duplication and srt writing, which is the part that saves you a second tool and a manual OCR step. The cloud OCR comparison is simpler. Baidu or Alibaba OCR would remove the local model download and the CUDA version matching, and would likely handle unusual fonts better, but it means uploading frames and depending on a paid service. VSE's stated position is local recognition with no API, and that is the whole reason to accept the PaddlePaddle install. Against a general OCR pipeline you assemble yourself, the difference is the subtitle-specific filtering: region selection, typoMap.json corrections, and repeated-line removal are already wired in.

## Licence, releases and what maintenance looks like

The project is Apache-2.0, which permits commercial use and modification provided the licence and notices are preserved; the LICENSE file sits at the repository root. That matters here because you are likely to bundle the tool into a translation or archiving workflow. Apache-2.0 does not resolve the separate terms of the models and runtimes you install alongside it, and the README does not discuss those, so check PaddlePaddle and PaddleOCR licensing if you redistribute a packaged build. On releases, the recent history is 2.0.0 in October 2023, 2.0.3 in April 2025, and 2.2.0 in April 2026, so the cadence is roughly annual with a long gap in the middle. The last push to the default branch was on 2026-04-09. That is recent enough that the code is not abandoned, but the version pinning in requirements.txt means an upgrade is not just pulling the newest release: moving PaddlePaddle or paddleocr forward can change OCR behaviour, and the README offers no migration notes between 2.0.3 and 2.2.0. If you need reproducible output, freeze your own environment rather than tracking the requirements file.

## Conclusion

Adopt video-subtitle-extractor if you have a batch of hard-subbed videos, a machine that can carry PaddlePaddle, and a preference for local OCR over uploading footage to an online service. Skip it if your subtitles are already a separate stream, since extraction is the wrong operation and a demuxer is enough. Before committing, verify the three things the README is explicit about: paths without Chinese characters or spaces, a PaddlePaddle 3.3.1 build that matches your CUDA or DirectML setup, and whether the fast mode's dropped lines are acceptable for your material.

## FAQ

### Is there a way to extract subtitles from a video with video-subtitle-extractor?

Yes, for hard-coded subtitles. The README describes a pipeline that extracts key frames, detects the text region, recognizes the text with local OCR, filters non-subtitle text and duplicate lines, and writes an srt or txt file. It states that no third-party API is needed.

### How do I extract subtitles with video-subtitle-extractor rather than VLC?

The README does not mention VLC. It says to click Open, select the video, adjust the subtitle region and click Run in the GUI, or run python ./backend/main.py for the command line version, after installing the dependencies for your hardware.

### Do I need to unzip anything to use video-subtitle-extractor?

The README recommends downloading a release archive and extracting it before running. Its second troubleshooting entry covers a 7z extraction error and says to upgrade the 7-Zip program to the latest version.

### Is video-subtitle-extractor the best free tool to remove subtitles from videos?

No. VSE extracts hard-coded subtitles into srt or txt files; it does not remove them from the picture. The README points to a separate project, video-subtitle-remover, for removing watermarks, station logos and the original hard subtitle.

## Sources

- [Issues](https://github.com/YaoFANGUK/video-subtitle-extractor/issues)
- [License: Apache-2.0](https://github.com/YaoFANGUK/video-subtitle-extractor/blob/main/LICENSE)
- [README](https://github.com/YaoFANGUK/video-subtitle-extractor/blob/main/README.md)
- [Releases](https://github.com/YaoFANGUK/video-subtitle-extractor/releases)
- [YaoFANGUK/video-subtitle-extractor on GitHub](https://github.com/YaoFANGUK/video-subtitle-extractor)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/yaofanguk-video-subtitle-extractor
