Model or dataset
WEIFENG2333/VideoCaptioner avatar
WEIFENG2333/VideoCaptioner

VideoCaptioner: an LLM subtitle pipeline that runs free without an API key

🎬 卡卡字幕助手 | VideoCaptioner - 基于 LLM 的智能字幕助手 - 视频字幕生成、断句、校正、字幕翻译全流程处理!- A powered tool for easy and efficient video subtitling.

16,115 stars1,399 forksPythonGPL-3.0

At a glance

What is it?
VideoCaptioner chains speech recognition, sentence segmentation, LLM correction, translation and video synthesis into one CLI and desktop app. Its free path (bijian ASR plus Bing translation) works with no configuration at all, which is the main reason to look at it.
Who is it for?
Adopt VideoCaptioner if you already have an OpenAI-compatible endpoint and want transcription, segmentation, translation and burning behind one command, or if the free bijian plus Bing path covers your language pair. Skip it if you need a fully offline pipeline or a permissively licensed component inside a closed product, since the licence is GPL-3.0.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 18 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Who VideoCaptioner is built for

The repository describes it as a large-language-model based video subtitle tool covering recognition, optimization, translation and synthesis in one pass. The practical audience is narrower than that sentence suggests. It fits people who process video in batches and want the subtitle chain automated: a talk recording that needs English captions, a Chinese interview that needs Japanese subtitles, a downloaded clip that needs text burned in. The pyproject classifiers list the intended audience as End Users :: Desktop, and the entry points confirm two front ends, a CLI and a PyQt5 desktop app, both shipped from the same package.

The project is not aimed at editors who want frame-level control over a single subtitle file. The pipeline makes decisions for you: speech recognition produces word-level timestamps, then segmentation and LLM correction reshape the text before translation. If your job is adjusting one cue by 200 milliseconds, a manual editor is the faster tool. VideoCaptioner is for the case where doing that by hand across a 40-minute video is the thing you are trying to stop doing.

The pipeline: from audio to a burned-in subtitle track

The README gives the data flow as a single line: audio or video input, speech recognition, subtitle segmentation, LLM optimization, translation, video synthesis. Each arrow is a separate stage that can also be invoked on its own.

Recognition is handled by one of several engines. The CLI table lists faster-whisper, whisper-api, bijian, jianying and whisper-cpp. The README marks bijian and jianying as free, and states that the free features need no configuration. The README also mentions word-level timestamps plus VAD (voice activity detection) as the mechanism behind recognition accuracy, and describes the segmentation step as LLM semantic understanding rather than fixed character or duration limits. Translation is described as context-aware with a reflection and refinement mechanism, which in practice means the model sees surrounding cues rather than each line in isolation. Batch concurrency is listed as a property of the processing.

That architecture has one consequence worth naming: the quality of everything after recognition depends on the segmentation step, and segmentation depends on an LLM. Without a configured model, the free path still transcribes and translates, but the semantic re-segmentation described in the README is not part of it.

Installing VideoCaptioner and running a first transcription

The README gives a single pip command that installs both the CLI and the desktop version. Python support in pyproject is >=3.10,<3.13, so a 3.13 interpreter will not satisfy the requirement.

bash
pip install videocaptioner

After installation, the free path needs no key. The README's first CLI example transcribes a local file with the bijian engine, which it lists as free.

bash
videocaptioner transcribe video.mp4 --asr bijian

You should get a subtitle file for the input video. To translate an existing subtitle file with the free Bing service, the README shows the subtitle subcommand with a translator and a target language.

bash
videocaptioner subtitle input.srt --translator bing --target-language en

The LLM stages (subtitle optimization and model-based translation) need an OpenAI-compatible endpoint. Configuration is set through the config command, one key at a time.

bash
videocaptioner config set llm.api_key <your-key>
videocaptioner config set llm.api_base https://api.openai.com/v1
videocaptioner config set llm.model gpt-4o-mini

The README states the precedence order explicitly: command-line arguments override environment variables prefixed with VIDEOCAPTIONER_, which override the config file, which overrides defaults. Running `videocaptioner config show` prints the values currently in effect, which is the fastest way to find out why a setting you thought you wrote is being ignored. Once keys are in place, the whole chain runs in one command.

bash
videocaptioner process video.mp4 --target-language ja

The desktop version starts with `videocaptioner-gui`, or with `videocaptioner gui`, or by running `videocaptioner` with no arguments at all. Windows users can alternatively download an installer from the releases page, and macOS users are given a one-line shell script in the README.

Where the free path stops and the paid path starts

The split is clear in the documentation and it is the first thing to understand before adopting. Recognition through bijian or jianying and translation through Bing or Google are free and require no configuration. LLM subtitle optimization and LLM translation require an API key and a base URL. The README lists three example providers: the project's own relay at api.videocaptioner.cn, SiliconCloud, and DeepSeek, and notes that any OpenAI-compatible service works.

So the honest framing is that VideoCaptioner is free to try and not necessarily free to use at full quality. If you skip the LLM configuration you lose the semantic re-segmentation and the context-aware translation that the README presents as the project's distinguishing features. What remains is a working transcriber and translator, which is useful but is roughly what several other tools already do.

There is a second boundary: the free ASR engines are network services. The README does not describe them as running locally, and the configuration section treats them as the no-setup option. If your material cannot leave your machine, the free engines are the wrong choice, and you would be looking at faster-whisper or whisper-cpp instead.

Limitations and cases where this is the wrong tool

The project classifies itself as Development Status :: 4 - Beta in pyproject.toml. Treat that as the maintainers' own description of stability rather than a marketing label, and plan around it if the pipeline sits in an automated job.

The dependency list is heavy for a subtitle tool: PyQt5 is pinned at 5.15.11 and PyQt-Fluent-Widgets at 1.8.4, and both are listed as core dependencies rather than under the gui optional group, which is empty. Installing the CLI therefore pulls the desktop stack with it. On a headless server or inside a slim container this is real weight, and the pinned Qt versions are the kind of thing that collides with an existing environment. A virtual environment is the sane default here.

The README does not document rollback, and it does not describe what happens when an LLM returns malformed output beyond the presence of json-repair in the dependencies, which suggests the project expects to repair imperfect model responses rather than prevent them. If your workflow cannot tolerate a run that needs re-running, that is a gap to test before depending on it.

Finally, the pipeline is linear and opinionated. If you need to keep the original cue boundaries and only translate, the `subtitle` subcommand is the right entry point and the full `process` command is the wrong one, because it will re-segment and re-optimize text you already approved.

How it compares to Whisper, Subtitle Edit and translation wrappers

Whisper is a speech recognition model, not a subtitle workflow. Using it directly means you get transcripts with timestamps and then write your own segmentation, translation and burning logic. VideoCaptioner wraps that layer and adds the LLM steps, which is the actual difference: it is an orchestrator, and the recognition engine is a swappable part of it. If you already have a Whisper pipeline you like, VideoCaptioner's value is the segmentation and translation stages, not the transcription.

Subtitle Edit is a desktop editor for subtitle files. It works on text you already have and gives you direct control over timing and formatting. VideoCaptioner starts earlier, from the audio, and makes the timing decisions for you. The two are not substitutes: an editor is what you reach for when the machine got it wrong, and VideoCaptioner is what you reach for when doing it by hand is not viable.

Translation wrappers around a single API are the closest functional overlap, and the difference is again the chain. A translation wrapper takes subtitle text and returns subtitle text. VideoCaptioner takes video and returns video with subtitles, with the segmentation and optimization steps in between, and it exposes each stage as its own subcommand so you can stop partway.

Licence and what upgrading actually costs

The project is licensed GPL-3.0, stated in both the README and the LICENSE file, and repeated in the pyproject license field and the classifier. For individual use and for internal tooling this is unremarkable. For embedding the code in a product you distribute, GPL-3.0 carries obligations that a permissive licence does not, and that is a question for your own legal review rather than something this article can settle.

On upgrades: the last push to the repository was on 2026-07-19, and the most recent tagged release in the list is v1.4.2 from 2026-05-24, with a separate ffmpeg-bin release on 2026-06-17 marked as a runtime dependency that should not be deleted. The gap between the last push and the last release means the master branch may carry changes that are not in a tagged version. If you install from PyPI you are tracking releases; if you clone and run from source you are tracking master, and the README's development section uses uv for that path.

bash
git clone https://github.com/WEIFENG2333/VideoCaptioner.git
cd VideoCaptioner
uv sync && uv run videocaptioner

That command also pulls the test suite into reach: `uv run pytest tests/test_cli/ -q` runs the CLI tests, and `uv run pyright` runs type checking. Running those before an upgrade is a cheap way to find out whether your environment still matches what the project expects.

Editorial conclusion

Adopt VideoCaptioner if you already have an OpenAI-compatible endpoint and want transcription, segmentation, translation and burning behind one command, or if the free bijian plus Bing path covers your language pair. Skip it if you need a fully offline pipeline or a permissively licensed component inside a closed product, since the licence is GPL-3.0. Before committing, run videocaptioner transcribe on a real clip and read the segmentation output, then run videocaptioner config show to confirm which values are actually in effect, because command-line arguments override environment variables and the config file.

Frequently asked questions

Do I need an API key to use VideoCaptioner?

No. The README states that the free features, bijian speech recognition and Bing or Google translation, need no configuration and work immediately after installation. An API key is only required for LLM subtitle optimization and model-based translation.

How do I install VideoCaptioner on macOS?

The README gives a one-line shell script for macOS under the alternative installation methods, and a pip install for the general case. The pip route installs both the CLI and the desktop version.

Which Python version does VideoCaptioner require?

The pyproject file specifies requires-python >=3.10,<3.13, and the classifiers list 3.10, 3.11 and 3.12. A 3.13 interpreter does not satisfy the requirement.

What speech recognition engines can VideoCaptioner use?

The CLI table lists faster-whisper, whisper-api, bijian, jianying and whisper-cpp. The README marks bijian and jianying as free, and the transcribe command takes the engine through the --asr flag.

Official sources

  1. License: GPL-3.0
  2. Project website
  3. README
  4. Releases
  5. WEIFENG2333/VideoCaptioner on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/weifeng2333-videocaptioner.svg)](https://hysenlabs.com/projects/weifeng2333-videocaptioner)