# JJYB AI VideoAutoCut: a Windows desktop workbench for narrated video editing

> A local-first Flask and PyWebView application that analyses footage, writes a commentary script, generates speech, syncs it to shots and exports a cut, with licensing terms that limit you to personal use.

**jianjieyiban/JJYB_AI_VideoAutoCut** — JJYB_AI 智剪 - 智能视频自动剪辑与AI解说工具（离线TTS、原创解说、混剪、AI配音）

- Repository: https://github.com/jianjieyiban/JJYB_AI_VideoAutoCut
- Stars: 1,067 · Forks: 204
- Language: HTML
- License: NOASSERTION
- Published: 2026-10-07 · Updated: 2026-10-07 · Language: en
- Canonical page: https://hysenlabs.com/projects/jianjieyiban-jjyb-ai-videoautocut

## A pipeline that starts from the footage and ends in a draft

The README describes the project as a local, desktop-first tool for video creators that organises material analysis, intelligent shot segmentation, commentary scripts, speech generation, audio and video sync, subtitles, remix and export into one workflow, so that assets do not have to be shuttled between separate tools. That framing matters, because most of what this application does exists somewhere else in some form. The claim here is about the ordering and the glue.

The stack is legible from the README's own table. Flask, Flask-SocketIO and PyWebView provide the desktop window and the server underneath it, with HTML, JavaScript and Layui in the frontend. Project data lives in SQLite, in a path the README points at. Video work goes through FFmpeg, ffprobe, OpenCV, MoviePy and ImageMagick. Audio uses librosa, SoundFile and Pydub. Text models sit behind adapters that cover OpenAI, Claude, Gemini, Kimi, Qwen and others. Speech covers Edge TTS, gTTS, pyttsx3, Azure, Volcengine and an optional Coqui path.

The repository has 1,067 stars and 204 forks, with 13 open issues, which is a healthy ratio for a project of this size. The default branch is `videoautocut`, the language GitHub reports is HTML, and the release history has one entry, v3.3.0, published 2026-08-09. There is no licence file the repository metadata can interpret, and the licence situation is discussed below, because for this project that is not a formality.

## Shot detection with a documented fallback chain

The commentary workflow begins with `SmartShotSegmenter`, and the README is specific about how it finds cuts. It first tries TransNetV2 for learned shot segmentation. When that is unavailable, it falls back to OpenCV frame differencing combined with histogram detection, and it does not impose a frame-count ceiling in either path.

The fallback is worth pausing on, because a deep model and a signal-processing heuristic are very different tools. The README presents them as a chain rather than a preference, which is an honest description of what a local application can do. TransNetV2 has to be present as an optional resource before it can be used at all, and the release notes confirm that optional pieces such as IndexTTS2, TransNetV2 and ImageMagick are prepared separately rather than shipped in the source archive.

Once shots exist, the script is generated by whichever visual or text model you configured. The README lists twelve narration styles in the script configuration, including a conversational default, a high-energy compilation style, a news broadcast register and a film commentary style. Alongside that sit four emotional tones, twelve commentary types and ten opening hook types, all selectable by hand.

There is also a style preset dashboard that applies a whole video type at once. Choosing film commentary, for example, is described as synchronising script style, emotional tone, commentary type, hook type and writing persona together. A separate writing persona panel offers twelve personas, four narrative points of view including first person and interior monologue, five speech-rate steps and four emotional intensity steps. That is a lot of dials, and the design answer is that most people should use the preset and only the curious should use the individual controls.

## The sync engine refuses rather than fakes a duration

`PuppetSyncEngine` is the component that matches voice segments to shots, and the README's description of it contains the most careful engineering language in the project. It works segment by segment and offers direct concatenation, trimming to the highlight, restricted slow motion, and merging short voice segments. Then it says something no marketing document would say: when a source shot is too short, even after slowing it down as far as safety allows, the engine refuses rather than padding the timeline with a frozen final frame.

The same discipline shows up in the target duration feature. You can ask for a finished length, and the result is checked against the real timeline and with FFprobe. When a safe speech rate or the footage you supplied makes the request impossible, the tool says it could not be met instead of producing something that looks compliant.

Original-audio intercutting works the same way. You choose to weave the original soundtrack back in either by proportion of the finished video or at intervals between segments, and set the volume balance. In interval mode the matcher prefers a successor shot from the same source or a temporally nearby one; in ratio mode it inserts near the same-source boundary. If no usable shot exists, you get a warning or a failure, and a shot is never reused to fill the gap.

This refusal behaviour is the single most useful thing to know about the project. Long jobs also support checkpointing, so a run interrupted partway can resume and stream progress back to the interface.

## Beat analysis, and failing when there are not enough beats

The remix mode has two shapes. The ordinary mode segments footage in batch, scores shots by image quality, and picks highlights, optionally matching voice segments to them by audio duration. The music-beat mode uses librosa to read the background track's beats and energy, and offers dense, medium and sparse beat densities, with an allowance for moderate slow motion in climactic passages.

A rhythm pack sits on top of this with twelve presets, including standard remix, fast-cut beat matching, high-energy trailer, plot progression, game highlight, Vlog flow, product showcase, knowledge breakdown, travel montage, food close-up, pet reaction and slow montage. The README says these presets coordinate shot selection, beat density, keyframes, transitions and timeline rules together, which is the difference between a preset that is really a set of defaults and one that is really a policy.

One change is documented as a deliberate correction: an advanced feature moved away from copying commentary-style script structure toward matching beat points and shot descriptions specifically for remix, with a low-confidence list and top-K reranking for shots. In other words, the project took a feature that was doing the wrong job and redirected it.

Target duration works here too, with the beat engine computing a maximum beat count from the requested length and FFprobe verifying the actual output. When the footage cannot supply enough beats, the README says it fails explicitly, and does not pad with repeated shots or freezes. The same JianYing draft export is available from both modes, producing video, audio, subtitle and background music tracks that the JianYing Pro desktop application can open directly.

## Windows, Python and a large optional dependency surface

The environment requirements are Windows 10 or 11, Python 3.10 through 3.13, and FFmpeg installed and callable, with ffprobe available too. ImageMagick is optional and bundled under `resource/imageMagick/`. An NVIDIA CUDA setup helps with segmentation, YOLO and Whisper tasks. IndexTTS2 runs in its own Python 3.10 environment with dependencies declared in `models/index-tts/pyproject.toml`, which keeps the heavy local voice-cloning stack out of the main environment.

Installing the main dependencies is one command:

```powershell
py -3 -m pip install -r requirements.txt
```

and verifying the environment uses the bundled checker:

```powershell
py -3 scripts/check_system.py
```

FFmpeg is checked the usual way:

```powershell
ffmpeg -version
ffprobe -version
```

The application starts either by double-clicking the bundled batch file or by running the entry point:

```powershell
py -3 frontend\app.py
```

It tries to open a PyWebView desktop window, and the README notes the same service is reachable at `http://127.0.0.1:5000` when the window is not in use. First-run configuration is deliberately split into three independent configuration groups, for text language models, visual models and image generation models, so that an API key for cover art does not have to live beside the key used for script writing. There is a test-connection button on each card.

The `requirements.txt` file is worth reading as a map of the project's ambitions. It pulls in MoviePy, scenedetect, pysrt, faster-whisper, ultralytics, torch and a set of cloud SDKs including openai, anthropic, google-generativeai, dashscope and zhipuai, with several speech and ASR engines present as commented alternatives. That breadth is the reason the checklist exists.

## The licence is the real constraint, and the README says so up front

The repository metadata reports no recognised licence. The README does not leave that as an accident: it carries an authorisation notice stating that the source code is published for personal study, research and self-use, and that without prior written permission from the author, commercial use, paid services, paid distribution, commercial deployment and use inside a commercial product are all prohibited. The v3.3.0 release notes repeat the same terms and add that downloading, installing or using the software constitutes agreement to the terms in the `LICENSE` file at the repository root.

This is a source-available arrangement with a personal-use restriction, not an open source licence, and treating it as MIT or Apache because the source is visible would be wrong. If you intend to run it inside a business, sell anything produced with it, or offer it as part of a paid offering, you need to contact the author first.

Two other honesty notes are worth carrying through. The project only processes text, video, music and reference audio that you have the rights to use, and the voice-cloning feature requires explicit authorisation from the owner of the reference audio, with the interface expecting a consent flag. Second, whether cloud models, external TTS or local voice cloning actually work depends on your API configuration, your runtime and whether the model weights are present. The README's framing is that availability is conditional, not guaranteed.

The repository structure is a flat list that matches those boundaries: `backend/`, `config/`, `frontend/`, `models/`, `outputs/`, `scripts/`, `storage/`, `tests/` and a `开发文档/` directory, plus `requirements.txt`, `RELEASE_NOTES.md` and the Windows launcher batch file. A new v3.3.0 release script was added for reproducible source archives, and the ignore rules were tightened to keep API keys, databases, logs, user material, generated video and large model weights out of version control.

## Conclusion

What makes this project worth reading is not that it calls itself an AI video tool, but how plainly it documents where automation fails. When a source shot is too short to cover a voice segment at safe slow motion, the sync engine refuses the timeline instead of freezing the last frame. When the beat engine runs out of usable cuts, the remix fails loudly rather than repeating a shot to hit a duration. When a cloud image model is unreachable, cover generation falls back to a local template and says why. These are the decisions that separate a tool you can trust with a batch of real footage from a demo, and they are written down rather than hidden. The cost is a Windows-only Python application with a large optional dependency surface, and the licensing is narrower than the word open source usually implies: the source is published for personal study, research and self-use, and commercial deployment needs written permission from the author.

## FAQ

### What operating system does JJYB AI VideoAutoCut run on?

Windows 10 and Windows 11, according to the README's platform badge and its environment requirements. The quick start instructions use the Windows `py -3` launcher, and the documented launch path is a bundled batch file or the Flask entry point, so there is no stated macOS or Linux path.

### Do I need an API key to use the commentary workflow?

Yes, at least one text model needs to be configured on first run with an API key, base URL and model name, and the README recommends the custom OpenAI-compatible option. Cloud visual analysis for the commentary mode needs a separately configured vision model, and cover generation needs its own image generation configuration. The three groups are kept independent so one key does not have to cover all of them.

### Is the voice cloning feature safe to use with any audio?

No. The README states that you must obtain explicit authorisation from the owner of the reference audio before using the voice replication feature, and that the interface requires a consent flag. The project also states generally that it only handles text, video, music and reference audio you have legal rights to.

### What happens when the footage is too short to cover the narration?

The sync engine refuses rather than padding. When a source shot is still too short after safe slow motion, the timeline is not generated with a frozen final frame to fake the length. The same principle applies to target durations and to beat-driven remix: the output is verified against the real timeline with FFprobe, and the job fails explicitly instead of repeating shots or freezing frames to reach the number.

### Can I use this project commercially or sell videos made with it?

The README's authorisation notice restricts the project to personal study, research and self-use, and prohibits commercial use, paid services, paid distribution and commercial deployment without prior written permission from the author. The v3.3.0 release notes restate the restriction, so the source is published for reading but not as an unrestricted open source licence.

## Sources

- [Issues](https://github.com/jianjieyiban/JJYB_AI_VideoAutoCut/issues)
- [jianjieyiban/JJYB_AI_VideoAutoCut on GitHub](https://github.com/jianjieyiban/JJYB_AI_VideoAutoCut)
- [README](https://github.com/jianjieyiban/JJYB_AI_VideoAutoCut/blob/videoautocut/README.md)
- [Releases](https://github.com/jianjieyiban/JJYB_AI_VideoAutoCut/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/jianjieyiban-jjyb-ai-videoautocut
