# video-recap-skills: turning a video into a Chinese narration recap inside Claude Code

> A set of six Agent Skills that chain scene detection, ASR, a vision model, script writing, TTS and ffmpeg into one recap pipeline, with an optional editable JianYing draft at the end. It runs locally with ffmpeg and a single MiMo key, and no GPU.

**zenstory-ai/video-recap-skills** — Claude Code / Codex skills that turn a video into a Chinese narration recap (视频解说): scene detection, ASR, VLM, script, TTS, ffmpeg assembly, optional editable JianYing / CapCut draft export (剪映草稿导出). Local ffmpeg + one MiMo key, no GPU. | 用 Claude Code skills 把视频做成中文解说成片，可选一键导出可编辑剪映草稿。

- Repository: https://github.com/zenstory-ai/video-recap-skills
- Website: https://zenstory.ai/video-recap
- Stars: 533 · Forks: 99
- Language: Python
- License: MIT
- Published: 2026-09-20 · Updated: 2026-09-20 · Language: en
- Canonical page: https://hysenlabs.com/projects/zenstory-ai-video-recap-skills

## The gap video-recap-skills fills between an editor and a chat window

Turning a long video into a short narrated recap is normally four jobs that live in different tools: finding the shots worth keeping, transcribing speech, writing a script that fits the runtime, and voicing plus mixing it. video-recap-skills packages that chain as Agent Skills for Claude Code, Codex CLI, OpenCode and OpenClaw. You give the agent a video path and a sentence of intent, and it produces a file named recap_<name>.mp4.

The target user is someone who already drives an agent harness and does not want to assemble a media pipeline by hand. The README is explicit that no GPU and no model download are involved; the heavy lifting goes to Xiaomi MiMo over an API for ASR, VLM and TTS, while the local machine only needs Python 3.10+ and ffmpeg on PATH. That is a deliberate trade: you accept per-call API cost and a network dependency in exchange for not running anything locally beyond ffmpeg.

## How the six skills chain understanding, planning and assembly

The README's flow diagram splits the work into four stages. Stage one is comprehension: scene detection, ASR and VLM analysis of the source. An optional background research step writes character relationships and plot context into background_research.json, which the README says makes the VLM more likely to identify who is who. Stage two is the agent acting as director and writer, choosing which cuts survive and drafting the narration. Stage three is voiceover through MiMo or Fish Audio. Stage four is assembly: mixing and subtitles, ending at the recap file.

There is a second path the diagram marks as 剪辑模式, cut-first mode. Instead of planning narration against the original timeline, the agent first cuts the long video into a finished short, then writes narration against that output timeline. The README's argument for this is alignment: the narration is written to the cut that exists, so the timeline matches by construction rather than by correction.

Multi-source work is handled through source_id. You pass several videos, segments are selected per source, and they are cut into one recap rather than separate summaries. Analysis results are stored as a filesystem material library of JSON and Markdown plus an index. The README is clear about what that library is not: it does not copy original media, does not create a database and does not do embeddings. Retrieval is the agent searching the filesystem.

## Installing video-recap-skills and running a first recap

Inside Claude Code the README gives two commands, a marketplace add followed by a plugin install. Both are typed as slash commands in the harness, not in a shell.

```text
/plugin marketplace add zenstory-ai/video-recap-skills
/plugin install video-recap-skills@video-recap
```

Before any of that, the local prerequisites have to be in place. The README lists Python 3.10+, ffmpeg on PATH, and a Xiaomi MiMo API key. Subtitles are burned in by default, so the ffmpeg build needs libass and the subtitles filter.

```bash
brew install ffmpeg                        # macOS
sudo apt install ffmpeg                    # Debian / Ubuntu
choco install ffmpeg                       # Windows, also scoop / winget

export MIMO_API_KEY=your-mimo-key          # macOS / Linux
export MIMO_TOKEN_PLAN_CLUSTER=cn          # tp-* key optional: cn | sgp | ams
```

On Windows PowerShell the README substitutes $env:MIMO_API_KEY="your-mimo-key". The MIMO_TOKEN_PLAN_CLUSTER variable only applies to tp-* keys and accepts cn, sgp or ams. Default endpoint is https://api.xiaomimimo.com/v1. The README states that a subscription is not required: sk-* keys bill per use, and it reports about 1.3 CNY for one complete video in its own run, with the caveat that cost varies with duration and call volume.

Once installed, the README suggests letting the agent verify the environment rather than checking by hand:

```text
检查 video-recap 的运行环境，告诉我 Python、ffmpeg/libass 和 MiMo 配置是否就绪。
```

A first real request is a single sentence containing the path, the intent and any background the model cannot infer. The README's own example is a full recap of an episode with burned-in subtitles.

```text
给 /path/to/video.mp4 做一个中文解说成片。这是《庆余年》第一集，主角是范闲，字幕烧进画面。
```

The agent then runs comprehension, planning, cutting, scripting, voicing and assembly, and writes recap_<name>.mp4. Two other request shapes are documented: cutting a long video into roughly ten minutes while keeping key original audio and reactions, and combining two episodes into one ten-minute recap along a single storyline rather than two separate summaries.

## Where the pipeline pushes back: cost, language and the missing API

The most concrete limitation is language. The project produces Chinese narration, the skill prompts shown in the README are Chinese, and the built-in Fish Audio voice is described as an entertainment-commentary timbre. Nothing in the README describes an English narration path. If your audience is not Chinese-speaking, this is the wrong tool, and adapting it means rewriting the skill instructions rather than flipping a flag.

The second constraint is cost shape. The README's figure of roughly 1.3 CNY per complete video comes from the project's own run and is explicitly qualified as varying with duration and call volume. Every ASR, VLM and TTS call goes to a remote service, so a long source or a retry loop multiplies that number. There is no documented offline mode. A MiMo outage or an exhausted quota stops the pipeline, because the local machine contributes only ffmpeg and the Python standard library.

The third is integration surface. The repository's top level contains skills/, tools/, scripts/ and tests/, and the README describes the deliverable as skills invoked through a harness. It does not document a supported Python API for calling the pipeline from your own service. If your requirement is a library you import in a backend job, this is the wrong project.

The MiMo quality review is a fourth case worth naming. The README says it is advisory only: at most one request per stage, it fails open, and it never modifies or blocks the final render. So it will not catch a bad cut for you. It comments, and the render proceeds.

## JianYing draft export and what ffmpeg still decides

The optional JianYing (CapCut) export is the feature that separates this from a one-shot renderer. The README describes a schema-driven multi-track draft where the original footage, narration, BGM, subtitles and local image overlays are all editable. Video, audio and image assets are packed into Resources/local by default with a material index, so the draft still resolves after a clone or a directory move.

The boundary is stated plainly: ffmpeg remains the standard for the final cut. The draft is for continuing work by hand, not for producing the deliverable. That ordering matters if you were hoping the export replaces rendering. It does not; it adds a manual editing step after the automated one.

Export is a separate step that consumes an existing timeline.json. The README's combined request asks for MiMo review both before synthesis and after the finished cut, plus an editable draft. If you only want the draft and not the review, the review is optional and can be left out of the request.

## Alternatives: video-recap-skills versus a general video toolkit skill

The obvious comparison is a general Claude video toolkit skill, the kind that wraps ffmpeg operations so the agent can trim, convert, extract frames or transcode on request. The difference is scope, not quality. A toolkit exposes primitives and leaves the editing decision to you and the model turn by turn; video-recap-skills encodes an opinionated pipeline with named stages, a cut-first mode, a filesystem material library keyed by source_id, and a defined output artifact.

That opinion is the point and also the cost. With a toolkit you can build something the author never imagined, including English narration or a different recap length policy, at the price of designing the flow yourself. With video-recap-skills you get scene detection, ASR, VLM, scripting, TTS and assembly already sequenced, and you inherit its assumptions: Chinese output, MiMo as the default provider, ffmpeg as the arbiter of the final file. If your workflow is already a hand-built ffmpeg chain, adopting this means giving up control of the middle for a faster first draft. Whether that trade is good depends on how often you produce recaps versus how often you need a one-off edit.

## Maintenance, licence and the upgrade surface you inherit

The repository is not archived and the last push was on 2026-09-16, so it is being worked on. The release cadence visible in the release list is roughly monthly to biweekly: v0.3.3 on 2026-06-27, v0.4.0 on 2026-07-26, v0.5.0 on 2026-09-05. The v0.5.0 notes name Fish Audio TTS, boundary validation and three real defect fixes. v0.4.0 covers multi-source cutting and QC, portable JianYing drafts and a content-driven creation flow.

That cadence has a cost. Skills are prompt-and-script artifacts, so an upgrade can change agent behaviour without changing any function signature you could pin. The release titles suggest the author treats defects and boundary checks as first-class work, which is reassuring, but there is no documented compatibility policy for timeline.json or the draft schema across versions. If you build anything on top of the exported draft, verify it after each upgrade rather than assuming the schema held.

The licence is MIT, which is permissive and places few obligations on redistribution or modification. Two caveats sit outside the code licence. The Xiaomi MiMo API and the Fish Audio API each carry their own terms, and the README points to Fish Audio's pricing, rate-limit and commercial-use pages, noting the default s2.1-pro-free model was free only for a period. Check those terms against your use before shipping anything commercial. Nothing here is legal advice.

## Conclusion

Adopt it if you already work inside Claude Code, Codex CLI, OpenCode or OpenClaw and want a recap assembled end to end with ffmpeg on the machine and one MiMo key. Skip it if you need a stable Python API to call from your own service, if your source audio is not Chinese, or if you want an English narration pipeline: the skill prompts and the output are built around Chinese 解说. Before committing, run the environment self-check the README suggests, confirm your ffmpeg build carries libass and the subtitles filter, and read the Fish Audio pricing and commercial-terms pages, since the default s2.1-pro-free model was only free for a period.

## FAQ

### What does video-recap-skills do with a video?

It runs scene detection, ASR and VLM analysis on the source, then plans the story and cuts, writes narration, generates voiceover through MiMo or Fish Audio, and assembles a mixed and subtitled file named recap_<name>.mp4. It can also export an editable JianYing draft, though the README states ffmpeg remains the standard for the final cut.

### Does video-recap-skills need a GPU or local models?

No. The README states no GPU and no model download are required, and that local runtime needs only the Python standard library and ffmpeg. ASR, VLM and TTS all go to Xiaomi MiMo over an API.

### How do I install video-recap-skills in Claude Code?

The README gives two slash commands: /plugin marketplace add zenstory-ai/video-recap-skills, then /plugin install video-recap-skills@video-recap. Codex CLI, OpenCode and OpenClaw each have their own documented install path.

### What video creation skills does Claude have?

In this project's case, six Agent Skills are registered, and the README says daily end-to-end production uses video-recap while video-script covers planning or script writing alone, with the remaining four handling tool stages.

## Sources

- [License: MIT](https://github.com/zenstory-ai/video-recap-skills/blob/main/LICENSE)
- [Project website](https://zenstory.ai/video-recap)
- [README](https://github.com/zenstory-ai/video-recap-skills/blob/main/README.md)
- [Releases](https://github.com/zenstory-ai/video-recap-skills/releases)
- [zenstory-ai/video-recap-skills on GitHub](https://github.com/zenstory-ai/video-recap-skills)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zenstory-ai-video-recap-skills
