# Seven confirmation steps before a single frame is rendered

> srt-whiteboard-animation is a Chinese-language skill package that turns SRT subtitles into a hand-drawn whiteboard animation, one approval step at a time. The design lives in an annotation file, and the renderer is told what to draw by flags rather than by code.

**geeklee/srt-whiteboard-animation** — 将 SRT 字幕做成暖米黄纸张底的流式笔迹白板手绘动画 skill：mask 分区遮罩编排 + stream 连续笔迹（ink→color）。

- Repository: https://github.com/geeklee/srt-whiteboard-animation
- Stars: 3,975 · Forks: 638
- Language: Python
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/geeklee-srt-whiteboard-animation

## The documentation is Chinese only and the release title is too

The README of this repository is written in Chinese, from the title SRT 白板动画 Skill down to the closing lines about the author, and the tree contains no translated copy of it. That is worth stating at the top because almost everything else a reader needs is a script name, a flag or a JSON key, and those are the same in either language. The project's own description in the repository metadata is also Chinese: it describes turning SRT subtitles into a stream-pen-stroke whiteboard animation on a warm beige paper background, with mask based region orchestration and ink followed by color. The single release is titled in Chinese as well, v1.0.0, described as the first public release of the SRT whiteboard animation skill, and it was published on 2026-07-27.

## Seven approval steps stand between a subtitle and an MP4

The workflow section is the heart of the document and its stated key is 字幕驱动、逐步确认, subtitle driven and confirmed step by step. The reason given is cost: each step waits for confirmation so that rendering money is not spent while the storyboard, the line art or the annotations are still unfinished. The seven steps are parse the SRT and propose a storyboard, generate line art in one consistent style after approval, create annotations from the subtitles and the original image and load the preview console, generate region and direction check images, adjust regions, narrative order, timing and subtitle association in the console and save, render each scene to MP4, and merge the scenes once each is approved. Six human checkpoints for a video, which is a lot for a tool whose selling point is automation.

## Scene length has a target and two bounds

The parser is told how long a scene should be with three separate numbers:

```bash
python scripts/parse_srt.py <字幕.srt> --target-sec 30 --min-sec 25 --max-sec 35
```

The capability list describes the same thing as splitting scenes on a suggested 25 to 35 seconds, and the flag names show the intent more precisely: 30 is the target the splitter aims for, while 25 and 35 are the bounds it will not cross. A subtitle file is therefore expected to be narratable in half-minute chunks, and the output of this step is a proposal rather than a decision, since the workflow requires the storyboard to be confirmed before line art is generated. The next step in that proposal is the illustration strategy, with the rule that each scene expresses only one core idea.

## The annotation file holds the actual design

Almost everything the renderer needs arrives as annotation.json, and the format section specifies it in detail. Each element uses integer pixel coordinates from the original image and is tied to a subtitle event through three fields: sequence for drawing order, subtitle for the line it corresponds to, and narrativeRole for what it does in the story. Regions are expected to be ordered by a fixed progression, scene setup, then key character or object, then action or change, then reaction or result. The reveal block carries direction, startMs, durationMs, maskPaddingPx and protectedRegions, and protectedRegions is how one element declares the area a later element must not appear in early. direction and handPath exist only for the rectangle proxy the preview console draws; the real strokes in the finished video are generated by the stream renderer.

## Rendering is chosen by two flags and one hand image

The render step takes an image, an annotation file, an output path and a hand asset, then two flags that decide how the drawing looks: `--ink-path grid` for the line art and `--color-fill contour-wipe` for adding colour. The quality check list names a third option, `--ink-path skeleton`, to be used when the line art is clear enough for it. The check list is also the closest thing to a specification of what correct output means: the first frame must be clean beige with no lines showing early, canvas must match the original image size, every region must be integer pixels inside the canvas, sequence and startMs must agree with the subtitle order, regions that have not started and protected areas must not appear in middle frames, the pen tip should stay close to the current stroke, and each scene must hold a full frame for at least half a second at the end.

## The environment is a script that prints ENV_PY

Dependency isolation is handled by a preparation script rather than by a lock file or a requirements file, and the first two commands are:

```bash
python scripts/prepare_env.py --check
python scripts/prepare_env.py
```

The first is a check and the second actually prepares. The document then says that on success the first command outputs `ENV_PY=<路径>`, and that later rendering should use that interpreter to keep dependencies isolated. That convention shows up in the render and merge commands, which both start with `<ENV_PY>` rather than with python, so a renderer that skips the preparation step has no interpreter to substitute and no obvious way to know which one is meant.

## Two file naming families for what looks like one example scene

The asset layout is prescribed per project directory:

```
assets/whiteboard/<项目名>/
├── scene-01-<名称>.png
├── scene-01-<名称>.annotation.json
├── scene-01-<名称>-whiteboard.mp4
└── scene-01-<名称>-preview.mp4
```

The rule that follows is that the image and the annotation must share a name, with scene-01-demo.png matching scene-01-demo.annotation.json. The example directory does not quite follow one naming family. It holds scene-01-monkey-mountain.png, which is the file the README links as the original line art, and separately scene-01-monkey-mountain-banana.png with its own annotation file, plus two animated outputs, one named after the whiteboard variant and one after the stream variant. So the single worked example in the repository spans two base names, and anyone copying it has to decide which one is the canonical pair.

## A skill package with a Codex metadata file and one release

The repository is packaged as a skill rather than as a library. SKILL.md sits at the root holding the complete workflow and its constraints, and the tree listing in the README describes agents/openai.yaml as Codex metadata. There is no Python package to install: scripts/ holds five files, parse_srt.py for subtitles and storyboard suggestions, render_annotation_preview.py for annotation check images, render_stream_whiteboard.py as the stream stroke renderer, merge_scenes.py for combining scenes, and prepare_env.py for the environment. Assets hold the hand image and preview.html, the local editing console, which is opened directly in a browser and loads a scene directory through an open folder control. The release history is one entry, v1.0.0, and the contributing section asks that any change to the drawing logic be checked against real subtitles, annotations and finished video for mask protection, timing and final frames.

## Conclusion

This package suits a creator who wants a narrated explainer drawn to order and is willing to sit through seven approval rounds, since the workflow exists to protect rendering cost rather than to automate it. It does not suit batch work: the subtitle parser suggests a scene split, but every scene then waits for a person. Before starting, check three things. The documentation is Chinese only. The output carries no on-screen text, because the visual rules forbid scene text and labels even though subtitles drive the timing. And the environment is created by a script that prints ENV_PY, which every render command then expects you to use.

## FAQ

### What does the srt-whiteboard-animation project do?

It turns SRT subtitles into a whiteboard hand-drawn animation on a warm beige paper background, following the narrative order of the subtitles. Each element enters in turn, the pen lays continuous ink inside a region, colour is added progressively, and the result is exported as MP4.

### Is there an English version of the srt-whiteboard-animation documentation?

No. The README is written in Chinese and the repository tree contains no translated copy of it. The single release is titled in Chinese too, v1.0.0, described as the first public release of the SRT whiteboard animation skill.

### What does the annotation file control?

Regions, timing, subtitle association and overlap protection. Each element carries integer pixel coordinates from the original image plus sequence, subtitle and narrativeRole, and a reveal block holding direction, startMs, durationMs, maskPaddingPx and protectedRegions, which is where an element declares the area a later element must not appear in early.

### How long should each scene be?

The README suggests 25 to 35 seconds per scene, and the parser takes three numbers for it: --target-sec 30 as the aim, with --min-sec 25 and --max-sec 35 as the bounds it will not cross.

### How is the Python environment prepared for rendering?

By running `python scripts/prepare_env.py --check` and then `python scripts/prepare_env.py`. The first command prints `ENV_PY=<path>`, and the render and merge commands are written to be run with that interpreter so dependencies stay isolated.

### How are several scenes merged into one video?

Render each scene first with render_stream_whiteboard.py, then run `<ENV_PY> scripts/merge_scenes.py --inputs scene1.mp4 scene2.mp4 scene3.mp4 --output final.mp4`. The quality check list requires the merge order to match the subtitle storyboard and each scene to hold a full frame for at least 0.5 seconds at its end.

## Sources

- [geeklee/srt-whiteboard-animation on GitHub](https://github.com/geeklee/srt-whiteboard-animation)
- [Issues](https://github.com/geeklee/srt-whiteboard-animation/issues)
- [License: MIT](https://github.com/geeklee/srt-whiteboard-animation/blob/main/LICENSE)
- [README](https://github.com/geeklee/srt-whiteboard-animation/blob/main/README.md)
- [Releases](https://github.com/geeklee/srt-whiteboard-animation/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/geeklee-srt-whiteboard-animation
