Seven confirmation steps before a single frame is rendered
将 SRT 字幕做成暖米黄纸张底的流式笔迹白板手绘动画 skill:mask 分区遮罩编排 + stream 连续笔迹(ink→color)。
At a glance
- What is it?
- srt-whiteboard-animation is a Chinese-language skill package that turns SRT subtitles into a hand-drawn whiteboard animation, one approval step at a time. The design lives in an annotation file, and the renderer is told what to draw by flags rather than by code.
- Who is it for?
- This package suits a creator who wants a narrated explainer drawn to order and is willing to sit through seven approval rounds, since the workflow exists to protect rendering cost rather than to automate it. It does not suit batch work: the subtitle parser suggests a scene split, but every scene then waits for a person.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 70 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The documentation is Chinese only and the release title is too
The README of this repository is written in Chinese, from the title SRT 白板动画 Skill down to the closing lines about the author, and the tree contains no translated copy of it. That is worth stating at the top because almost everything else a reader needs is a script name, a flag or a JSON key, and those are the same in either language. The project's own description in the repository metadata is also Chinese: it describes turning SRT subtitles into a stream-pen-stroke whiteboard animation on a warm beige paper background, with mask based region orchestration and ink followed by color. The single release is titled in Chinese as well, v1.0.0, described as the first public release of the SRT whiteboard animation skill, and it was published on 2026-07-27.
Seven approval steps stand between a subtitle and an MP4
The workflow section is the heart of the document and its stated key is 字幕驱动、逐步确认, subtitle driven and confirmed step by step. The reason given is cost: each step waits for confirmation so that rendering money is not spent while the storyboard, the line art or the annotations are still unfinished. The seven steps are parse the SRT and propose a storyboard, generate line art in one consistent style after approval, create annotations from the subtitles and the original image and load the preview console, generate region and direction check images, adjust regions, narrative order, timing and subtitle association in the console and save, render each scene to MP4, and merge the scenes once each is approved. Six human checkpoints for a video, which is a lot for a tool whose selling point is automation.
Scene length has a target and two bounds
The parser is told how long a scene should be with three separate numbers:
python scripts/parse_srt.py <字幕.srt> --target-sec 30 --min-sec 25 --max-sec 35The capability list describes the same thing as splitting scenes on a suggested 25 to 35 seconds, and the flag names show the intent more precisely: 30 is the target the splitter aims for, while 25 and 35 are the bounds it will not cross. A subtitle file is therefore expected to be narratable in half-minute chunks, and the output of this step is a proposal rather than a decision, since the workflow requires the storyboard to be confirmed before line art is generated. The next step in that proposal is the illustration strategy, with the rule that each scene expresses only one core idea.
The annotation file holds the actual design
Almost everything the renderer needs arrives as annotation.json, and the format section specifies it in detail. Each element uses integer pixel coordinates from the original image and is tied to a subtitle event through three fields: sequence for drawing order, subtitle for the line it corresponds to, and narrativeRole for what it does in the story. Regions are expected to be ordered by a fixed progression, scene setup, then key character or object, then action or change, then reaction or result. The reveal block carries direction, startMs, durationMs, maskPaddingPx and protectedRegions, and protectedRegions is how one element declares the area a later element must not appear in early. direction and handPath exist only for the rectangle proxy the preview console draws; the real strokes in the finished video are generated by the stream renderer.
Rendering is chosen by two flags and one hand image
The render step takes an image, an annotation file, an output path and a hand asset, then two flags that decide how the drawing looks: `--ink-path grid` for the line art and `--color-fill contour-wipe` for adding colour. The quality check list names a third option, `--ink-path skeleton`, to be used when the line art is clear enough for it. The check list is also the closest thing to a specification of what correct output means: the first frame must be clean beige with no lines showing early, canvas must match the original image size, every region must be integer pixels inside the canvas, sequence and startMs must agree with the subtitle order, regions that have not started and protected areas must not appear in middle frames, the pen tip should stay close to the current stroke, and each scene must hold a full frame for at least half a second at the end.
The environment is a script that prints ENV_PY
Dependency isolation is handled by a preparation script rather than by a lock file or a requirements file, and the first two commands are:
python scripts/prepare_env.py --check
python scripts/prepare_env.pyThe first is a check and the second actually prepares. The document then says that on success the first command outputs `ENV_PY=<路径>`, and that later rendering should use that interpreter to keep dependencies isolated. That convention shows up in the render and merge commands, which both start with `<ENV_PY>` rather than with python, so a renderer that skips the preparation step has no interpreter to substitute and no obvious way to know which one is meant.
Two file naming families for what looks like one example scene
The asset layout is prescribed per project directory:
assets/whiteboard/<项目名>/
├── scene-01-<名称>.png
├── scene-01-<名称>.annotation.json
├── scene-01-<名称>-whiteboard.mp4
└── scene-01-<名称>-preview.mp4The rule that follows is that the image and the annotation must share a name, with scene-01-demo.png matching scene-01-demo.annotation.json. The example directory does not quite follow one naming family. It holds scene-01-monkey-mountain.png, which is the file the README links as the original line art, and separately scene-01-monkey-mountain-banana.png with its own annotation file, plus two animated outputs, one named after the whiteboard variant and one after the stream variant. So the single worked example in the repository spans two base names, and anyone copying it has to decide which one is the canonical pair.
A skill package with a Codex metadata file and one release
The repository is packaged as a skill rather than as a library. SKILL.md sits at the root holding the complete workflow and its constraints, and the tree listing in the README describes agents/openai.yaml as Codex metadata. There is no Python package to install: scripts/ holds five files, parse_srt.py for subtitles and storyboard suggestions, render_annotation_preview.py for annotation check images, render_stream_whiteboard.py as the stream stroke renderer, merge_scenes.py for combining scenes, and prepare_env.py for the environment. Assets hold the hand image and preview.html, the local editing console, which is opened directly in a browser and loads a scene directory through an open folder control. The release history is one entry, v1.0.0, and the contributing section asks that any change to the drawing logic be checked against real subtitles, annotations and finished video for mask protection, timing and final frames.
Editorial conclusion
This package suits a creator who wants a narrated explainer drawn to order and is willing to sit through seven approval rounds, since the workflow exists to protect rendering cost rather than to automate it. It does not suit batch work: the subtitle parser suggests a scene split, but every scene then waits for a person. Before starting, check three things. The documentation is Chinese only. The output carries no on-screen text, because the visual rules forbid scene text and labels even though subtitles drive the timing. And the environment is created by a script that prints ENV_PY, which every render command then expects you to use.
Frequently asked questions
What does the srt-whiteboard-animation project do?
It turns SRT subtitles into a whiteboard hand-drawn animation on a warm beige paper background, following the narrative order of the subtitles. Each element enters in turn, the pen lays continuous ink inside a region, colour is added progressively, and the result is exported as MP4.
Is there an English version of the srt-whiteboard-animation documentation?
No. The README is written in Chinese and the repository tree contains no translated copy of it. The single release is titled in Chinese too, v1.0.0, described as the first public release of the SRT whiteboard animation skill.
What does the annotation file control?
Regions, timing, subtitle association and overlap protection. Each element carries integer pixel coordinates from the original image plus sequence, subtitle and narrativeRole, and a reveal block holding direction, startMs, durationMs, maskPaddingPx and protectedRegions, which is where an element declares the area a later element must not appear in early.
How long should each scene be?
The README suggests 25 to 35 seconds per scene, and the parser takes three numbers for it: --target-sec 30 as the aim, with --min-sec 25 and --max-sec 35 as the bounds it will not cross.
How is the Python environment prepared for rendering?
By running `python scripts/prepare_env.py --check` and then `python scripts/prepare_env.py`. The first command prints `ENV_PY=<path>`, and the render and merge commands are written to be run with that interpreter so dependencies stay isolated.
How are several scenes merged into one video?
Render each scene first with render_stream_whiteboard.py, then run `<ENV_PY> scripts/merge_scenes.py --inputs scene1.mp4 scene2.mp4 scene3.mp4 --output final.mp4`. The quality check list requires the merge order to match the subtitle storyboard and each scene to hold a full frame for at least 0.5 seconds at its end.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/geeklee-srt-whiteboard-animation)