# hand-drawn-explainer-video-nikola: a Codex Skill for Chinese stroke-by-stroke explainer videos

> A Python-based Codex Skill that turns a topic, a script or an SRT into a hand-drawn Chinese explainer video, with two production routes and a hard rule against passing off pans of static art as real drawing. The install is short, but the voiceover and rendering dependencies are where projects stall.

**hi-nikola/hand-drawn-explainer-video-nikola** — 中文手绘知识讲解视频 Codex Skill：逐笔故事、双语义岛、让怪诞小黑动起来与程序动画

- Repository: https://github.com/hi-nikola/hand-drawn-explainer-video-nikola
- Stars: 362 · Forks: 44
- Language: Python
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/hi-nikola-hand-drawn-explainer-video-nikola

## The gap this fills: Chinese explainer videos that are actually drawn, not panned

Most tools that advertise a hand-drawn look animate a finished illustration. The picture slides, scales or fades in, and the pen never touches the page. This repository takes the opposite position in writing: it states that it does not pass off a pan of a hand-drawn still as stroke-by-stroke drawing. That sentence is the whole reason the project exists.

The intended user is someone producing Chinese knowledge explainers who wants the drawing to happen on screen, one stroke at a time, while a voice explains. The README frames the output as a real MP4 plus subtitles, a timeline and editable source assets, not a preview clip. It is a Codex Skill, so the workflow is driven by prompts inside Codex rather than by a standalone GUI.

There is a second audience: people who only want prompts. The repository can emit self-contained prompts for Flow, Nano Banana or other image and video models. The README is careful to call that a delivery scope and not a third production route, which is a distinction worth keeping straight when you plan work.

## Two routes, and the picture structure is not a third one

The README lays out exactly two production routes. The first is stroke-by-stroke story animation. The second is program animation built from HTML, SVG and GSAP. Everything else is a modifier on the first route.

Picture structure and visual style are described as composable options inside the stroke-by-stroke route. You can use a single scene, a multi-act story, or the left-right dual semantic island arrangement, where the left side is drawn and explained first, then the right side is drawn while the earlier content stays on screen. Visual style is separate: Q-version characters suit biographies and historical figures, while the Xiaohei style suits abstract arguments, methods and metaphors. Both styles work with either picture structure.

The program animation route is a different mechanism entirely. Elements such as cards, arrows, magnifiers and subtitle bars are composed as independent HTML and SVG objects and driven along the real narration timeline by GSAP and HyperFrames. The README's comparison table is blunt about the boundary: a hand-drawn-looking element animation is not real stroke drawing, and an image fade, a flying card or an SVG motion must not be presented as ink being laid down.

## Installing the Skill and running the first check

The README gives two installation paths: let Codex install the repository, or clone it manually into the Codex skills directory. On Windows the clone target is the user profile's .codex/skills folder, and the second command runs the setup check script.

```powershell
git clone https://github.com/hi-nikola/hand-drawn-explainer-video-nikola.git "$env:USERPROFILE/.codex/skills/hand-drawn-explainer-video-nikola"
python "$env:USERPROFILE/.codex/skills/hand-drawn-explainer-video-nikola/scripts/setup_check.py"
```

After that, say the README, prompt mode is already usable. Scripts need Python 3.10 or newer. On Windows, if the default python is older, the README suggests py -3.12 or another interpreter you have confirmed.

Stroke-by-stroke video needs two more steps: preparing the vendored environment and running the preflight with a report file.

```powershell
cd "$env:USERPROFILE/.codex/skills/hand-drawn-explainer-video-nikola"
python vendor/srt-whiteboard-animation/scripts/prepare_env.py
python scripts/stroke_story_preflight.py --report preflight-stroke-story.json
```

Full MP4 output requires FFmpeg and FFprobe. Program animation requires Node.js, a browser and HyperFrames. The README points to docs/INSTALL.md, docs/CONFIGURATION.md and docs/TROUBLESHOOTING.md for the rest. A first real use is to paste one of the three trigger prompts from the README, for example the 45-second Chinese explainer prompt that asks for left-then-right drawing, Q-version characters, a representative shot first, and then a real MP4, SRT, timeline and editable assets.

## What the stroke-by-stroke backend actually does

The drawing runtime is vendored under vendor/srt-whiteboard-animation/ with an MIT licence and a provenance note, and the README states it contains no second Skill. According to the README, the technique is bitmap line-art extraction plus skeleton and stream continuous strokes plus regional masks. The canvas is drawn by semantic region: line work first, colour added afterwards, and content already drawn stays visible.

That regional mask idea is what makes the dual semantic island layout possible. Instead of clearing the frame between beats, the renderer keeps earlier regions and extends the drawing into a new one. The README also notes that the core stroke rendering does not need a downloaded local neural network, model weights or a GPU; with a usable source image or line art and narration, it can be finished with local traditional image algorithms. Network access is needed the first time Python or npm dependencies are installed, and external services are needed when generating new illustrations from a script or synthesising Liu Fei's voice.

## Voiceover is the dependency that decides whether you finish

The default narration is Volcengine speech synthesis using the Liu Fei voice, zh_male_liufei_uranus_bigtts, with the seed-tts-2.0 resource model. The README instructs users to open the Doubao speech synthesis 2.0 service in the Volcengine console, and includes a screenshot of where that toggle lives. The .env.example file in the repository root carries a single optional key, VOLCENGINE_TTS_API_KEY, left empty unless you choose that provider.

This is the sharpest constraint in the project. The README says system speech, edge-tts and other low-quality local text-to-speech should not be quietly treated as a finished substitute, because they differ noticeably from the verified Liu Fei output. If there is no Volcengine authorisation and no acceptable narration supplied by the user, the instruction is to keep the finished picture, subtitles and timeline, state the narration gap, and not silently swap in a system voice while claiming completion.

That is an honest rule and an inconvenient one. A user with no Volcengine access gets a partially delivered project rather than a finished video, by design. The repository also warns that cloud narration, image generation and video services may cost money, and advises a dry run, reuse of matching cache, and no blind retries.

## Where it is the wrong tool, and what to use instead

If your content is a flowchart, a rule set, a comparison or anything with exact labels and numbers, the stroke-by-stroke route is the wrong choice and the README says so through its own routing table: flow cards, relationship diagrams, precise labels and independent element motion all point to program animation. Trying to draw a precise number stroke by stroke wastes time and produces worse typography than deterministic SVG text.

The natural alternative for that job is the same repository's program animation route rather than a different project, but if you want a purpose-built whiteboard workbench instead of a Codex Skill, the README's acknowledgements name ChenShuo2004/cs-board for a complete whiteboard video workbench with skeleton strokes and a production pipeline. The difference in approach is structural: cs-board is a workbench you operate, while this project is a Skill you invoke from prompts inside Codex and which returns a defined deliverable set. For story-to-comic conversion and line-art/colour pairing, the acknowledgements also point to gnipbao/story-to-handdrawn-video, and for SRT semantic regions and continuous canvas strokes, to geeklee/srt-whiteboard-animation, which is the upstream of the vendored runtime here.

One more boundary: the README says the repository does not include the 7 minute 29 second full commercial cut referenced in its case notes. The 36 second project in the repository is a trimmed public sample for download and study.

## Licence split, upgrade cost and what maintenance looks like

Licensing is layered, and the layers matter. The Skill, original scripts and documentation are Apache-2.0. The vendored runtime under vendor/srt-whiteboard-animation/ keeps its upstream MIT licence. Example media is offered under a CC BY 4.0 media note, limited to the parts the author has the right to license. Third-party software, personal names, product names and trademarks are not re-licensed by any of this, and THIRD_PARTY_NOTICES.md is where the boundaries are recorded. None of that is legal advice; if you plan to publish commercially, read the three licence files and the third-party notices yourself.

On maintenance, the last push to the default branch was on 2026-09-03, the same day v1.1.0 was released; v1.0.0 was released on 2026-09-02. The repository is not archived. Those are the only maintenance facts available here, and they say nothing about how quickly issues are answered.

Upgrade cost is mostly environmental rather than code. The stroke runtime is vendored, so moving to a newer upstream means re-vendoring rather than a package bump. The program animation route pulls GSAP through npm install because the browser file is not committed, for third-party licence reasons, so a fresh checkout needs network access before it can render. Python is pinned at 3.10 or newer, and the media requirements live in requirements-media.txt.

## Conclusion

Adopt it if you already work inside Codex, need Chinese narration and want deliverables you can edit afterwards: MP4, SRT, timeline and source assets. Do not adopt it if you have no Volcengine speech authorisation and no acceptable narration of your own, because the README explicitly refuses to substitute a system voice and call the job done. Before committing to a full video, run the preflight report and produce one representative shot with the Q-version or Xiaohei style you intend to use, since the repository's own notes say an earlier 9:16 single-canvas sample was dropped because character, prop and background lines interfered with each other.

## FAQ

### What is an explanatory video in the context of hand-drawn-explainer-video-nikola?

In this project it is a Chinese knowledge explainer where part of the content is spoken and part is drawn on screen, delivered as a real MP4 with subtitles, a timeline and editable assets. The repository distinguishes that from a pan or zoom over a finished hand-drawn still, which it refuses to present as stroke-by-stroke drawing.

### How do I install hand-drawn-explainer-video-nikola?

Clone the repository into the Codex skills directory and run scripts/setup_check.py, which the README shows as a PowerShell one-liner pair. Prompt mode works after that; stroke-by-stroke video additionally needs prepare_env.py from the vendored runtime and a preflight report, and full MP4 output needs FFmpeg and FFprobe.

### Does hand-drawn-explainer-video-nikola need a GPU or downloaded model weights?

The README states that the core stroke rendering does not require downloading a local neural network or model weights and does not require a GPU. With a usable source image or line art and narration, it says the work can be completed with local traditional image algorithms.

### Can I use a system voice instead of Volcengine in hand-drawn-explainer-video-nikola?

The README advises against automatically substituting system speech, edge-tts or other low-quality local text-to-speech, saying they differ noticeably from the verified Liu Fei voice. If no Volcengine authorisation and no acceptable narration are available, it says to keep the picture, subtitles and timeline and state the narration gap rather than silently swapping in a system voice.

## Sources

- [hi-nikola/hand-drawn-explainer-video-nikola on GitHub](https://github.com/hi-nikola/hand-drawn-explainer-video-nikola)
- [Issues](https://github.com/hi-nikola/hand-drawn-explainer-video-nikola/issues)
- [License: Apache-2.0](https://github.com/hi-nikola/hand-drawn-explainer-video-nikola/blob/main/LICENSE)
- [README](https://github.com/hi-nikola/hand-drawn-explainer-video-nikola/blob/main/README.md)
- [Releases](https://github.com/hi-nikola/hand-drawn-explainer-video-nikola/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/hi-nikola-hand-drawn-explainer-video-nikola
