Model or dataset
Mr-funny/hbg-classical-poem-silk-video avatar
Mr-funny/hbg-classical-poem-silk-video

hbg-classical-poem-silk-video: an Agent Skill that turns classical Chinese poems into vertical silk-scroll videos

Agent Skill for turning Chinese classical poems into vertical Chinese-art videos with ImageGen stills, Docker I2V, calligraphy captions, retained ambience, BGM and final MP4 QA.

363 stars54 forksShellMIT

At a glance

What is it?
The Skill splits a poem into one image per line (or per couplet), generates stills with the agent's built-in ImageGen, animates them through a Docker I2V runtime, and writes calligraphy captions over retained ambience. Its real subject is restraint: locked camera, one action per scene, and QA on the encoded MP4 rather than the preview.
Who is it for?
Adopt it if you already run Codex or Claude Code, keep the hbg-gemini-flow-suite Docker runtime available, and want per-line poem scenes where motion comes from the painted objects instead of a post-production zoom. Skip it if you cannot complete the runtime's one-time user-controlled authorization, if you need a fully offline pipeline, or if your footage is not a Chinese painting in the first place.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 59 days ago.
What is it written in?
Mainly Shell, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem: painted scenes that move too much

Image-to-video models are generous with motion. Point one at a Chinese painting and the mist thickens into auspicious clouds, a temple gets rebuilt between frames, birds multiply, and a walking horse grows a fifth leg. The README lists these as the recurring failures the Skill is built around, and its framing is blunt: the goal is not that the painting moves, it is that it does not move arbitrarily.

The project is an Agent Skill, a directory of instructions and scripts that a coding agent loads on demand. It targets people who already write prompts inside Codex or Claude Code and want a repeatable procedure for a specific output: a 1080x1920 vertical video in which each line of a poem becomes one scene, captions are written character by character in a brush typeface, and the ambience generated alongside each shot is kept rather than replaced. The repository is Shell-first, MIT licensed, and its only release so far is v0.1.0, tagged on 2026-08-03, the same day as the last push to main.

How the pipeline runs from poem to MP4

The README's flowchart is the clearest statement of the data flow: full poem, then imagery and emotional arc, then a split into one scene per line or per couplet, then a Chinese-painting style channel, then ImageGen stills, then Docker I2V, then frame sampling at the start, middle and end of each shot, then removal of the white model watermark, then character-by-character calligraphy captions, then original ambience plus BGM, then cross-dissolves for both picture and sound, then final MP4 QA.

Four rules govern the motion prompts. The camera is locked, so movement must come from objects already in the painting. Actions are anchored: wings stay around a body centre, ripples spread from a contact point, willow leaves do not leave the branch. Each scene gets exactly one primary action, because several simultaneous actions push the model into recomposing the frame. And the deliverable is the encoded MP4, not the browser preview.

The style layer is where the project differs from a single-look template. Under one Chinese-painting quality bar it switches channels per scene: ink wash, gongbi fine-line, blue-green shanshui, figure-and-horse. The documented example, Qian Tang Hu Chun Xing, uses four scenes from eight lines, each about nine seconds, joined by three 1.2-second dissolves, at 1080x1920, 30fps, H.264 with AAC, for a 32.44-second film.

Installing the Skill into Codex or Claude Code

The README offers four routes. The recommended one is to paste a natural-language install prompt into an agent that supports SKILL.md and let it find the global skills directory, back up any older copy, verify that SKILL.md, agents, references, scripts and font assets are present, and run shell syntax, Python compile and skill validation checks. Two scripted routes follow. For Codex:

bash
curl -fsSL https://raw.githubusercontent.com/Mr-funny/hbg-classical-poem-silk-video/main/install.sh | sh

That lands the Skill in ${CODEX_HOME:-~/.codex}/skills/classical-poem-silk-video. For Claude Code, the same installer takes a flag:

bash
curl -fsSL https://raw.githubusercontent.com/Mr-funny/hbg-classical-poem-silk-video/main/install.sh | sh -s -- --claude

The manual route clones the repository and copies the Skill directory into place:

bash
git clone https://github.com/Mr-funny/hbg-classical-poem-silk-video.git
mkdir -p ~/.codex/skills/classical-poem-silk-video
cp -R hbg-classical-poem-silk-video/skill/classical-poem-silk-video/. \
  ~/.codex/skills/classical-poem-silk-video/

The stills come from the agent's built-in ImageGen, so no separate image service is needed. Animation does need the Docker runtime. The README names Mr-funny/hbg-gemini-flow-suite, container name gemini-flow-suite, with the workspace mounted at /workspace and outputs at /data/outputs. You complete a one-time user-controlled authorization in that runtime first, then check the ground:

bash
skill/classical-poem-silk-video/scripts/check_prerequisites.sh

After that, a first real request is just a sentence to the agent, naming the poem, the scene split, the colour arc and the motion sources. The README's example asks for Feng Qiao Ye Bo, one scene per line, night shifting from cold blue to the warm of a fishing lamp, no post-production zoom, with motion attributed to moonlight, ripples, crows, frost mist and the visible cause of a bell.

QA on the encoded file, not the preview

The most opinionated part of the project is where it puts verification. The README states plainly that the final encoded MP4 is the delivery object and that frames must be re-extracted from it, because a browser preview can look clean while the encoded file carries black frames. The QA script takes a video and an output directory:

bash
skill/classical-poem-silk-video/scripts/final_media_qa.sh \
  final.mp4 qa/final

It writes ffprobe.json, blackdetect.log, silencedetect.log and volumedetect.log, and extracts frames at 10, 25, 50, 75 and 90 percent plus the last frame, along with a contact sheet. For targeted checks such as hoof contact or bird count at a transition midpoint, you supply a TSV of timestamps and labels as a third argument. This is a real constraint, not a formality: inspecting a dissolve midpoint means pulling a frame the sampling percentages would miss.

Two details in the QA list are worth flagging because they are easy to get wrong. The white model watermark in the lower right is treated separately from a traditional red seal, which is kept as picture content. And the caption font, Ma Shan Zheng, is bundled, with the right column written first and the left column after, characters revealed one at a time.

Where the Skill stops being the right tool

The runtime boundary is the first limit. Image-to-video and the optional watermark cleanup depend on the hbg-gemini-flow-suite container, which is a separate project with its own licence and is not redistributed here. If you cannot run Docker, or cannot complete the runtime's user-controlled authorization, the pipeline stops at stills. The README is explicit that the Skill itself does not include, upload or print cookies, API keys or browser profiles, and that generation only calls Docker, but that also means the Skill is not self-contained.

The second limit is subject matter. Every motion rule assumes a Chinese painting with recognisable anchors: architecture, mountain forms, tree trunks, shorelines, a human torso. Applied to photographic footage or to a Western illustration, the instruction to lock the camera and animate only existing objects is not obviously useful, and the style-channel vocabulary has nothing to map onto.

Third, the QA is inspection, not repair. The README's repair example asks the Skill to check four I2V segments for rebuilt architecture, duplicated birds, malformed horse legs, falling willow leaves, white watermark, caption safe area and black-frame transitions, then redo only the failing shots. Nothing in the repository suggests the checks are automated pass or fail gates; the scripts produce logs and frames for a human or an agent to read. If you expect a single command that guarantees a clean render, this is not that.

Compared with assembling the same video by hand

The obvious alternative is doing it yourself in ffmpeg and a diffusion model: generate stills, run each through an I2V endpoint, then concatenate with xfade and mix audio. That approach gives you total control and no dependency on a Docker suite. It also gives you no motion vocabulary. The four-part prompt structure documented here, static anchors, a local action region, a stable ending, and explicit anti-hallucination prohibitions, is the part you would otherwise reinvent per shot, and the failure list at the top of the README is a fair description of what happens when you do not.

A second alternative is a slideshow tool that applies a slow Ken Burns push to a single painting. The README calls this out directly, describing the Skill as not a template for slowly pushing a camera over an old painting. The difference is where the motion originates: a push is applied to the whole frame after the fact, while this pipeline asks the model to move named objects inside the frame and then verifies that the frame was not rebuilt. That distinction is the whole product, and it is also why the QA step samples specific timestamps instead of just checking that the file plays.

Maintenance, licensing and what a fork inherits

The repository was pushed on 2026-08-03 and is not archived, so it is recent rather than proven. There is one release, v0.1.0, and the README invites issues and pull requests. No upgrade path, migration guide or rollback procedure is documented; the installer's only stated concession to versioning is that the agent should back up an existing copy before updating. If you pin the Skill for a production workflow, that backup step is the mechanism you have, and you should keep your own copy of a working skills directory.

Licensing is split three ways, and the split matters. The Skill, its scripts and its documentation are MIT. The bundled Ma Shan Zheng font is under the SIL Open Font License 1.1, with the text at skill/classical-poem-silk-video/assets/OFL.txt, which is a different licence with its own conditions on redistribution. The hbg-gemini-flow-suite runtime is a separate project under its own licence and is not redistributed in this repository. If you ship a video made with the pipeline, the output is yours to reason about, but if you repackage the Skill, the font travels with a licence that is not MIT. The README also states that only the white model watermark is touched, and only when the user asks and applicable terms allow it, while traditional red seals and intentional marks must be preserved. That is a policy statement in the project, not legal advice.

Editorial conclusion

Adopt it if you already run Codex or Claude Code, keep the hbg-gemini-flow-suite Docker runtime available, and want per-line poem scenes where motion comes from the painted objects instead of a post-production zoom. Skip it if you cannot complete the runtime's one-time user-controlled authorization, if you need a fully offline pipeline, or if your footage is not a Chinese painting in the first place. Before committing, run check_prerequisites.sh, confirm SKILL.md, agents, references, scripts and the font assets all landed in your skills directory, and read runtime-contract.md to see exactly what crosses the Docker boundary.

Frequently asked questions

What is hbg-classical-poem-silk-video?

It is an Agent Skill for Codex, Claude Code and other agents that support SKILL.md. It turns a Chinese classical poem into a 1080x1920 Chinese-art video by splitting the text into one scene per line or per couplet, generating stills with the agent's built-in ImageGen, animating them through a Docker I2V runtime, and adding brush-style captions over retained ambience.

How do I install hbg-classical-poem-silk-video into Codex?

The README gives a one-line installer that pipes install.sh to sh, which places the Skill in ${CODEX_HOME:-~/.codex}/skills/classical-poem-silk-video. For Claude Code the same script takes the --claude flag, and a manual route clones the repository and copies the skill directory into your skills folder.

Does hbg-classical-poem-silk-video need Docker?

Yes for the animation stage. Stills come from the agent's built-in ImageGen, but image-to-video and the optional white watermark cleanup depend on the hbg-gemini-flow-suite container, named gemini-flow-suite, with the workspace at /workspace and outputs at /data/outputs. The README says to complete a one-time user-controlled authorization in that runtime before running check_prerequisites.sh.

Official sources

  1. Issues
  2. License: MIT
  3. Mr-funny/hbg-classical-poem-silk-video on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mr-funny-hbg-classical-poem-silk-video.svg)](https://hysenlabs.com/projects/mr-funny-hbg-classical-poem-silk-video)