Open-source project
Endless1936/book-video avatar
Endless1936/book-video

book-video refuses to replace a good render until the new one passes

A natural-language workflow for creating atmospheric short book videos with AI-generated visuals, HyperFrames, GSAP animation, voiceover timing, subtitles, and BGM.

351 stars68 forksJavaScriptNOASSERTION

At a glance

What is it?
A natural-language workflow for atmospheric book videos, driven by an agent rather than a CLI, whose real engineering is a five-state episode file, non-zero exits on a bad render, and subtitles timed from FFmpeg silence boundaries instead of Whisper.
Who is it for?
book-video suits a creator who publishes repeatedly and is willing to work inside a Chinese toolchain, since voiceover either comes from an agent environment with Doubao reference voice cloning or from an MP3 exported elsewhere and handed over by path. It does not suit a one-off video, and it does not clear rights for you, because Apache-2.0 covers the code, docs and templates while the four bundled background tracks and the reference scripts sit outside it.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

You describe a book, the agent runs the pipeline

book-video is a workflow for making short atmospheric videos that sell a book, and its interface is a conversation. You open Codex and describe the video you want in plain language, without learning a command or reading a script. On first use the agent inspects and prepares the local environment and asks you only where a decision is genuinely yours.

The steps it follows are specific. It confirms the title, author and edition, prints the full voiceover copy in a code block so it can be copied straight into JianYing, with the book title on the first line, and saves the same text as a local script file. You review the copy in the conversation, and only after approval does it generate the atmosphere images and video frames. The project description names HyperFrames and GSAP animation as the motion layer, with voiceover timing, subtitles and background music assembled around them.

Asking for a recommendation works too. Requests like recommending five books suited to short videos, or asking for a video about loneliness and self-growth, are ordinary entry points. Asking to replace the copy, the images or the audio starts a new proposal rather than mutating the current one in place.

Five states per episode and no terminal failure

Every episode keeps a state file at `episodes/<book title>/workflow-state.json`, and what makes it interesting is that it has no terminal failure state. Artifacts are recorded as executable, valid, degraded, pending diagnosis or stale, arranged by dependency rather than by step number, so a run can stop anywhere and resume without pretending the work is finished.

The file is also disposable. If it is missing or corrupt, it is rebuilt from the artifacts and validation reports that already exist, which means the state is a cache of what happened rather than the only record of it.

Three subcommands cover the cases you actually hit. In the commands below, `<书名>` stands for the book title:

bash
npm run workflow -- status "<书名>"
npm run workflow -- next "<书名>"
npm run workflow -- repair "<书名>"

`status` shows where the episode stands, `next` tells the agent what to do next, and `repair` re-checks the artifacts against the checks. Script confirmation is an explicit `approve`, and `deliver` is used once the finished video is embedded back into the conversation. Every other step is inferred from the artifacts and their validation reports, which is why the same command works after an interrupted render.

A bad render exits non-zero and leaves the old video alone

The failure design is the part worth borrowing. A failing command keeps a non-zero exit status, so a broken artifact cannot quietly take effect, and at the same time it prints `BOOK_VIDEO_DIAGNOSTIC` and writes the detail to `tmp/last-workflow-diagnostic.json`. The agent reads that file, checks inputs or dependencies, fixes them, and retries the smallest step that failed rather than restarting the episode.

The other half is that a previously valid finished video is not replaced until the new candidate passes. For a workflow whose output is a published video, that ordering is the whole difference between a broken render costing you a retry and a broken render costing you an episode.

The scripts behind it are plain Node with no dependencies declared, so there is nothing to install to run the checks:

json
"scripts": {
  "check": "node scripts/check.mjs",
  "test": "node scripts/run-tests.mjs",
  "workflow": "node scripts/workflow-state.mjs",
  "smoke:timings": "node scripts/tests/smoke-timing-fallback.mjs"
}

`smoke:timings` is named for the timing path, which is where the next section starts.

Subtitles are timed from silence, not from the transcript

Voiceover length and subtitle timing are the two things that go wrong in narrated video, and this workflow handles them by refusing to let the speech model supply timestamps. FFmpeg detects silence boundaries in the audio to produce a timing reference, and `script.csv` is the source of truth for subtitle text. Whisper is present only to check what a text-to-speech engine actually said, which is a content check rather than a timing one.

That split is a judgement worth naming. Whisper timestamps are good enough for a transcript and poor enough to drift against a synthetic voice whose pauses come from punctuation rather than from a person, so trusting them produces subtitles that lag the narration. Silence boundaries keyed to the real audio, combined with copy you already approved, keep the two in agreement.

The smoke test named `smoke:timings` covers the fallback behaviour of that path, and the timing logic sits alongside the other steps in `scripts/` rather than in a separate tool.

Voiceover depends on an agent environment that may not be there

Step four is where this workflow is most exposed. If the agent environment explicitly supports free Doubao reference voice cloning, the spoken audio is generated with the project's reference voice. If it does not, you export the voiceover yourself from JianYing, save it as an MP3, and hand the file path to the agent.

So the audio step has two entirely different shapes depending on your environment, and only one of them is automated. The fallback is not a degraded mode, it is a hand-off: a person records or generates the track, exports a file, and tells the workflow where it is. Everything downstream, the silence-based timing and the subtitles, then works the same either way, which is the correct place for the seam.

For anyone outside that ecosystem the practical consequence is that you supply your own text-to-speech output and the workflow still gives you aligned subtitles. What you lose is the reference voice, and with it the consistency between episodes that a fixed voice provides.

Apache-2.0 covers the code and almost nothing else

The licensing is layered and the README is explicit about the layers. Code, documentation and the reusable templates are Apache-2.0, copyright prototech as the organisation and endless as the online name, with 未济 also named in the copyright line. The bundled 得意黑 font follows its upstream SIL OFL 1.1 licence, and its licence file travels with the template at `templates/shared-video-template/body/fonts/LICENSE.txt`. Apache-2.0 is stated not to cover third-party tools, models, image generation services or media assets.

Two categories sit outside the grant and matter commercially. The reference copy was compiled from third-party voiceover, so it is not covered and the rights need checking before redistribution; the repository keeps the copy and its analysis but not the original videos. Then there are four default background tracks in `assets/bgm/`, 城南花已开, 红色高跟鞋, 起风了 and 如愿, shipped as study material with the note that Apache-2.0 does not apply and the copyright stays with the respective rights holders.

The practical reading: any episode you publish with the default music needs clearance you have to obtain, and GitHub's own licence field for this repository resolves to no recognised identifier, so the README and the LICENSE file are the authority rather than the badge.

What is in the repository, and what the agent still needs

The tree is small and content-shaped: `assets/`, `data/`, `docs/`, `episodes/`, `scripts/`, `templates/`, plus `AGENTS.md`, `LICENSE` and `README.md`. The reusable template lives under `templates/shared-video-template/`, the per-episode state under `episodes/`, and the deeper production rules are in `AGENTS.md` and `docs/book-video-playbook.md` rather than in the README.

The credits are worth keeping in mind as provenance. The workflow follows a post by Serena, and the copy-structure research is based on three videos published by the Douyin creator 听页, with the written analysis under `assets/reference-videos/`.

Against editing in a conventional timeline editor, the difference in approach is ownership of the artefacts. A conventional editor gives you one file and a timeline you adjust by hand; here the video is the output of a pipeline whose steps are re-runnable, whose state is written down, and whose previous good result is protected until a new one validates. That is worth the agent dependency only if you make more than one episode.

Editorial conclusion

book-video suits a creator who publishes repeatedly and is willing to work inside a Chinese toolchain, since voiceover either comes from an agent environment with Doubao reference voice cloning or from an MP3 exported elsewhere and handed over by path. It does not suit a one-off video, and it does not clear rights for you, because Apache-2.0 covers the code, docs and templates while the four bundled background tracks and the reference scripts sit outside it. Before your first episode, replace the default music, confirm the font licence file travels with your template, and read episodes/<book>/workflow-state.json after the first failed render rather than re-running the step by hand.

Frequently asked questions

What is the book-video project?

A reusable workflow for making atmospheric short videos that sell a book, driven by natural language through Codex rather than by a command line. It covers choosing the book, writing the voiceover copy, generating atmosphere images and frames, aligning narration, producing subtitles and assembling music.

How do I start a book video in book-video?

Open Codex and describe the video in natural language, or ask it to recommend five books suited to short videos. On first use it checks and prepares the local environment and asks you where a decision is needed, then confirms the title, author and edition before writing the copy.

How does book-video produce subtitle timings?

FFmpeg detects silence boundaries in the audio to build a timing reference, with `script.csv` as the source of truth for subtitle text. Whisper is used only to check what a text-to-speech engine said and does not provide timestamps.

What commands does book-video expose?

Four npm scripts with no declared dependencies: `check`, `test`, `workflow` and `smoke:timings`. The workflow script takes `status`, `next`, `repair`, `approve` and `deliver`, with the book title as the argument, for example `npm run workflow -- status "<书名>"`.

Can I publish a book-video episode with the bundled music?

Not on the strength of the project licence alone. Code, docs and templates are Apache-2.0 and the bundled font is SIL OFL 1.1, but the four tracks in `assets/bgm/` are explicitly not covered by Apache-2.0 and the copyright stays with the respective rights holders, as does the reference copy compiled from third-party voiceover.

Official sources

  1. Endless1936/book-video on GitHub
  2. Issues
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/endless1936-book-video.svg)](https://hysenlabs.com/projects/endless1936-book-video)