# Image Story Video Wizard: a confirmation-gated Codex and WorkBuddy skill for audio-first image-story videos

> The skill walks a project through sixteen named stages, from BRIEF to FINAL_RENDER, and stops at four confirmation gates before it spends money on voice or image generation. It is a workflow controller, not a renderer.

**aaronyi97/image-story-video-wizard** — A confirmation-gated Codex and WorkBuddy skill for audio-first image-story video production.

- Repository: https://github.com/aaronyi97/image-story-video-wizard
- Stars: 351 · Forks: 58
- Language: Python
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/aaronyi97-image-story-video-wizard

## What the skill actually controls, and what it does not

Image Story Video Wizard is a skill definition for Codex and WorkBuddy, written in Python, MIT licensed. Its job is to keep a long production process in order. The README describes the core idea directly: instead of handing the user one large process, the skill judges which stage the project is in, does the work it can do alone, and asks the user only for materials or confirmations at key points. The README states plainly that if all you need is a single script, one image, or ordinary editing, you do not need to call this skill. That is an unusually honest scope statement, and it is the right way to read the project. This is not a video generator. It is a controller that sequences other tools and refuses to move forward until a human agrees. The audience is people producing narrated image-story videos: AI book summaries, history and emotional stories, audio stories, podcast audio paired with still images, and static-image narrative videos that need character consistency, a fixed visual style, subtitles and text cards.

## The five-line turn format and the sixteen-stage state machine

Every turn the skill produces the same five labelled items, quoted from the README: what stage is running now, what the skill will do, what the user needs to supply, what will be delivered when it finishes, and what the next step is after confirmation. That fixed shape is the mechanism. It means the user never has to reconstruct where the project stands.

The full stage list runs START, BRIEF, BENCHMARKS, WRITING_PACK, SCRIPT, VOICE, STORYBOARD, VISUAL_STYLE, CHARACTER_ANCHORS, IMAGE_PROMPTS, IMAGE_GENERATION, ASSET_QC, MUSIC, PREVIEW, FINAL_RENDER, FEEDBACK. State lives in a file called PROJECT_STATE.json, which the README says records the current stage, confirmed decisions, artifact paths, and the next question the user has to answer. Because the state is a file rather than conversation memory, a project can be moved to another host and resumed.

The gates are the interesting part. The README lists four: no writing pack before the benchmark direction is confirmed; no full voice generation before the script is final; no batch image generation before the visual style, text style and sample images are confirmed; no final render before the preview is approved. There is also a failure rule: when a step fails, the skill returns only to the earliest affected stage rather than restarting the project. That is a real design decision about cost. Voice and image generation are the expensive stages, and all four gates sit in front of them.

## Installing the skill and starting a first project

The README gives two installation routes. The conversational one is to send Codex a sentence instructing it to use skill-installer to install from the repository URL, with the skill at the repository root and the install name image-story-video-wizard. Codex loads the skill on the next turn. The terminal route clones the repository into the Codex skills directory:

```bash
# Codex
git clone https://github.com/aaronyi97/image-story-video-wizard.git ~/.codex/skills/image-story-video-wizard
```

The README notes that if WorkBuddy also loads from a local skill directory, you can install the same repository there or link to the copy under Codex. After installation, the README's example opening message is to ask for a new project on a given topic and to have the skill drive the process forward. The skill should reply with the five labelled items for the START stage and tell you what it needs first. The README also shows the validation commands, which are worth running before you trust the install:

```bash
python3 -m unittest discover -s tests -v
python3 /path/to/skill-creator/scripts/quick_validate.py .
```

The README states that the automated checks cover skill structure, the 16 stages, the five-part guidance format, key tutorial requirements, state machine consistency, and the state file needed to initialize and resume a project. Note the second command depends on a skill-creator script outside this repository, so you need that path to exist.

## The external services are checked, not assumed

The tutorial route in the README is concrete about which outside tools it expects. For long scripts it suggests finishing benchmark study and the writing pack in Codex, then handing off to Kimi K3 inside WorkBuddy, selecting Max mode, the highest thinking intensity and 1M context when the page offers them. For voice it names Doubao Seed-TTS 2.0, with a specific method: test five to ten voices on the same roughly 20-second passage, confirm the voice, then adjust speed. For images it suggests three to five sample images first, then locking the visual style, the character master and the text style, with the manual route defaulting to one prompt and one image per new conversation. For editing it prefers HyperFrames to assemble the assets, preview first, render after confirmation.

What matters here is the honesty clause. The README says WorkBuddy, Kimi, Doubao, image generation and rendering are live external capabilities, that the skill checks whether each is available when entering the corresponding stage, and that it will not claim a capability participated or completed when it was not actually invoked. That is the correct behaviour for a skill whose value depends on other services, and it also tells you where the failure modes live.

## Where this skill is the wrong tool

The first limitation is stated by the project itself: if you only need one script, one image, or normal editing, the skill is unnecessary overhead. Sixteen stages and four gates exist to protect expensive batch operations, and there is nothing to protect if you are not doing batch work.

The second limitation is that the skill does not render anything. Every capability that produces audio, images or video comes from outside: WorkBuddy, Kimi, Doubao Seed-TTS, an image generator, HyperFrames. The README says the skill checks availability at the relevant stage, which means a stage can simply stop when a service is unreachable. If you want one program that takes a picture and returns a movie, this is the wrong layer entirely.

The third is throughput. Four confirmation gates mean four waits for a human, and the README's manual image route is one prompt and one image per conversation. That is a deliberate cost control, but it makes the skill unsuitable for unattended pipelines or for anyone who wants to generate a hundred images in one pass. The README does not document a way to bypass the gates, and it does not document rollback behaviour beyond the rule that a failed step returns to the earliest affected stage.

## How it differs from one-shot image-to-video tools

Tools such as Magic Hour or MiriCanvas take an image, or a set of images, and produce motion or a video clip. The difference in approach is not quality, it is where the human sits. A one-shot converter puts the user at the start and the end: upload, wait, download. Image Story Video Wizard puts the user at four points in the middle, and its stages are about deciding things rather than rendering them. BENCHMARKS, WRITING_PACK, VISUAL_STYLE and CHARACTER_ANCHORS have no counterpart in a converter, because a converter has no opinion about your script or your character consistency. Conversely, the wizard has no motion synthesis of its own, so it is not a replacement for those tools. If your problem is that a still image looks static, a converter addresses it. If your problem is that your tenth video drifts from your first in voice, style and character, the wizard's gates and its PROJECT_STATE.json address that.

## Maintenance, upgrade cost and the MIT licence

The repository is not archived. The last push was on 2026-09-15, two days before this article, so the code is recent. There are no retrieved releases, which means updates arrive as commits on main rather than as tagged versions. For a skill installed by git clone into ~/.codex/skills/image-story-video-wizard, upgrading means pulling that directory; the README does not document a version pinning scheme, and it does not document rollback, so if you modify the skill locally you own the merge. Because the skill drives external services whose interfaces change independently, a pull can change which stages behave as expected even when the skill's own stage list is unchanged.

The licence is MIT, which permits commercial use and modification. The README does not discuss the terms of the external services the tutorial route names, and those are governed by their own agreements, not by this repository's MIT licence. Nothing here is legal advice; if you plan to sell the output, check the terms of the voice, image and rendering services separately.

## Conclusion

Adopt it if you already produce narrated image-story videos and want the stage order and confirmation points written down, and you are prepared to run the external tools it points at yourself. Do not adopt it if you want a single tool that turns a picture into a finished movie, or if you need unattended batch rendering, because every gate waits for a human. Before you start, verify that your host loads skills from ~/.codex/skills, that the 16 stage names in SKILL.md match the ones the README lists, and that the external voice, image and rendering services you intend to use are reachable from your machine.

## FAQ

### How do I turn a picture into a movie with Image Story Video Wizard?

The skill does not convert images to motion itself. Its IMAGE_GENERATION and FINAL_RENDER stages coordinate external image generation and editing tools such as HyperFrames, with a preview gate before the final render. You supply the pictures and confirm the preview.

### How can I create a video story with Image Story Video Wizard?

Install the skill into Codex or WorkBuddy, then ask it to start a new project on your topic and let it drive. It runs from START through BRIEF, BENCHMARKS, WRITING_PACK, SCRIPT and onward, asking you only for materials or confirmations at each gate.

### Can AI convert an image to a video with audio using Image Story Video Wizard?

Audio and images are separate stages in this skill. VOICE handles narration through an external service such as Doubao Seed-TTS 2.0, and the image stages run separately, so the audio is not generated from the image. The skill checks whether each external capability is available before entering its stage.

### How can I create a video using pictures with Image Story Video Wizard?

The README's route is to lock the visual style, character master and text style using three to five sample images, then generate the rest, assemble with HyperFrames, preview, and render only after you approve. The manual image route defaults to one prompt and one image per conversation.

## Sources

- [aaronyi97/image-story-video-wizard on GitHub](https://github.com/aaronyi97/image-story-video-wizard)
- [Issues](https://github.com/aaronyi97/image-story-video-wizard/issues)
- [License: MIT](https://github.com/aaronyi97/image-story-video-wizard/blob/main/LICENSE)
- [README](https://github.com/aaronyi97/image-story-video-wizard/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/aaronyi97-image-story-video-wizard
