# Agent Storyboard keeps the script, the shots, and the voice in one table

> Agent Storyboard is a local web service that a Codex or Claude Code agent writes into through MCP tools, so a video project keeps its script, shot table, generated images, motion clips, and voiceover in one place on your own machine. The repo is named codex-storyboard, the plugin is named agent-storyboard, and the voiceover path sends your lines to an online service.

**Yuuhann1999/codex-storyboard** — 本地多项目 Codex 视频分镜工作台，支持图片/视频生成任务、HyperFrames 与 Remotion 自动回填。Local multi-project storyboard workspace for Codex.

- Repository: https://github.com/Yuuhann1999/codex-storyboard
- Stars: 353 · Forks: 50
- Language: JavaScript
- License: MIT
- Published: 2026-09-15 · Updated: 2026-09-15 · Language: en
- Canonical page: https://hysenlabs.com/projects/yuuhann1999-codex-storyboard

## One table, because every window switch loses context

The problem statement is unusually concrete. A script lives in a document, the storyboard lives in a spreadsheet, images get made in a web generator, voiceover happens in yet another tool, and assets scatter across folders. Switch windows and the context goes with them. Agent Storyboard collects all of that into a single local storyboard table, and the row is the unit of work: shot type, media, duration, dialogue, visual description, generation method, asset preview, and notes all sit on one line.

Four claims follow from that design. The agent writes into the table directly, so a whole project can be created from one sentence without anyone driving a browser. Results fill themselves back into the right shot, so nobody matches filenames by hand. Your only job is acceptance, since every automated result is visible and editable in the table. And the data stays on your machine, with projects, scripts, and assets stored locally.

The effect is that the table becomes the review surface rather than a documentation artifact. Long projects switch to a compact layout that fits about seven shots on one screen, and shots whose pacing is off are flagged in place instead of being discovered during an export.

## The agent writes through MCP, never through the browser

The workspace is a local web service, default address `http://127.0.0.1:43218`, and the agent talks to it through MCP tools instead of clicking. The project states the boundary plainly: the agent works through these tools, does not need to control a browser, and does not directly edit data files. `open_storyboard` starts or connects to the workbench and returns the link. `create_storyboard_project` creates a complete project with every shot and an optional `DESIGN.md` in one call.

The generation tools carry most of the weight. `list_storyboard_generation_tasks` reads what is pending, and `generate_storyboard_image` does four things in one call: claim the task, generate through Codex, validate the result, and write it back into the shot. Long running work such as motion video goes through the claim, complete, fail, and heartbeat tools, which is how a task survives an agent session that ends mid-render.

Two tools exist purely as brakes. `plan_broll_motion` is a confirmation gate for a B-roll motion plan, and `manage_storyboard_audio` handles voice generation, version selection, dialogue alignment, and applying shot duration. `inspect_storyboard_environment` reports on voiceover, FFmpeg, Whisper, and the local agent before any of that is attempted.

## Three generation modes, and the queue feeds itself back

Every shot picks one of three generation modes. AI image generation builds the picture from the visual description: Codex uses its own built-in image generation, and Claude Code calls the local Codex through `generate_storyboard_image`. HyperFrames or Remotion motion renders subtitles, infographics, and transitions as code, with the agent rendering locally and needing the matching plugin plus a render toolchain. Manual assets are your own footage, uploaded as images or video.

The loop around those modes is the part worth tracing. You give a topic and requirements, the agent writes the whole storyboard project through MCP, you check and adjust it in the table, shots go into a generation queue one at a time or in a batch, the agent produces images, motion video, or voice, and the result fills back into its shot, which returns the shot to you for review. The flow deliberately cycles rather than terminating, because the human check is the loop's exit condition.

Nothing in that loop depends on the agent having a working browser session. That is the practical difference from a browser-driving approach, where a dialog in the wrong window silently ends the run.

## Shot filenames carry the shot index

Reordering is drag and drop or `Alt + ↑/↓`, with duplication, insert below, and undo for an accidental delete. The detail that makes reordering safe is filename handling: asset filenames include the shot number, and they resync automatically after a reorder so assets do not overwrite each other. In a workspace where a long project will be reordered several times before it is finished, that is the difference between a rename chore and a non-event.

The table stays searchable by dialogue, visual description, or notes, which matters once a project has more shots than fit on screen. The top bar shows live counts for generating, queued, and failed, and clicking a count jumps to the shots in that state, so a failure is a place you can go rather than a line you have to hunt for.

The wider project surface is small and concrete: create, rename, duplicate, search, and delete projects; aspect ratios `9:16`, `16:9`, `3:4`, `4:3`, and `1:1`; a presentation mode that previews shot by shot full screen for checking pace; export to Markdown, HTML, Word, or plain text; and local assets with zoomed preview, replacement, and deletion.

## Voiceover is aligned by transcribing it back

Voice generation happens on the script page rather than shot by shot. You describe the voice style, or upload a reference clip to clone the timbre, and the result is then recognised with local Whisper, aligned against each shot's dialogue, and applied to the shot duration in one click. Aligning by transcription rather than by arithmetic is what lets the durations settle to the actual audio instead of an estimate.

The dependency table is honest about what breaks. The voice service itself is VoxCPM, called over the network from Node, so narration needs connectivity. FFmpeg handles voiceover transcoding and duration reading, installed with `brew install ffmpeg`. Whisper does the alignment, installed with `brew install whisper-cpp`, and the table is explicit that voiceover still works without it, since only the alignment step needs it. Codex CLI is optional too, letting Claude Code generate images through Codex when you would rather not use the agent's own image path or upload something yourself.

Python is not required. The environment check exists to report only these items rather than a long list of theoretical requirements, which is a sign the author expects people to run with a partial setup.

## Twelve styles land in DESIGN.md, and a shot can override them

Visual consistency is handled with a file rather than a per-shot prompt. There are 12 built-in visual styles, and selecting one writes it into the project's `DESIGN.md`. Every image and video generated afterwards follows that specification, which is what stops a nine-shot project from drifting into nine different looks.

The precedence rule is the interesting part: an explicit requirement on a single shot wins over the project-wide specification. That ordering makes sense for a production tool, since the whole point of a style guide is a default rather than a cage, and it means a one-off exception does not require rewriting the guide for the other eight shots.

Covers are handled in their own workbench and are managed separately for vertical 9:16 and horizontal 16:9, each with templates, a reference image, and a prompt. The asset overview lists only assets that actually exist, so progress is visible without reading the table, and dark mode plus a card layout for narrow windows is what lets the table live inside an agent sidebar instead of demanding a full screen.

## Voice chunk size is an exposed quality dial

The environment variables are unusually well explained, which is where most of this project's real behaviour is written down. `AGENT_STORYBOARD_VOICE_CHUNK_CHARS` caps how many characters go into one voiceover segment, and the tradeoff is spelled out: a larger value produces fewer seams between segments but makes the timbre drift more in the later parts, and a smaller value stays steadier at the cost of more seams. The default is 800.

The rest of the variables cover the parts you would otherwise have to read the source for. `AGENT_STORYBOARD_PORT` sets the port, default 43218. `AGENT_STORYBOARD_DATA_DIR` sets the data directory, and `AGENT_STORYBOARD_FFMPEG` points at the binary, with the same pattern for FFPROBE, Whisper, and the Whisper model. `AGENT_STORYBOARD_VOXCPM_URL` points at a self-hosted voice service if you would rather not use the hosted one, and `AGENT_STORYBOARD_CODEX` sets the Codex CLI path.

`AGENT_STORYBOARD_CODEX_MODEL` names the image model, defaulting to whatever Codex is configured with, and it has a fallback written into the behaviour: if the chosen model does not work with a ChatGPT account, image generation retries with gpt-5.5. Legacy `CODEX_STORYBOARD_*` names still work, and the new names take priority, which is the migration path for anyone who configured the older version.

## Named codex-storyboard, installed as agent-storyboard

The repository name and the product name disagree, and you will meet both in the first five minutes. Installing for Codex takes two commands:

```bash
codex plugin marketplace add Yuuhann1999/agent-storyboard
codex plugin add agent-storyboard@agent-storyboard
```

The Claude Code path mirrors it with `claude plugin marketplace add Yuuhann1999/agent-storyboard` and `claude plugin install agent-storyboard@agent-storyboard`. After either install you open a new conversation so the MCP tools reload, and `/mcp` in Claude Code confirms the connection. Node.js 18 or newer is required, plus Codex or Claude Code. The same rename shows up in the data directory, where an existing `~/.codex-storyboard/` is still used rather than moved or deleted.

For local development the project is cloned and run straight from source:

```bash
git clone https://github.com/Yuuhann1999/agent-storyboard.git
cd agent-storyboard
npm start            # http://127.0.0.1:43218，数据在仓库内 ./data
npm run check        # 语法检查
npm test             # 单元与集成测试
```

The package is private, version 0.8.0, and has no runtime npm dependencies at all: the frontend is plain HTML, CSS, and JavaScript, and the local service uses only the Node standard library. That is why `npm start` needs nothing installed first, and why the only real maintenance task is syncing the root `server.mjs` and `public/` into the copy bundled under the plugin, a consistency check that `npm test` enforces. There are no GitHub releases, the license is MIT, and the last push is dated 2026-09-30.

## Conclusion

Agent Storyboard suits someone who keeps losing context between a script document, a spreadsheet, an image generator, and a voiceover tool, and who wants the agent to do the writing while a human keeps the final say in one reviewable table. It does not suit a team that needs hosted collaboration, a rendered timeline, or a guarantee that generated assets never leave the machine. Before the first project, check which of VoxCPM, ChatGPT image generation, and FFmpeg you are willing to depend on, because each one is optional in a different way.

## FAQ

### What does the Agent Storyboard workspace replace?

It collects the script, the storyboard table, generated images, motion clips, voiceover, and assets into one local storyboard table. The stated motivation is that every switch between windows loses context, and that results fill back into the right shot without matching filenames by hand.

### Do I need Python, FFmpeg, or Whisper to use Agent Storyboard?

Python is not needed. FFmpeg handles voiceover transcoding and duration reading and is installed with brew install ffmpeg, while Whisper aligns voiceover to the dialogue with brew install whisper-cpp, and voiceover can still be generated without it.

### Where does Agent Storyboard keep my data, and what leaves the machine?

Projects, scripts, and assets are stored locally in ~/.agent-storyboard/ by default, an existing ~/.codex-storyboard/ keeps being used, and the local service listens only on 127.0.0.1 and rejects cross-origin requests. Voiceover sends dialogue to the VoxCPM online service, and Codex image generation sends prompts to ChatGPT's image service.

### How do I install the Agent Storyboard plugin into Codex?

Run codex plugin marketplace add Yuuhann1999/agent-storyboard followed by codex plugin add agent-storyboard@agent-storyboard, then open a new conversation so the MCP tools reload. The plugin starts the local workbench itself and prints a link, by default http://127.0.0.1:43218.

### What is plan_broll_motion for in Agent Storyboard?

It is the confirmation gate for a B-roll motion plan. Longer work such as motion video is handled through the claim, complete, fail, and heartbeat_storyboard_generation_task tools so a task survives an agent session that ends mid-render.

## Sources

- [Issues](https://github.com/Yuuhann1999/codex-storyboard/issues)
- [License: MIT](https://github.com/Yuuhann1999/codex-storyboard/blob/main/LICENSE)
- [README](https://github.com/Yuuhann1999/codex-storyboard/blob/main/README.md)
- [Yuuhann1999/codex-storyboard on GitHub](https://github.com/Yuuhann1999/codex-storyboard)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/yuuhann1999-codex-storyboard
