# Vox Director: an agent skill that turns one topic into a paper-collage explainer video

> Vox Director is a Python agent skill that generates Vox-style paper-collage videos end to end on the Atlas Cloud API plus local ffmpeg. It is opinionated about where the look comes from and where the human stays in the loop.

**Alisa0808/vox-director** — Turn one topic into a finished Vox-style paper-collage explainer/ad video — automated end to end on Atlas Cloud + ffmpeg. An agent skill.

- Repository: https://github.com/Alisa0808/vox-director
- Stars: 2,106 · Forks: 319
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/alisa0808-vox-director

## The problem Vox Director solves, and for whom

Producing a paper-collage explainer by hand is a chain of separate jobs: write a script, design a poster per beat, animate each poster, record narration, pick music, cut captions, then assemble. Each step lives in a different tool, and the visual consistency that makes the format read as a single piece comes from a human holding the style in their head across all of them.

Vox Director collapses that chain into one repository. The README describes it as an agent skill that takes a one-line topic and returns an mp4, running on the Atlas Cloud API plus local ffmpeg. The audience is narrow on purpose: people who already work inside a coding agent such as Claude Code or Codex, are comfortable exporting an API key, and want a generated first cut rather than a blank timeline. If you do not have a coding agent, the project has no interface for you.

The format itself is defined in the README: hand-cut paper cut-outs, torn edges, tape, halftone dots, newspaper clippings, bold flat color per beat, and big cut-out headlines. That is a specific look, not a general video generator, and the repository is built around reproducing it consistently.

## The beat map, the two gates, and what the agent actually runs

Everything in a project hangs off a single beats.json file. The README lays the flow out as six stages: pick a narrative arc and write the beat map, run a style bake-off, generate one collage poster per beat, animate each poster, produce narration and music, then assemble with ffmpeg including music ducking under the voice-over and burned-in captions and watermark.

Two of those stages are human gates. Gate 1 is approving the beat map before anything expensive runs. Gate 2 is picking a look by eye from three or four rendered themes. The rest is automated. That placement is the interesting design decision: the project spends its human attention budget on structure and style, the two things that are hardest to fix after generation, and automates the mechanical steps.

The README states a principle the pipeline is built around: the look is born in the image step, and if the poster is not a rich collage, nothing downstream saves it. Motion is added afterward. By default an AI video model animates the whole poster. An optional local keyframe engine instead cuts the poster into parts and drives them frame by frame, which the README describes as pixel-exact and free of content filters, and which it recommends for real people.

Model IDs are not hardcoded in a way you can trust. The skill fetches the live list from GET https://api.atlascloud.ai/api/v1/models before running, which is a sensible hedge against provider churn.

## Installing Vox Director as a Claude Code skill

The README gives two install paths. Option A clones the repository directly into the Claude Code skills directory. Option B downloads the packaged vox-director.skill file and installs it through the Claude skills UI. For non-Claude agents, the README points at AGENTS.md as the entry point, which then leads to SKILL.md.

Option A is one command:

```bash
git clone https://github.com/Alisa0808/vox-director.git ~/.claude/skills/vox-director
```

After that, export your Atlas Cloud key. The README points at atlascloud.ai/console/api-keys for obtaining one:

```bash
export ATLASCLOUD_API_KEY="sk-..."
```

Requirements listed in the README are a coding agent, the Atlas Cloud key, ffmpeg plus ffprobe, and Python 3 with Pillow for caption and watermark overlays. On macOS the README suggests brew install ffmpeg, and pip install pillow for the Python dependency.

First real use is a prompt to your agent, not a CLI invocation. The README's own example is: "Make me a Vox-style collage video introducing Mexican street food, English, 16:9, 15 seconds." The agent should respond with a draft beat map for your approval, then a style bake-off, then generation, and finally write out/<project>/final.mp4. If your agent starts generating before showing you a beat map, the skill has not been picked up correctly.

## A-roll and C-roll: the same engine, two other inputs

B-roll is the default path: a topic goes in, everything is generated. The repository also documents two reuse modes that matter more than they first appear.

A-roll starts from a talking-head video you already have. The README says it is ASR-segmented into beats and re-styled into the collage look while keeping the real face, lip-sync and gestures frame for frame, using google/gemini-omni-flash/video-edit with automatic retry on seedance-2.0/reference-to-video. That is a different product from a text-to-video generator: it is a restyling pass over existing footage.

C-roll starts from a single still photo, a selfie or a product shot. The subject is cut out as a photographic sticker and never redrawn, and each beat's poster is generated around it via google/nano-banana-2/edit. The README notes narration can be cloned into the subject's own voice using bytedance/seed-audio-1.0.

The repository ships example beat files that show the shape of the input: examples/money-15s.beats.json, examples/money-60s-9x16-english.beats.json, examples/tang-30s.beats.json, examples/ronaldo-9x16-kling.beats.json, and examples/cr7-act.elements_spec.json. Reading those before writing your own is the fastest way to learn the schema, since the README does not reproduce it inline.

## Where Vox Director breaks down

The most consequential limitation is stated by the project itself: the look is decided in the image step. If the generated poster is a weak collage, the animation, narration and captions will not rescue it. That means the quality ceiling is set by an image model and by the prompt structures in references/prompt-guide.md, not by anything you control in post. A beat with a bad poster is a re-roll, not an edit.

The second limitation is the API dependency. Every generation stage except assembly runs through Atlas Cloud. There is no documented offline or local-generation path for keyframes, motion, narration or music. If your source material cannot leave your environment, A-roll and C-roll are effectively closed to you, and the README does not describe a self-hosted alternative.

Model drift is a real operational risk. The README states plainly that model IDs drift, which is why the skill fetches the live model list before running. A fetch that returns a changed or missing ID is a failure mode you should expect to encounter rather than treat as exceptional. The README does not document rollback or a pinned-version mode.

Finally, the project is an agent skill, not an application. There is no GUI, no queue, no project dashboard. If your agent misreads the workflow, you debug a markdown file and a set of scripts, and the README does not describe a validation or dry-run mode for that.

## How it differs from a timeline editor or a general text-to-video model

The obvious alternative is doing this in a conventional editor such as DaVinci Resolve or After Effects with a collage asset pack. The difference is not quality, it is where the work goes. An editor gives you frame-level control and no per-beat generation cost, but every poster is manual and the style consistency depends on you. Vox Director trades that control for a scripted, repeatable pipeline with two checkpoints. If you need to nudge a cut by four frames, the editor wins outright.

The closer alternative is a general text-to-video model used directly. Those produce continuous generated footage; Vox Director produces a collage poster per beat and then animates it. That intermediate poster is the whole point. It is a reviewable artifact you approve at Gate 2 before spending on motion, and it is what keeps the output reading as one designed piece rather than a sequence of unrelated clips. A general model gives you no such checkpoint and no per-beat visual consistency guarantee.

For A-roll specifically, the comparison is video restyling tools. Vox Director's claim is frame-for-frame preservation of face, lip-sync and gestures. That is a stronger constraint than a stylization filter, and it is the part most worth testing on your own footage before trusting it.

## Maintenance, licence and the cost of upgrading

The repository is MIT licensed, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are retained. That is the standard MIT bargain and it is permissive enough for most commercial video work. This is not legal advice; if you are embedding the skill in a product, have your own counsel read the LICENSE file at the repository root.

The maintenance picture is mixed. The last push was on 2026-08-11, roughly five weeks before this writing, and the repository is not archived. There are no retrieved releases, so there is no versioned artifact to pin against. package.json declares version 1.0.0, but the install paths in the README are a git clone or a packaged .skill file, neither of which gives you a tagged upgrade channel.

Upgrade cost is dominated by the external surface. Because model IDs drift and the skill resolves them at runtime, an upgrade can change your output without any change to the repository. The practical defence is to keep your project's beats.json files under version control and re-run the style bake-off after any pull, so you can see whether the look shifted before committing to a full render. The README does not describe a changelog or migration notes.

## Conclusion

Vox Director fits teams and solo creators who already have a coding agent, an Atlas Cloud key and ffmpeg installed, and who want a repeatable collage explainer pipeline rather than a timeline editor. It is the wrong tool if you need frame-accurate editorial control, if you cannot send source footage or photos to a third-party API, or if you want a hosted editor instead of a scripted workflow. Before committing, verify three things in your own environment: that your agent actually picks up SKILL.md, that the live model list from GET https://api.atlascloud.ai/api/v1/models still contains the IDs the workflow names, and that the A-roll and C-roll paths behave acceptably on your own footage.

## FAQ

### What is Vox Director and who is it for?

It is an agent skill that turns one topic into a finished Vox-style paper-collage explainer or ad video, covering script, collage keyframes, motion, voice-over, music and captions. It is aimed at people already working inside a coding agent such as Claude Code or Codex who have an Atlas Cloud API key and local ffmpeg.

### How do I install Vox Director?

The README gives two options: clone the repository into ~/.claude/skills/vox-director, or download the packaged vox-director.skill file and install it through your Claude skills UI. Non-Claude agents start from AGENTS.md, which leads to SKILL.md.

### What do I need before running Vox Director?

The README lists a coding agent, an Atlas Cloud API key, ffmpeg plus ffprobe, and Python 3 with Pillow for caption and watermark overlays. The API key is exported as the ATLASCLOUD_API_KEY environment variable.

### Does Vox Director work with videos or photos I already have?

Yes. The README describes an A-roll path that ASR-segments an existing talking-head video and re-styles it while keeping the real face, lip-sync and gestures frame for frame, and a C-roll path that cuts the subject out of a single still photo as a photographic sticker and generates each beat's poster around it.

## Sources

- [Alisa0808/vox-director on GitHub](https://github.com/Alisa0808/vox-director)
- [Issues](https://github.com/Alisa0808/vox-director/issues)
- [License: MIT](https://github.com/Alisa0808/vox-director/blob/main/LICENSE)
- [README](https://github.com/Alisa0808/vox-director/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/alisa0808-vox-director
