# gbro-collage-broll's three gates sit in front of one fixed 720x1280 clip

> This is an agent skill that turns a five second voiceover line into a halftone paper-collage animation, and its own documentation says the interesting part is not the prompt but the three approval gates it forces before any video is billed. That design is well argued. The wiring around it is rougher: Gate 2 depends on a Codex-only tool while the install line offers a Claude skills directory, the dependency install writes into another project's home directory, and the whole README is in Chinese.

**pyang5166/gbro-collage-broll** — 半调纸拼贴 B-roll 生成 skill：三闸门审批，Gemini Omni Flash 首尾帧组装动画 | Editorial halftone paper-collage B-roll agent skill

- Repository: https://github.com/pyang5166/gbro-collage-broll
- Stars: 1,314 · Forks: 114
- Language: Python
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/pyang5166-gbro-collage-broll

## The gates are the product, and the project says the prompt is not

The README states plainly that the core of this skill is not a prompt template but a forced three-stage approval, so that your attention goes to aesthetic judgement instead of generation cost. The stages are worth naming precisely.

Gate 1 is metaphor confirmation and it produces no images and no video at all. It outputs only the visual metaphor plan, broken into core meaning, key objects, base colour and assembly order. Nothing is billed, because nothing has been generated.

Gate 2 is still frame confirmation. Only after Gate 1 passes does it generate the colour collage still plus a contact sheet, then wait again.

Gate 3 is the video. Once the still is approved it runs `gemini-omni-flash-preview` for a first and last frame assembly animation and attaches the full QA, which is per-second frame extraction, first frame empty-field verification and last frame comparison.

Batch mode allows partial approval, so only the confirmed items advance to the next stage. That is the detail that makes the gate design workable on a multi-line script rather than only on a single line.

## Gate 2 depends on a Codex tool, and the install line offers a Claude directory

The requirements table has one row for the agent environment, and it says Gate 2 still frame generation depends on the built-in `image_gen` tool in the Codex environment.

The installation section then says to put the whole directory into your agent skills directory, giving `~/.agents/skills/` or `~/.claude/skills/` as the two examples, and the only command it offers is a clone into the first of those.

So one half of the documentation assumes Codex and the other offers a Claude Code path. Since Gate 2 is the gate that produces the still every later decision depends on, an agent running somewhere without `image_gen` has no described route through the pipeline. The trigger vocabulary in the usage section is language-neutral and the model is Gemini, so nothing else in the skill is Codex-specific except that one requirement.

There is exactly one agent interface file in the tree, `agents/openai.yaml`, described as the Codex interface configuration, which suggests Codex is the intended home and the Claude directory is offered out of habit rather than tested.

## The dependency install writes into another project's directory under your home

There is no installer script. Installation is a single clone:

```bash
git clone https://github.com/pyang5166/gbro-collage-broll.git ~/.agents/skills/gbro-collage-broll
```

Everything else is deferred to first use. On first trigger the skill runs `scripts/check_setup.sh` itself and gives configuration guidance for whatever is missing, which is a sensible design for something distributed as a directory rather than as a package.

The requirements are `GEMINI_API_KEY` created in Google AI Studio, with video generation billed by usage, Python 3.10 or newer, `google-genai` 2.10.0 or newer, and ffmpeg with ffprobe for the first and last frame handling, the audio strip and the contact sheet.

The detail to notice is where the Python side lands. The skill guides creation of a shared virtual environment at `~/hyperframes-projects/.omni-venv/`, a hardcoded absolute path whose directory is named for a different project. So installing this skill creates state outside its own directory, in a path shared with something else on your machine, and the self-check script is what reconciles it if that path is already there.

## The deliverable is one fixed shape: 720x1280, five seconds, 24fps, silent

The stated default output is 9:16 at 5 seconds, 720x1280, 24fps, with no audio, as an MP4 that can be laid directly under a voiceover. The input is equally narrow, described as a voiceover line of about five seconds.

That is a coherent pair. The skill is built for the vertical short-form case where a spoken line and a picture occupy the same five seconds, and the aspect ratio, duration and frame rate are all part of the same assumption.

What is not documented is how to depart from it. There is no documented override for resolution, frame rate, duration or aspect ratio anywhere in the requirements, the workflow or the FAQ. ffmpeg is listed for first and last frame processing, audio stripping and contact sheet generation rather than for rescaling the finished clip, so post-processing is left to whatever ffmpeg does next.

The visual target is equally specific, and it is the opposite of the default motion language. The look is a strong flat solid-colour paper field, black and white halftone photo cutouts and coloured cardstock accents, with elements sliding in one at a time from an empty field and seating themselves in a stop-motion texture rather than fading in or pushing slowly.

## The first frame check has a known cosmetic failure, and the fix is another tool

Gate 3 attaches QA that includes first frame empty-field verification, alongside per-second frame extraction and last frame comparison. The empty field is doing real work in this design: it is what makes the assemble-from-nothing motion legible, since the first frame is what tells the viewer the shot starts on bare paper.

The FAQ then describes the case where a little bit of paper shows at the edge of the first frame. It says a slight amount is acceptable, and that if you need a strict empty field you should use an editable-timeline animation tool to patch the front section.

That is a candid answer and also an admission. The QA gate verifies the property, and the property still fails cosmetically often enough to have its own FAQ entry, with the remedy being an external editor rather than a parameter on this skill.

The third FAQ entry is shorter and is about model choice. The default is fixed to `gemini-omni-flash-preview`, and it only switches when another model is explicitly named. So the model is pinned by default and changeable deliberately, which is the opposite arrangement from the three gates, where partial approval is the default and strictness is opt in.

## Three gates, four evals, and a legacy script that ships switched off

The documented structure is a skill directory holding `SKILL.md` as the main document with the three-gate protocol, the prompt template and the QA standard, one agent interface file at `agents/openai.yaml`, an eval file at `evals/evals.json` described as four gate behaviour evals, and four scripts: `check_setup.sh`, `generate_video.py`, `upload_file.py` and `generate_veo_first_last.py`.

Three counts do not line up and all three are worth a look. There are three described gates and four evals, so one gate carries more than one behaviour or the fourth covers something outside the three-stage story, and the eval file would settle it. The last script, `generate_veo_first_last.py`, is labelled the old Veo path, kept for compatibility only and unused by default, which means the default model is Gemini Omni Flash while a Veo path is still shipped in the tree. And the repository root also holds `assets/`, which the structure listing does not account for.

None of that is alarming. It is the shape of a small skill that grew: one compatibility path kept, one directory not yet written into the docs, and one eval count ahead of the documented protocol.

## The documentation is Chinese only, and the trigger phrases are Chinese

The repository has a single `README.md` and no English version and no translation directory, while the repository description carries both a Chinese sentence and an English one.

That asymmetry matters more here than it would for most projects, because the skill is invoked by phrase. The documented usage is to say a Chinese voiceover line to your agent, and the trigger words are `collage b-roll`, `纸拼贴 b-roll`, `半调拼贴`, `拼贴风格配画面` and `gbro-collage-broll`. Four of the five are Chinese, and the Chinese ones carry the actual intent, halftone collage and collage-style visual pairing, rather than just transliterating the package name.

So an English-only user has one line of description and a repository of Chinese prose, with the key operational detail being which phrases to type. The install command, the environment variables, the file paths and the output specification are all language-neutral and readable; the gate protocol and the trigger vocabulary are not.

The MIT licence is unaffected by any of this, and the last push to the default branch is dated 2026-07-15.

## Conclusion

Use gbro-collage-broll if you are cutting vertical short-form video with a spoken track underneath and you want a stop-motion assembly look rather than a fade or a slow push, and if you have a Codex environment available, because Gate 2 depends on its built-in `image_gen` tool and there is no fallback path described. Three things to settle first. The output is fixed at 9:16, 5 seconds, 720x1280, 24fps and silent, so any other shape needs work this skill does not document. The Gate 3 QA checks that the first frame is empty, and the FAQ admits a sliver of paper at the edge can still appear and points you at an external timeline tool to patch it. And the install is a plain `git clone` into a skills directory, with the Python dependency landing in a venv under someone else's project path.

## FAQ

### What are the three approval gates in gbro-collage-broll?

Gate 1 is metaphor confirmation and generates nothing, outputting only the core meaning, key objects, base colour and assembly order. Gate 2 generates the colour collage still plus a contact sheet and waits again. Gate 3 runs `gemini-omni-flash-preview` for the first and last frame assembly animation and attaches QA covering per-second frame extraction, first frame empty-field verification and last frame comparison. In batch mode only confirmed items advance.

### What does gbro-collage-broll need before it will run?

Clone the repository into a skills directory with `git clone https://github.com/pyang5166/gbro-collage-broll.git ~/.agents/skills/gbro-collage-broll`. It also needs a GEMINI_API_KEY from Google AI Studio, since video generation is billed by usage, Python 3.10 or newer, google-genai 2.10.0 or newer, and ffmpeg with ffprobe. Gate 2 still frame generation depends on the built-in image_gen tool in a Codex environment. The skill runs scripts/check_setup.sh on first trigger.

### What format does gbro-collage-broll output?

The default deliverable is a 9:16 MP4 at 5 seconds, 720x1280, 24fps, with no audio, intended to sit under a spoken track. The motion style is elements sliding in and seating themselves from an empty paper field, a stop-motion texture rather than a fade-in or a slow zoom. No override for resolution, duration or aspect ratio is given.

### Can gbro-collage-broll use a video model other than Gemini Omni Flash?

Not by default. The model is fixed to `gemini-omni-flash-preview` and it only switches when another model is explicitly specified. A legacy Veo script, scripts/generate_veo_first_last.py, ships in the repository for compatibility but is unused unless you ask for it.

## Sources

- [Issues](https://github.com/pyang5166/gbro-collage-broll/issues)
- [License: MIT](https://github.com/pyang5166/gbro-collage-broll/blob/main/LICENSE)
- [pyang5166/gbro-collage-broll on GitHub](https://github.com/pyang5166/gbro-collage-broll)
- [README](https://github.com/pyang5166/gbro-collage-broll/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/pyang5166-gbro-collage-broll
