# The Agentic Video Editor's retry loop is a model scoring a model, gated on an undefined composite

> A command-line editor that points at a folder of clips, takes a creative brief, and runs four agents, three of which are the same model, with the fourth rendering through FFmpeg. The retry mechanism is the interesting part and it is fully configurable, which is the problem: the gate that ends the loop is a composite score whose formula the documentation never states. Two documented pipelines disagree on the threshold.

**poseljacob/agentic-video-editor** — AI-powered video editor that turns raw footage and a creative brief into a polished ad using an ensemble of AI agents (Google Gemini + FFmpeg)

- Repository: https://github.com/poseljacob/agentic-video-editor
- Stars: 496 · Forks: 63
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/poseljacob-agentic-video-editor

## The loop ends on a composite score whose formula is never stated

Five dimensions are scored on a zero to one scale, and the fifth is the one that matters.

Adherence asks whether the edit followed the brief. Pacing asks whether shot durations and energy are balanced. Visual quality asks whether cuts are clean and transitions smooth. Watchability asks whether a viewer would watch to the end. And Overall is described as a composite score.

The pipeline gate is set on Overall. The shipped pipeline manifest reads:

```yaml
  - agent: reviewer
    retry_if:
      metric: overall
      threshold: 0.65           # retry if overall score < 0.65
      max_retries: 2            # up to 2 retries (3 total passes)
      feedback_target: director # send reviewer feedback back to director
```

So the number that decides whether the pipeline runs again is a composite of a model's four judgements about a video. Two of those four are not observable in the render. You cannot measure pacing from a file, and you cannot measure whether a viewer would watch to the end from a file. They are the model's predictions.

And how the four become one is not documented. A mean, a weighted mean, a minimum, a median: these give materially different retry behaviour, and a threshold of 0.65 means something different under each. The documentation gives the composite a name and nothing else.

The custom-pipeline example compounds the uncertainty by choosing a different number, 0.7, with three retries instead of two. Both values are presented without justification and neither is described as a default. So the first thing anyone tuning this has to do is discover what score their own footage normally earns, with no reference point to start from.

## Retry feedback goes to the Director, so a bad cut triggers re-selection of shots

Read the feedback target again: it is the Director.

The pipeline order is Director, Trim Refiner, Editor, Reviewer. If the Reviewer's complaint were about a cut boundary, the natural fix would be to re-run the Trim Refiner, which is the agent that exists for exactly that. The configuration sends the feedback to the Director instead.

The Director is the expensive step and the least local one. It searches the footage index, selects shots, decides their order, and writes the plan with trims and text overlays. Re-running it produces a different selection, which means the Trim Refiner's work is discarded and the Editor re-renders, and the previous plan is thrown away entirely.

That is a defensible choice when the failure is selection, which is the Reviewer's most common complaint in practice: a video that does not follow the brief usually got the wrong shots, not badly trimmed shots. But it means the pipeline has one retry resolution rather than per-step resolution, and the operator cannot redirect a visual-quality complaint at the trimming agent without editing the manifest.

The retry arithmetic is also worth knowing before you budget anything. The comment on the shipped manifest says two retries means three total passes, so a maximum-retry setting of three means four renders. Each retry is saved as a separate file, named with an incrementing version suffix, in the output directory, so a four-pass run leaves four movies and four sets of scores rather than one movie that improved.

Nothing in the documentation says which of those four is the deliverable. A user has to pick.

## The documented pip install cannot install the dev group, because dev is not an extra

The setup section gives two ways to install, and only one of them does what the surrounding text implies.

The first is a virtual environment and an editable install with a development extra:

```bash
python -m venv .venv
source .venv/bin/activate   # or .venv\Scripts\activate on Windows
pip install -e ".[dev]"
```

The second is a single sync command using the other package manager, followed by activating the same environment.

Now look at what the manifest declares. There is no optional-dependencies section at all. The development requirement appears under a dependency-groups table, holding one entry, a test runner pinned to version 9.0.3 or newer. Dependency groups and installable extras are different mechanisms, and pip resolves the extra form from the optional-dependencies table, which does not exist.

So the first command installs the project and does not install the test runner. The second command does install it, because that tool installs declared dependency groups by default. The two documented methods are not equivalent, and the one presented as the plain option is the one that misses.

The cost is small in absolute terms, since the development group is a single package and the README tells you the test command separately. It is worth knowing because the failure is silent: pip reports success, and the missing test runner surfaces later as a command not found when you follow the instructions for running tests.

A lockfile is committed for the second path, so the reproducible install is the one that works.

## A web server is a runtime dependency for an interface the documentation tells you not to use

The dependency list has twelve entries, and three of them exist only for the web interface.

The web interface is described as pre-alpha, work in progress, explicitly not the recommended way to use the project, with rough edges, missing features and breaking changes expected. The README is unusually clear about this, and states twice that the command line is the supported path.

Still in the runtime dependencies, not in an optional group, are a web framework, an ASGI server with its standard extras, and a WebSocket library. So every install of the command-line tool pulls in a web server, its reload and HTTP machinery, and a WebSocket implementation for a component the same document says not to use.

The interface itself is not small, which makes the dependency weight harder to justify. It is described as a traditional non-linear editor layout: a project picker, separate source and program monitors, a drag-and-drop timeline, a media browser, an inspector, and a radar chart for the review scores.

Running it needs a second toolchain on top of the Python environment: a Node runtime of 18 or newer, that ecosystem's package manager, and two processes running at once, an ASGI server on one port with reload enabled and a frontend dev server on another. No production run path is documented, only the two development servers.

There is also a behavioural difference between the two interfaces that goes unremarked. The backend exposes a REST surface covering jobs, projects, footage, feedback and full create, read, update and delete on edit plans. So in the web interface a human can rewrite the plan the Director produced. The command-line path has no equivalent, because on that side the plan is whatever the agent decided. The same pipeline produces different artefacts depending on which front end you use.

## The wheel ships a top-level package called src

Two configuration lines decide the installable identity of this project, and both point at a directory literally named src.

The console script is declared as a dotted path into that package, and the wheel target lists that one package. Nothing renames it on the way out.

So the installed distribution contains a top-level importable package called src. That is a known collision hazard rather than a stylistic choice: the name is not namespaced, so on any environment where two installed distributions both ship a top-level src, they fight over the same location on the import path, and which one wins depends on installation order rather than on anything the user controls.

The architecture section makes it easy to see why the name ended up this way. The directory holding the entry point also holds the agents, the schemas, the pipeline runner, the tool functions and the web backend. The web frontend lives inside it too, under a nested studio directory with its own package manifest, which is why the top-level listing shows no JavaScript project at all and a second toolchain is invisible until you read the README's optional section.

The entry point itself is fine. A console script mapped to a click command in the entry module is the conventional shape, and the click dependency in the list confirms it.

The collision is the part to raise before installing this alongside anything else in a shared environment, because it fails at import time with a message that points at the wrong package rather than at the collision.

## The footage index is cached between runs and no invalidation rule is given

The first stage of every run scans the footage folder, detects scenes, transcribes speech and writes an index file, and that index is described as cached between runs.

That stage is the expensive one. Scene detection runs an OpenCV-backed detector over every clip, and transcription runs a Whisper model over the audio. Both are orders of magnitude slower than everything downstream, which is presumably why the cache exists.

But no invalidation rule is documented. Not a modification-time check, not a file count check, not a flag to force a rebuild. The key is the footage directory path and nothing else is said.

The failure mode follows from that. Add a clip to the folder, run again, and the index does not mention it, so the Director cannot select it and the edit is silently missing footage. Remove a clip and its scenes stay in the index, so the Director may select shots from a file that is no longer there and the render produces a gap or an error. Neither case is loud.

The same issue applies to edited footage. Re-export a clip with different content under the same filename and the index still describes the old version, so the transcription the Director reasons over does not match the pixels being rendered.

This is the failure mode most likely to cost someone an afternoon on a first real run, because the tool produces a finished-looking movie either way. The documentation says the index is cached and stops there.

## Three of the four agents are the same model behind one API key

The agent table lists four roles with four power sources, and reading the power column is the fastest way to understand the system's real shape.

The Director is credited to one vendor's agent development kit. The Trim Refiner is credited to that same model. The Editor is credited to FFmpeg and MoviePy, which is not a model at all. The Reviewer is credited to the same model again, and it watches the rendered video.

So three of four agents are instances of one model, and the fourth is a renderer. The project description calls this an ensemble of AI agents, and the summary sentence above the architecture says the system is built with one vendor for intelligence and FFmpeg for rendering. Both statements are accurate, and together they are narrower than the word ensemble suggests.

The practical consequences are all about that one dependency. The environment example file contains exactly one variable, a key for that vendor's model service. There is no configuration for an alternative model, no option to swap the reviewer, and no local model path. Every run therefore requires that service, and the Reviewer's video consumption means the multimodal input path is on the critical path for every pass, not just the first.

The cost profile follows the same shape. A four-pass run on the shipped settings is four Director calls, four Refiner calls and four Reviewer calls that each ingest a video, against three renders. The agents are the expensive part of the loop, not the rendering.

Whether the Reviewer runs through the same agent kit as the Director is not stated; only the Director's row names it, while the other rows say only the model.

## One style file, one pipeline file, and a music mood parameter with no consumer

Two directories carry the configuration surface, and each contains exactly one example file.

The pipeline directory holds a single manifest, the one for user-generated-content direct-to-consumer advertising. The style directory holds a single file, which defines a 30-second structure of five beats: a hook, a problem, a solution, social proof, and a call to action.

That is a complete worked example for one genre, which makes it a good tutorial and a thin library. The mechanism is general; the inventory is one.

The style file is described as carrying segment durations, pacing rules, text overlay placement and music mood. Three of those four connect to something in the documented pipeline: durations and pacing feed shot selection, and overlay placement matches the plan the Director writes. Music mood does not.

Nothing in the four agents generates, selects or adds music. There is no audio stage, no music step in the manifest, and no music tool in the tool list. So a documented style parameter has no consumer in the shipped pipeline, which means either the pipeline is incomplete relative to the style schema or the schema is aspirational.

Overlay placement has its own gap. The plan carries text overlays, and a subtitle library is in the dependency list, but nothing in the documentation configures a font, a brand colour, a safe area or a logo. The brief schema has no field for any of them either: it takes a product name, an audience, a tone, a duration and a style reference. So a campaign that has a call-to-action segment structurally has nowhere to put the call-to-action text.

One last wrinkle in how configuration arrives. The brief schema includes a style reference field, while the worked command passes the style as its own flag. Both routes exist and nothing says which takes precedence, or whether they must agree.

## Conclusion

Use it if you want to see an agent-driven edit pipeline end to end and are prepared to tune the retry threshold against your own footage, since the gate is the only quality control and it is a number you set. Do not adopt it expecting reproducible output, because the Director's shot selection is model-driven and the Reviewer's score decides whether it runs again. Verify first how the composite score is computed and what a typical good edit scores on your material, since no baseline is given and two shipped examples disagree.

## FAQ

### How does the agentic video editor decide to retry an edit?

A Reviewer agent watches the rendered video and scores it on five zero-to-one dimensions: adherence, pacing, visual quality, watchability and an overall composite. The pipeline manifest gates on the overall metric, with a threshold and a retry count you set, and the shipped example uses 0.65 with two retries, meaning three passes.

### What does the agentic video editor need before it will run?

Python 3.11 or newer, FFmpeg installed and on the path, and a Google AI API key placed in an environment file. There is no offline mode and no alternative model configuration, so all three are required. FFmpeg must be a system install rather than something the package manager provides.

### Which agents does the agentic video editor use and what powers them?

Four agents. The Director searches the footage index, selects shots and writes an edit plan, powered by Google Gemini through the agent development kit. The Trim Refiner adjusts cut points, the Editor renders through FFmpeg and MoviePy, and the Reviewer watches the finished video and scores it. Three of the four are the same model.

### How do I install the agentic video editor for development?

Clone the repository, create a virtual environment and install it in editable mode. Note that the documented pip command requests a development extra, but the manifest declares the test runner as a dependency group rather than an extra, so that form does not install it. The alternative single-command sync with the other package manager does install it, and a lockfile is committed for that path.

### Is the AVE Studio web interface for the agentic video editor ready to use?

It is described as pre-alpha and work in progress, explicitly not the recommended way to use the project, with rough edges, missing features and breaking changes expected. The command line is the supported path. Running the web interface additionally requires a Node runtime of 18 or newer and that ecosystem's package manager, with a backend and a frontend dev server running at once.

## Sources

- [Issues](https://github.com/poseljacob/agentic-video-editor/issues)
- [License: MIT](https://github.com/poseljacob/agentic-video-editor/blob/main/LICENSE)
- [poseljacob/agentic-video-editor on GitHub](https://github.com/poseljacob/agentic-video-editor)
- [README](https://github.com/poseljacob/agentic-video-editor/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/poseljacob-agentic-video-editor
