CLI tool
browser-use/video-use avatar
browser-use/video-use

video-use: a Claude Code skill that edits video from a transcript instead of frames

video-use gives coding agents commands for cutting, arranging, captioning, and rendering video projects.

27,585 stars3,254 forksPythonMIT

At a glance

What is it?
video-use is an MIT-licensed Python skill that gives coding agents cut, caption and render commands for raw footage. Its core bet is that an LLM should read a transcript, not watch frames.
Who is it for?
Adopt video-use if you already run Claude Code, Codex, Hermes or Openclaw on a machine with ffmpeg and you have talking-head, interview or tutorial footage where filler words and dead space are the main editing problem. Skip it if your footage needs frame-level visual judgement, if you cannot send audio to ElevenLabs, or if you want a GUI where you drag clips on a timeline.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

The editing problem video-use actually targets

Most editing time on talking-head footage is not creative. It is deleting umm, uh and false starts, tightening the silence between takes, and making sure no cut pops. video-use is built for exactly that job. The README lists cutting filler words and dead space as the first capability, and the design principles state that cuts come from speech boundaries and silence gaps, with audio treated as primary and visuals following it. That ordering explains the whole tool. If your edit decisions are driven by what was said, this fits. If they are driven by what is on screen, it does not.

The intended user is someone who already has a coding agent open in a terminal. Setup is a prompt pasted into Claude Code, Codex, Hermes, Openclaw or any agent with shell access, and the README says the agent handles the clone, dependencies, skill registration and the one-time API key prompt. You point the agent at a folder of raw takes and type something like "edit these into a launch video". There is no timeline UI, no preset menu and no content-type assumption. The README states it works for talking heads, montages, tutorials, travel and interviews.

Reading video instead of watching it: the two-layer model

The claim that carries this project is in the README: the LLM never watches the video, it reads it. Two layers make that work.

Layer one is always loaded. One ElevenLabs Scribe call per source returns word-level timestamps, speaker diarization and audio events such as (laughter), (applause) and (sigh). All takes pack into a single roughly 12KB takes_packed.md, which the README calls the LLM's primary reading view. The format is compact, and the README shows an example line with a take identifier, a duration, a phrase count, and bracketed time ranges prefixed by a speaker label.

Layer two is on demand. A command called timeline_view renders a filmstrip plus waveform plus word labels as a PNG for any time range. The README says it is called only at decision points: ambiguous pauses, retake comparisons, cut-point sanity checks. The token argument is stated plainly in the README as a comparison, not a benchmark: a naive approach of 30,000 frames at 1,500 tokens each would be 45M tokens of noise, against 12KB of text plus a handful of PNGs. The README frames this as the same idea as browser-use giving an LLM a structured DOM instead of a screenshot, applied to video.

That is a real architectural choice with a real cost. Anything the transcript does not encode is invisible until someone calls timeline_view. A jump cut that is obvious to a human eye but not reflected in the speech is not something the transcript layer can flag.

The pipeline, the EDL and the self-eval loop

The README gives the pipeline as Transcribe, Pack, LLM Reasons, EDL, Render, Self-Eval, with a loop back from self-eval to fix and re-render, capped at three attempts. EDL stands for edit decision list: the agent's reasoning step produces a machine-readable cut list, and rendering is a separate stage that consumes it. That separation is what makes the self-eval loop possible at all, since the renderer can be re-run against a corrected list without redoing the reasoning.

The self-eval step runs timeline_view on the rendered output at every cut boundary. The README says it catches visual jumps, audio pops and hidden subtitles, and that you see the preview only after it passes. Two production details in the README support this: 30ms audio fades at every cut so you never hear a pop, and burned subtitles defaulting to 2-word uppercase chunks. The README also mentions animation overlays generated through HyperFrames, Remotion, Manim or PIL, spawned in parallel sub-agents with one per animation. Manim is listed as an optional dependency under the animations extra in pyproject.toml, so that path needs an extra install step.

Session state persists in project.md, which the README says lets next week's session pick up where you left off. Outputs go to <videos_dir>/edit/, keeping the skill directory clean.

Installing video-use and running a first edit

The README offers two paths. The fast one is a setup prompt pasted into an agent with shell access. The manual path is documented step by step, and that is the one worth following if you want to see what lands on disk.

First, clone the repository and symlink it into your agent's skills directory. The README gives both the Claude Code and Codex destinations, with the Codex line commented out:

bash
git clone https://github.com/browser-use/video-use ~/Developer/video-use
ln -sfn ~/Developer/video-use ~/.claude/skills/video-use        # Claude Code
# ln -sfn ~/Developer/video-use ~/.codex/skills/video-use       # Codex

Next, install dependencies. The project uses uv, with a pip fallback, and ffmpeg is marked required. yt-dlp is optional and only needed for downloading online sources:

bash
cd ~/Developer/video-use
uv sync                         # or: pip install -e .
brew install ffmpeg             # required
brew install yt-dlp             # optional, for downloading online sources

The brew commands are macOS-oriented as written. On Linux you would install ffmpeg through your distribution's package manager instead; the README does not cover that case.

Finally, create the environment file and add your key. The repository ships .env.example containing a single key, ELEVENLABS_API_KEY, with an empty value:

bash
cp .env.example .env
$EDITOR .env                    # ELEVE

The README's own snippet is truncated at that last comment, so treat the trailing token as a display artifact rather than a command. You need a key from elevenlabs.io/app/settings/api-keys.

With setup done, the first real use is short. Change into the folder holding your raw takes and start your agent there:

bash
cd /path/to/your/videos
claude    # or codex, hermes, etc.

Then ask for the edit in plain language, for example "edit these into a launch video". The README describes what should happen next: the agent inventories the sources, proposes a strategy, waits for your approval, and then produces edit/final.mp4 next to your sources. If you never get that file, the failure is somewhere in transcription, the EDL or the render stage, and the README does not document a rollback path for a bad render.

Where video-use is the wrong tool

The transcript-first design has a hard boundary. Because cuts are derived from speech boundaries and silence gaps, footage where the interesting edit is visual will not be served well. A skate montage, a product b-roll sequence or a scene where the right cut point is a gesture rather than a sentence gives the primary layer nothing to work with. timeline_view exists, but the README positions it as a check at decision points, not as the surface the agent reasons over.

There is a dependency risk too. Transcription runs through ElevenLabs Scribe, so audio leaves your machine and the workflow stops without a valid key. The README does not describe an offline or local transcription fallback. If your footage cannot be sent to a third-party API, this pipeline is closed to you as documented.

Cost and latency scale with source length in a way the README does not quantify. One Scribe call per source, plus a render, plus up to three self-eval and re-render cycles, is a lot of round trips for a long interview. The repository also carries no release beyond 0.1.0 in pyproject.toml and no retrieved releases, so expect to read helpers/ and SKILL.md rather than trust a stable interface. The README explicitly tells the agent to always read helpers/ because that is where the editing scripts live, which is a fair signal that the scripts are the real documentation.

How this differs from an agent-driven editor like ffmpeg-mcp

The closest alternative in kind is an MCP server that exposes ffmpeg operations as tools, letting an agent call trim, concat and overlay directly. The difference is where the reasoning lives. With a tool server, the agent plans each operation itself and there is no intermediate artifact between the plan and the render. With video-use, the LLM's reasoning is committed to an EDL before rendering starts, and the render is a separate stage that the self-eval loop can challenge and re-run.

That changes what failure looks like. A tool-server agent that misjudges a cut point usually produces a bad file with no check between intent and output. video-use's loop runs timeline_view on the rendered output at every cut boundary and, per the README, shows you the preview only after it passes. The trade-off is that you inherit the project's own production rules: 12 hard rules, per the README, with artistic freedom elsewhere. An agent calling ffmpeg primitives has no such rules and no self-eval, but it also has no pipeline to fight when your footage does not fit the model.

Editorial conclusion

Adopt video-use if you already run Claude Code, Codex, Hermes or Openclaw on a machine with ffmpeg and you have talking-head, interview or tutorial footage where filler words and dead space are the main editing problem. Skip it if your footage needs frame-level visual judgement, if you cannot send audio to ElevenLabs, or if you want a GUI where you drag clips on a timeline. Before committing, verify three things yourself: that install.md and SKILL.md actually exist at the paths the README names, that ELEVENLABS_API_KEY is populated so the Scribe transcription step can run, and that the render step produces edit/final.mp4 next to your sources on a short test clip. The repository carries no version tag beyond 0.1.0 in pyproject.toml, so pin the commit you install.

Frequently asked questions

What is video-use used for?

It gives coding agents commands for cutting, arranging, captioning and rendering video projects. The README lists cutting filler words and dead space, auto color grading, 30ms audio fades at cuts, burned subtitles and animation overlays.

How do I use video-use?

Paste the setup prompt from the README into an agent with shell access, then point that agent at a folder of raw takes and describe the edit you want. The README says the agent inventories the sources, proposes a strategy, waits for approval, and produces edit/final.mp4 next to your sources.

Does video-use work with Codex as well as Claude Code?

The README's setup prompt names Claude Code, Codex, Hermes, Openclaw or any agent with shell access, and the manual install shows a Codex symlink target alongside the Claude Code one. The symlink line for Codex is commented out in the README snippet.

What does video-use need installed before it can edit anything?

The manual install runs uv sync (or pip install -e .) and installs ffmpeg, which the README marks required. yt-dlp is optional and only for downloading online sources, and Manim is an optional extra for animation overlays.

Does video-use send my footage anywhere?

Transcription uses one ElevenLabs Scribe call per source, and the README's setup requires an ELEVENLABS_API_KEY in a .env file. The README does not describe an offline transcription option.

Official sources

  1. Official README
  2. Project repository
Community notes

Community notes