# bradautomates/claude-video: giving Claude Code the ability to watch a video

> The /watch skill downloads a video, extracts frames, transcribes the audio, and hands both to Claude. It is a vision and transcript pipeline for agent hosts, not a video generator, and its frame budget decides whether an answer is useful.

**bradautomates/claude-video** — Give Claude the ability to watch any video. /watch downloads, extracts frames, transcribes, hands it all to Claude.

- Repository: https://github.com/bradautomates/claude-video
- Stars: 17,868 · Forks: 1,841
- Language: Python
- License: MIT
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/bradautomates-claude-video

## What /watch actually solves, and who it is built for

Claude can read a webpage, run a script, and browse a repository. What it cannot do by default is watch a video. The README frames the gap directly: paste a YouTube link and the model either guesses from the title or pulls a transcript that, in the project's words, is missing 90% of what's on screen. A transcript alone tells you what was said, not what was shown, and for screen recordings, ad creative, launch videos and bug reproductions the visual track is the part that matters.

The project is a skill, not a standalone application. It is aimed at people who already drive Claude through a host that supports skills: Claude Code first, then Codex, Cursor, Copilot, Gemini CLI, or any of the 50+ Agent Skills hosts the README links to. Within that audience the intended jobs are concrete. Analyze someone else's content and ask what hook they opened with. Diagnose a bug from a screen recording someone sent you. Summarize a long video. Strip the hype out of an update video. Turn a playlist into per-video notes.

The scope is deliberately narrow. This is video analysis, not video generation and not video editing. If you arrived looking for a tool that produces clips, the repository does not do that, and the search traffic around "claude video generation" is pointed at something this project is not.

## The pipeline: captions first, frames second, Whisper only as fallback

The mechanism is a six-step data flow, and the ordering is the interesting part. You paste a video and a question. yt-dlp checks captions first: at `transcript` detail, a captioned URL returns without downloading video at all. Otherwise, or when Whisper needs audio, it downloads only what the run needs. ffmpeg then extracts frames at the chosen detail level. `efficient` decodes keyframes only and the README describes it as near-instant. `balanced` and `token-burner` prefer scene-change frames and fall back to a duration-aware uniform sampler when scene detection under-produces. JPEGs are 512px wide by default and clamped to 1998px tall for Claude Read compatibility.

The transcript comes from one of two places. First, yt-dlp pulls native captions, manual or auto-generated, from the source: free, instant, and in the README's own phrasing, accurate-ish. If there are none, the fallback extracts a mono 16 kHz 64 kbps mp3 clip of roughly 480 kB per minute and ships it to Whisper, either Groq's `whisper-large-v3` (preferred, described as cheaper and faster) or OpenAI's `whisper-1`.

Frames and transcript are then handed to Claude. The script prints frame paths with `t=MM:SS` markers and the transcript with timestamps, and Claude reads each frame in parallel because JPEGs render directly as images in its context. The answer is grounded in the frames and the audio rather than in the title or a description. Cleanup is the last step: the script prints a working directory, and if you are not asking follow-ups, Claude removes it. That detail matters more than it looks, because it means a one-shot question leaves nothing behind while a follow-up conversation keeps the extracted assets around.

## Installing the skill and running a first /watch

The recommended path is the Claude Code plugin marketplace, which the README says auto-updates. Two commands, run inside Claude Code, register the marketplace and install the skill.

```bash
/plugin marketplace add bradautomates/claude-video
/plugin install watch@claude-video
```

For Codex, Cursor, Copilot, Gemini CLI, or other Agent Skills hosts, the README gives a single npx command. The `-g` flag installs globally for your user so the skill is available across all projects; dropping it scopes the install to the current project.

```bash
npx skills add bradautomates/claude-video -g
```

There is no configuration step to get started. On first run the skill installs `yt-dlp` and `ffmpeg` via `brew` on macOS, and on Linux and Windows it prints the exact commands rather than running them, so you install the two binaries yourself before the first successful run.

A first real use is a URL plus a question. The README's own example asks about a specific moment, which is the shape of question this tool answers best.

```
/watch https://youtu.be/dQw4w9WgXcQ what happens at the 30 second mark?
```

What you should see is a Frames line reporting how many frames were selected from how many candidates, along with the transcript, followed by Claude's answer. A local file works the same way: `.mp4`, `.mov`, `.mkv` and `.webm` are the formats the README names.

```
/watch bug-repro.mov what's going wrong?
```

When you already know where the interesting part is, the README says to pass `--start` and `--end`. Focused mode gets denser per-second budgets, capped at 2 fps, which is far more useful than a sparse pass over a long video.

## The frame budget is the real constraint, not the download

Token cost is dominated by frames, and the README is unusually candid about it: every frame is an image and image tokens add up fast. The auto-fps logic exists so a run does not blow its context budget on a sparse scan of a 30-minute video that a focused 30-second window would have answered better.

The published budget table is the number to internalize. A clip of 30 seconds or less gets roughly 30 frames, which is dense enough to cover basically every key moment. Between 30 seconds and a minute it is about 40 frames. One to three minutes gets about 60. Three to ten minutes gets about 80, which the README itself calls sparse but workable. Past ten minutes the capped modes stop at 100 frames and the script emits a "sparse scan" warning, telling you to re-run focused or switch to `--detail token-burner` for full uncapped coverage.

The practical consequence is that this tool is at its best on short material or on a named window inside long material, and at its weakest when you ask a broad question about a long video and accept the default. The warning is the honest signal here. If you see it, the run has already told you the answer may be built on a thin sample. That is a design trade-off rather than a defect, but it is the trade-off that determines whether you get a useful answer, and the README does not pretend otherwise.

## Deduplication: dropping near-identical frames before they are billed

Frame selection can surface near-identical frames even when it is working correctly. The README's example is a screen recording that holds one slide for 90 seconds, which produces a dozen frames each billed as a separate image. The dedup pass exists to drop them before they reach Claude, and it runs by default on every frame mode. `--no-dedup` turns it off.

The implementation is deliberately small. One ffmpeg call scales each extracted JPEG to a 16x16 grayscale thumbnail. Everything after that is pure-stdlib Python, with no image libraries involved. For each frame the script computes the mean absolute difference against the last frame that was kept, which is the average per-pixel brightness change on a 0 to 255 scale. If that difference is at or below the threshold of `2.0`, the frame is a near-duplicate and is dropped. Otherwise it is kept and becomes the new reference. The frame-budget cap applies after dedup, so the budget is spent on distinct frames rather than on duplicates.

Two choices in that design are worth calling out. Comparing against the last kept frame rather than the immediately previous one catches slow fades that never trip a frame-to-frame threshold. And the threshold is deliberately low and measures absolute brightness rather than structure, so a one-line code diff, a terminal scrolling a single row, or two differently colored flat slides all survive the pass. The Frames line reports the collapse, for example `6 selected from 14 candidates (… 8 near-duplicates dropped …)`. On always-moving footage nothing is dropped and you pay what you would have paid anyway, so the pass has no downside in that case.

## Where /watch is the wrong tool

The clearest limitation is the one the project states outright: this is not video generation and not video editing. It reads video, it does not produce or modify it. Anyone searching for a Claude video editor will find nothing here.

The second limitation is cost and dependency shape. Captions cover most public videos for free, but a Whisper API key is required whenever a video has no captions, and that is a paid third-party call. If your material is private, uncaptioned, and you cannot send audio to Groq or OpenAI, the fallback is unavailable to you.

The third is the long-video case already described. Above ten minutes the capped modes give you 100 frames and a warning, and `token-burner` trades that cap for token spend. Neither is free, and a broad question about a 40-minute video is the scenario where this tool is most likely to disappoint.

The fourth is platform friction. Zero config applies to macOS, where `brew` installs `yt-dlp` and `ffmpeg` on first run. On Linux and Windows the skill prints the commands instead of running them, so the first run is not unattended. The README does not document rollback or uninstall steps for the plugin or the skill, and it does not document what happens when `yt-dlp` fails against a site it does not support.

## How it compares to a plain transcript tool

The obvious alternative is a transcript-only workflow: pull captions with yt-dlp and hand the text to Claude, skipping ffmpeg and the frame extraction entirely. The difference in approach is exactly the gap the project was built to close. A transcript pipeline is cheaper, faster, and needs no image tokens at all, and for a podcast or a talking-head interview it may be sufficient. What it cannot do is answer a question about what is on screen. The README's stated failure mode is a transcript missing 90% of what is on screen, and that is the case a frame pipeline exists to cover.

The second alternative is a general-purpose video understanding API, where you upload the file and a hosted model returns a description or answers questions about it. That approach moves the frame selection and token budgeting to the vendor and asks you to trust its sampling. This project keeps the sampling in your hands: you choose the detail mode, the start and end window, and whether dedup runs, and you can see the Frames line reporting what was actually selected. The trade is that you also carry the responsibility for choosing well, and the sparse-scan warning is the consequence of choosing badly.

A third option is simply watching the video yourself, which remains the correct answer when the question is judgment-heavy and the video is short enough that the overhead of setting up a pipeline exceeds the time saved.

## Maintenance, licence, and what an upgrade costs you

The repository is not archived. Its last push was on 2026-07-01, which is the same date as the v0.2.0 release, and before that v0.1.3 landed on 2026-05-08 and v0.1.2 on 2026-04-24. That is a release cadence of roughly one minor or patch version every few weeks across the visible window, with the most recent work arriving alongside a minor version bump. The README does not document a deprecation policy, a support window, or a migration guide between versions, so an upgrade path from v0.1.x to v0.2.0 is not described in the project's own files.

The project is MIT licensed, which permits commercial use, modification, and redistribution provided the licence and copyright notice are preserved. That covers the skill's own code. It does not cover the third-party services the skill calls: Whisper through Groq or OpenAI is governed by those providers' terms, and yt-dlp's downloads are subject to the terms of the sites it supports. Nothing in the repository changes those obligations, and this is not legal advice.

Upgrade cost is low in the Claude Code path, since the README states the plugin auto-updates via the marketplace. The npx path and manual installs are yours to refresh. The dependency surface is small, which is the main reason upgrades should be uneventful: yt-dlp, ffmpeg, and optionally a Whisper key. The risk sits in yt-dlp, whose extractors track site changes continuously, so a stale install is more likely to break on a specific site than the skill itself is to break.

## Conclusion

Adopt bradautomates/claude-video if you already run Claude Code or another Agent Skills host and your questions are about what is on screen in a specific video, especially short clips or a named time window. Do not adopt it if you want video generation or editing, and do not expect it to be free at scale: captions are free, but frames cost image tokens and Whisper is a paid API fallback. Before relying on it, verify that yt-dlp and ffmpeg install cleanly on your platform, that the videos you care about have captions so the Whisper path stays unused, and that the default frame budget matches how long your videos are. The 100-frame cap above ten minutes is the number to check first.

## FAQ

### Can Claude view a video?

Not on its own, which is the gap this project fills. With the /watch skill installed, Claude downloads the video, extracts frames with ffmpeg, pulls a timestamped transcript from captions or Whisper, and reads each frame as an image before answering.

### Can Claude generate video?

No. bradautomates/claude-video is a video analysis skill: it downloads, extracts frames, transcribes and reads. The README describes no generation or editing capability.

### Is the Claude video skill free?

Captions cover most public videos at no cost, and yt-dlp and ffmpeg are free tools. A Whisper API key is only needed when a video has no captions, and that path calls Groq's whisper-large-v3 or OpenAI's whisper-1, which are paid services. Frames also consume image tokens in Claude's context.

### How to use Claude video?

Install the skill, either through the Claude Code plugin marketplace or with npx skills add bradautomates/claude-video -g, then paste a URL or local path with a question. The README's example is /watch https://youtu.be/dQw4w9WgXcQ what happens at the 30 second mark?

### What is claude video?

It is the /watch skill from bradautomates/claude-video, an MIT-licensed Python Agent Skill that gives Claude Code and other Agent Skills hosts the ability to watch a video by extracting frames and transcribing audio.

## Sources

- [Official README](https://github.com/bradautomates/claude-video#readme)
- [Project repository](https://github.com/bradautomates/claude-video)
- [Release notes](https://github.com/bradautomates/claude-video/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/bradautomates-claude-video
