Model or dataset
HUANGCHIHHUNGLeo/claude-real-video avatar
HUANGCHIHHUNGLeo/claude-real-video

claude-real-video: scene-aware keyframes and transcripts for LLMs

Let Claude (or any LLM) actually watch a video — scene-aware, deduplicated frames + transcript, from a URL or local file. Runs locally, MIT.

2,190 stars196 forksPythonMIT

At a glance

What is it?
claude-real-video is a Python CLI that turns a video URL or local file into deduplicated scene-change frames plus a transcript, so an LLM can read a video instead of just its captions. The processing runs locally, the licence is MIT, and the interesting design choice is frame selection by scene change rather than a fixed sampling rate.
Who is it for?
Adopt claude-real-video if you already have ffmpeg and Python 3.10 or newer and you want a local, MIT-licensed way to hand an LLM the frames that changed plus a transcript, with no upload of the source video.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap claude-real-video is aimed at

Paste a YouTube link into most chat tools and the model reads the transcript, not the picture. The README states that Claude will not take a video file at all, and that Gemini samples video at a fixed interval of 1 fps by default, which means fast cuts fall between samples. claude-real-video takes the opposite position: it extracts every scene change rather than a fixed quota, discards near-duplicate frames, transcribes the audio, and writes a folder that any LLM can read. The audience is engineers who already drive a coding agent and want to paste a link rather than describe a video from memory. It is also positioned as a general-purpose keyframe extractor for people who are not doing LLM work at all, since scene detection and dedup need no model download.

Scene-change extraction instead of fixed-rate sampling

The mechanism is a pipeline rather than a model. A URL is fetched with yt-dlp; a local path is read directly. ffmpeg then seeks and decodes, and frame selection is driven by scene-change detection with a dedup pass on top. The README gives one concrete comparison: for a 58-second clip, fixed 1 fps sampling yields 58 frames, while crv keeps the 26 that actually differ, and --grid packs those into 3 contact sheets. Whisper handles the audio side and writes transcript.txt and transcript.json. Output lands in crv-out/ as frames/*.jpg, frames.json with per-frame timestamps, the transcript files, and a MANIFEST.txt. The MANIFEST is the part that matters for the workflow: it is the index you drop into the model alongside the images.

Two flags bend the selection rule. --adaptive picks frames against a rolling neighbourhood instead of a fixed threshold, which the README recommends for slow pans and gradual morphs where no single frame spikes. --text-anchors forces extra frames at subtitle-cue timestamps so each spoken segment gets a matching visual, capped at one forced frame per second, with scene detection left untouched. That second flag is the honest admission that scene detection alone is the wrong rule for lecture slides and talking heads.

Installing claude-real-video and running a first analysis

The package is on PyPI and requires Python 3.10 or newer. The README's install line pulls the whisper extra, which brings openai-whisper in for transcription. A second command installs the agent skill into Claude Code, Cursor, Codex, Copilot, Gemini CLI and other agentskills.io-compatible hosts, so the agent can invoke it from a pasted link.

bash
pip install "claude-real-video[whisper]"
npx skills add HUANGCHIHHUNGLeo/claude-real-video

If you only want the CLI, the README says the pip install is enough. Point crv at a URL and it writes the output folder described above.

bash
crv "https://www.youtube.com/watch?v=..."

Expect crv-out/ to appear with frames/, frames.json, transcript.txt, transcript.json and MANIFEST.txt. From there you paste the frames and MANIFEST.txt into Claude, ChatGPT or Gemini and ask your question. To inspect what the model will receive before spending tokens, add --viewer, which writes a local viewer.html containing the video, the keyframe grid and the transcript; the README says it needs no network and no extra installs. For a browser flow without a terminal, crv-web opens a local page in Traditional Chinese, Simplified Chinese or English where you paste a link or file path and click Analyze.

Where the design breaks down

The frame budget is real and it is spent on visual change. A 1920-wide screen recording downscaled to 640px will still be searched correctly, but the text that made the moment worth capturing is gone. The README's own remedy is --frame-width 1600, which is a manual trade of token cost against legibility rather than something the tool decides for you. --text-anchors has a harder constraint: it needs a sidecar .srt or .vtt file or an embedded subtitle track, and captions burned into the pixels cannot be detected at all. If your source is a screen recording with hardcoded subtitles and no subtitle file, that flag does nothing.

Speaker labels carry their own install. --speakers runs a local diarization model of about 45 MB that downloads once with no account or token, but it lives behind pip install "claude-real-video[speakers]" and requires sherpa-onnx. The README also draws a line between what the free tool does and what it does not: the free version lets a model see the video, while crv Pro, a one-time $29 add-on sold under the listing name llm-real-video Pro, adds how it was shot (cut rhythm, camera moves) plus a timestamped timeline of gestures, expressions, voice pitch shifts, emotion and sound events. If your question is about delivery rather than content, this tool is the wrong half of that split.

How it compares with sending video to a hosted model

The nearest alternative is Gemini's native video input, which the README describes as uploading the file to Google and sampling at a fixed interval. The difference is not accuracy in the abstract; it is where the sampling decision is made. Fixed-interval sampling is content-blind, so a two-second cut in a fast edit can be missed entirely, and the whole file leaves your machine. claude-real-video inverts both: selection follows scene change, and decoding, dedup, transcription and diarization all run locally, with only the frames and text you later paste going to a provider. The cost is that you assemble the prompt yourself and you carry ffmpeg, yt-dlp and a Whisper install. For a one-off question about a public video, the hosted route is less setup. For a private recording, a long call, or a workflow where you want to quote source timecodes, the local pipeline is the one that keeps the file in place.

Maintenance, licence and upgrade cost

The repository is not archived and the last push was on 2026-08-31, so it is current as of that date. The release history shows small, frequent fixes rather than long quiet periods: v0.10.3 changed a caption header so it no longer opens the transcript, v0.10.2 made URL runs use the source's own captions, v0.10.1 was window fixes. pyproject.toml pins yt-dlp to >=2026.8.19 with the [default,deno] extras, and the comment in the file explains why: YouTube changed player checks in August 2026 and 2026.07.x returns HTTP 403 on video data, while the deno extra bundles a JS runtime so URL downloads work without separate setup. That is the upgrade cost in practice. URL extraction depends on a dependency that has to track a moving target, so an old virtualenv is the most likely thing to break, and local-file runs do not share that exposure.

The licence is MIT, declared in both the LICENSE file and pyproject.toml. That covers the free package. crv Pro is a separate paid product sold on Capafy and through Lemon Squeezy, so the MIT terms do not extend to it, and the README does not describe its licence. Treat the free tool and the add-on as separate purchases and separate obligations.

Editorial conclusion

Adopt claude-real-video if you already have ffmpeg and Python 3.10 or newer and you want a local, MIT-licensed way to hand an LLM the frames that changed plus a transcript, with no upload of the source video. Do not adopt it if you need the model to reason about camera movement, tone of voice or sound events, since the README reserves that for the paid crv Pro add-on, or if your content is burned-in captions with no sidecar subtitle file, because --text-anchors cannot see them. Before relying on it, verify the yt-dlp version floor in an environment where URL downloads previously returned HTTP 403, and confirm ffmpeg is on PATH, since the README does not document a bundled fallback.

Frequently asked questions

Does Claude do video?

The README states that Claude will not take a video file at all, which is the gap claude-real-video fills: it converts a URL or local file into scene-change frames plus a transcript that you paste into the model.

Is the Claude video free?

The claude-real-video package is MIT licensed and free to install from PyPI. The README describes a separate paid add-on, crv Pro, at a one-time $29, which adds shot analysis and a timeline of gestures, expressions, voice pitch shifts, emotion and sound events.

Can Claude ingest video?

Not directly, according to the README. claude-real-video works around that by extracting keyframes and a transcript locally and handing the resulting folder to any LLM, including Claude.

How do you tell if a video is real?

The README does not cover authenticity or deepfake detection, so claude-real-video makes no claim about it. What it does provide is per-frame timestamps in frames.json, which let you point at a specific moment in the source.

Official sources

  1. HUANGCHIHHUNGLeo/claude-real-video on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/huangchihhungleo-claude-real-video.svg)](https://hysenlabs.com/projects/huangchihhungleo-claude-real-video)