Clipify: a Claude Code skill that cuts long videos into 9:16 clips with face-following pans
Claude Code skill: turn long videos into social-ready clips. Auto-find funny moments, cut, reframe to 9:16 with face-tracking, and burn opus-style captions.
At a glance
- What is it?
- Clipify is a local Claude Code skill for talking-head footage: Whisper finds the punchlines, ffmpeg motion energy decides who is speaking, and captions are burned in as ASS subtitles. It is free and offline, but it is a macOS-shaped script collection, not a finished product.
- Who is it for?
- Adopt Clipify if you already run Claude Code on macOS, your footage is a static-camera interview or podcast, and you are willing to install ffmpeg and Whisper yourself. Do not adopt it if you need a hosted editor, batch processing on a schedule, or Windows and Linux support without editing the ffmpeg flags.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 47 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 10, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Clipify solves, and who it is actually for
Turning a 45-minute interview into three vertical clips is mostly mechanical work. You watch or scrub for the good 30 seconds, cut it, reframe a 16:9 frame into 9:16 without cutting off a face, and then type the words back in as captions. Clipify targets that loop and nothing else. It is a Claude Code skill, which means it is not a standalone application with a window and a timeline. You install it into Claude Code's skills directory and invoke it as a slash command, and the conversation with the model is the interface: it proposes candidate clips, asks which one to cut, asks for the aspect ratio, and asks for a caption style.
The stated audience is narrow on purpose. The README says it is "built for talking-head dialogue (interviews, podcasts, two-person setups)" and that the author uses it to clip long-form videos for LinkedIn and TikTok. That framing matters more than the feature list. A skill that assumes a static camera and one or two visible faces will behave badly on handheld footage, gameplay, or a panel of five people. If your source material is a screen recording with a webcam bubble in the corner, the face-pan logic has nothing meaningful to track.
The economics are the other half of the pitch. The README claims "No cloud APIs. Runs entirely on your machine. No OpenCV." For anyone who has priced per-minute transcription and rendering, that is the difference between a tool you can run on every raw recording and one you ration.
How the pipeline works: Whisper, motion energy, ASS captions
Clipify is four Python scripts plus a prompt file. SKILL.md is what Claude Code reads; the scripts under scripts/ do the deterministic work. analyze.py produces a speaker timeline from two region-of-interest motion files, build_pan.py turns that timeline into an ffmpeg crop x-expression with hard cuts, build_ass.py generates ASS captions from Whisper JSON, and audio_align.py finds the offset of a sub-clip inside a longer source.
The interesting design decision is in the face-pan. There is no face detection model. The README explains the assumption directly: "Camera is static within a single clip." You eyeball each face's mouth and chin area as a rectangle on one sample frame, ffmpeg computes per-frame motion energy inside each rectangle using frame differencing, and whichever rectangle moves more at a given moment is treated as the speaker. That timeline becomes a list of x-coordinates, and the crop jumps between them as hard cuts rather than gliding.
This is cheap and it is honest about being cheap. The README calls the cost "a few seconds of ffmpeg per clip" and says it "works surprisingly well." The failure mode is equally clear: a listener who nods, a hand gesture crossing the wrong rectangle, or a cut to a wide shot all break the assumption that motion equals speech. There is no smoothing pass described, and no confidence score, so a wrong pan is something you catch by watching the output.
Captions are burned in rather than exported as a sidecar. build_ass.py writes ASS subtitle files in three named styles (opus, karaoke, minimal), and the README says you can paste a reference image to match a style. Burning means the words are in the pixels, so you cannot fix a typo after render without re-rendering, and you cannot hand the clip to an editor who wants a separate caption track.
Installing Clipify and running a first clip
Clipify installs by cloning the repository into the Claude Code skills directory. The README gives this as the whole install step, followed by a restart of Claude Code, after which /clipify becomes available as a slash command.
git clone https://github.com/louisedesadeleer/clipify.git ~/.claude/skills/clipifyBefore that, the requirements section lists four things: macOS, Claude Code, ffmpeg with libx264, and Whisper plus numpy for Python 3. The README gives the install commands for the dependencies.
brew install ffmpeg
pip install openai-whisper
pip install numpyThe macOS requirement is not cosmetic. The README states the skill "uses VideoToolbox for hardware-accelerated decode" and that it "works on Linux/Windows if you remove `-hwaccel videotoolbox` flags." So a Linux user is not blocked, but they are editing script flags rather than running an install path the author supports.
Once installed, the usage loop is conversational. You type the slash command, paste a video path when asked, and the skill transcribes and proposes three to five candidates with timestamps and titles. It then asks which to cut, which aspect ratio (9:16, 16:9, or 1:1), and, if you chose 9:16 from 16:9 with two faces, whether you want a pan or a split-screen. Finally it asks for a subtitle style. Rendered clips land in a clipify_out/ directory next to the source video, which is worth knowing before you run it on a file in a synced folder.
/clipifyThe README claims roughly 20 seconds of work for a 20-second clip on Apple Silicon. That is the author's figure, not a benchmark you can plan capacity around, and it says nothing about transcription time for the full source video, which is usually the longer wait.
Where Clipify breaks, and when it is the wrong tool
The static-camera assumption is the load-bearing one. Every part of the pan logic depends on it, and the README states it as a precondition rather than a limitation to work around. Footage from a moving camera, a gimbal walk-and-talk, or a multi-camera edit where the cut points change the framing will produce a speaker timeline that does not correspond to anything.
The second constraint is language and content. Nothing in the README describes how Whisper is configured, which model size is used, or whether non-English audio is handled. The candidate-finding step scans the transcript for "punchlines, reversals, awkward pauses, and audio peaks," which is a reasonable heuristic for English comedy and interview banter and a much weaker one for technical explanation, where the valuable segment is often the calm middle of an answer rather than a peak.
The third is that this is a skill, not a service. There is no queue, no watch folder, no scheduled batch job. Each clip is a conversation. If you need to process forty recordings overnight, you are writing that orchestration yourself on top of scripts the README does not document as a public API. The README also does not document rollback, version pinning, or how the skill behaves when Whisper fails on a corrupt audio track, so error handling is something you discover rather than read.
Finally, the repository is small: a prompt file, four scripts, an assets directory, and a licence. The last push was on 2026-08-24 and the only release is v0.1.0 from 2026-05-05. That is a working personal tool, not a project with a support surface.
Clipify versus the hosted clipping tools people compare it to
The obvious alternative is a hosted auto-clipper, the kind of product that takes a YouTube link and returns captioned vertical clips through a browser. Those tools remove the install entirely: no ffmpeg, no Whisper, no Python, no macOS. They also typically run their own speech models, which means they can accept a URL rather than a local file and can process on a schedule. The trade is that your footage leaves your machine, you are subject to per-minute pricing or a subscription, and the caption styling is whatever the product offers rather than a style you can match to a reference image.
A second alternative is doing it by hand in a general editor with a transcription feature. That keeps full control of the cut and the framing, and it is the right answer when a clip needs real editorial judgement, like trimming a sentence mid-answer or matching a b-roll insert. It costs the thing Clipify is trying to save, which is the time spent scrubbing and reframing.
The distinction that matters is not quality, since the README makes no comparative quality claim. It is where the work happens. Clipify runs locally, produces burned-in captions, and hands the decision points back to you through a chat prompt. A hosted clipper runs remotely and hands you a finished file. If your constraint is data residency or per-minute cost, Clipify's approach is the one that fits. If your constraint is that you do not want to install ffmpeg and a speech model, it is the one that does not.
Licence, maintenance and what an upgrade costs you
Clipify is MIT licensed, with a LICENSE file at the repository root and a one-line licence section in the README. MIT permits commercial use, modification, and redistribution provided the copyright notice and permission notice are kept. That is the standard reading, not legal advice; if you are embedding the scripts in a product, read the LICENSE file itself rather than this summary.
The practical licence question is dependency licensing, not Clipify's. Whisper and ffmpeg are separate projects with their own terms, and the README links to Whisper's repository rather than vendoring it. ffmpeg builds vary in which codecs and libraries they include, which affects what you can legally ship in a binary but not what you can run locally.
Upgrade cost is low in the ordinary case, because there is no package to update. You cloned the repository into ~/.claude/skills/clipify, so updating means pulling the branch again. The risk is that SKILL.md is the prompt the model reads, and a change there can alter behaviour without any version bump you would notice. With one release, v0.1.0, there is no changelog history to diff against, so the honest position is that you should read SKILL.md after any pull rather than assume the interface is stable.
Editorial conclusion
Adopt Clipify if you already run Claude Code on macOS, your footage is a static-camera interview or podcast, and you are willing to install ffmpeg and Whisper yourself. Do not adopt it if you need a hosted editor, batch processing on a schedule, or Windows and Linux support without editing the ffmpeg flags. Before trusting it on a paid job, verify two things: that Whisper's transcript matches your speakers' names and accents well enough for the candidate list to be useful, and that the motion-energy pan picks the right face on your own lighting, because the README describes the method but documents no accuracy figure.
Frequently asked questions
How do I use Clipify?
Install it by cloning the repository into ~/.claude/skills/clipify, restart Claude Code, then type /clipify and paste a video file path when asked. The skill transcribes the video, proposes three to five candidate clips, and asks you to choose the cut, the aspect ratio, and the caption style.
What is Clipify?
Clipify is a Claude Code skill that turns long videos into social-ready clips. According to the README, it transcribes with Whisper, finds clip-worthy segments, reframes 16:9 to 9:16 with face-following pans, and burns in word-by-word captions.
What does Clipify do?
It proposes three to five candidate moments from a transcript, cuts your chosen one, reframes it to 9:16, 16:9 or 1:1, and adds subtitles in an opus, karaoke or minimal style. Final clips are written to a clipify_out/ directory next to the source video.
Is Clipify free?
The repository is MIT licensed and the README states it uses no cloud APIs and runs entirely on your machine, so there is no per-minute or subscription charge described. You supply your own ffmpeg, Whisper and Python environment.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/louisedesadeleer-clipify)