Open-source project
HITsz-TMG/VideoClaw avatar
HITsz-TMG/VideoClaw

VideoClaw: an AI director pipeline from a single idea to a finished film

🚀 AI 全自动化视频生成员工 | Your First AIGC Coworker. Chat an Idea. Get a Film. 🦞

1,826 stars274 forksPythonMIT

At a glance

What is it?
VideoClaw is a Python multi-agent system that turns a one-line idea into script, character art, storyboards, reference images, video clips and a final cut. It is for people who want editable intermediate assets rather than a single black-box render.
Who is it for?
VideoClaw fits teams that already have API keys for LLM, image and video models and want a staged, inspectable pipeline they can stop and correct between stages. It does not fit anyone expecting a single prompt-to-MP4 box, or anyone unwilling to pay per-call model costs across six stages.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 36 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap VideoClaw targets: black-box text-to-video versus a production line

Most text-to-video tools return one file. You type a prompt, you wait, you get a clip, and if the protagonist's face changes between shots you start over. VideoClaw's README draws that contrast explicitly: it describes itself as not a single-point text-to-video tool but a production line covering script planning, character and scene design, storyboard planning, reference image generation, video generation and post-production editing. The intended user is someone producing narrative content, short drama series in particular, where continuity across shots matters more than the quality of any single frame.

The repository's own demo is a series, not a clip: an eight-episode short drama titled around a programmer who is laid off and buys back his former company, generated as six episodes plus two written through the continuation feature. That shape tells you who the project is for. If your output is thirty seconds of abstract motion, the staging overhead is wasted. If your output is episode seven of a serial, the staging is the product.

How the six-stage pipeline passes assets forward

The mechanism is sequential and asset-driven. Stage one takes a creative title and a project synopsis and produces a structured multi-scene script with narration and dialogue. Stage two extracts character and scene features from that script and generates reference concept art in a consistent style. Stage three decomposes each scene into consecutive visual shots with camera angle, action description and reference content. Stage four renders a reference base image per shot, controlling lighting and composition. Stage five calls video generation models to turn storyboard images into moving segments. Stage six aggregates the clips and exports a finished video.

The design claim in the README is that each stage decides the next, and that every key node is visualizable, editable and resumable. That is the real architectural difference: the intermediate assets are first-class outputs, not scratch data. Continuation works at the script layer, which is why the project can extend a series by writing two more episodes rather than regenerating eight.

Stage five is configurable in a way worth noting. Since the 2026/6/11 update, the main flow lets you choose between first-frame-to-video, first-and-last-frame-to-video, and reference-image-to-video, with a separate model configuration for each mode. That is a meaningful escape hatch: when a shot's motion looks wrong, you change the generation mode for that shot instead of changing the whole project's model.

Installing VideoClaw and running a first project

The README lists Python 3.9 or newer as the requirement and the repository root contains a video-claw directory alongside a FilmAgent directory. The README's news entry for 2026/5/8 states that WebUI configuration of APIs and default models was added together with one-click installation. The README does not publish a pip package name or a pinned release version, so the practical route is cloning the repository and following the install path inside it.

Start by cloning and entering the project directory:

bash
git clone https://github.com/HITsz-TMG/VideoClaw.git
cd VideoClaw

Because no package name is documented, inspect what the repository actually ships before running anything:

bash
ls video-claw

What you see there determines the install command. The README does not document a requirements.txt or an install script by name, so do not assume one exists.

The documented workflow after installation runs through the WebUI. The first configuration step is the global settings page, where you enter API keys and set default models. The README's homepage description lists exactly this: system overview, historical projects, new project creation, and global configuration for API keys and default models. Once keys are set, the first real use is stage one: enter a creative title and a project synopsis, and the system generates a structured multi-scene script with narration and dialogue. You then review that script before moving to character and scene design.

If you use OpenClaw, the README's headline integration is conversational: talk to OpenClaw directly with an instruction to generate a video for X, and the integration handles it. The README labels this as the third configuration method and links a ClawHub listing at clawhub.ai/hit-cxf/video-claw. The README does not document what the OpenClaw skill passes to the pipeline or how failures surface in that chat flow.

Where VideoClaw breaks down or is the wrong tool

Cost and latency scale with stages, not with output length. Six stages mean at least script generation, image generation for characters and scenes, image generation per shot, and video generation per shot. Every one of those is a paid API call to a hosted model. A ten-shot episode is a large number of billed calls before you see a single second of motion, and a rejected shot means paying for that shot again. The README does not publish cost estimates or token budgets, so you are budgeting blind until you run one project.

The pipeline is also only as good as its weakest model. If your video model produces flicker or identity drift, the earlier stages cannot fix it; they only give you cleaner inputs. The README's framing of a collaborator you can intervene with is accurate about control, not about quality.

There is a scope limit too. The README describes the project as aimed at creative video production and short drama specifically. Documentary footage, live-action capture, screen recordings, data visualisation and anything requiring real-world camera input fall outside the described pipeline entirely. If your source material is not generated, the six stages have nothing to work on.

Finally, the repository shows no releases. The README carries a version badge reading 1.0.0, but the release list is empty, which means upgrades arrive as commits on main rather than as versioned artifacts you can pin.

VideoClaw and FilmAgent: two directories, two approaches

The repository root contains a FilmAgent directory next to video-claw, and FilmAgent also appears in the search phrases people use around this project. FilmAgent is a research-oriented multi-agent filmmaking project, and its presence here suggests VideoClaw grew out of or alongside that line of work. The difference in approach is visible in the repository structure: FilmAgent is the research codebase, while video-claw is the application layer with a WebUI, global API configuration, project history and one-click export.

That distinction matters when you choose. If you want to read and modify the agent orchestration itself, the FilmAgent directory is where the underlying machinery lives. If you want to produce a series and stop between stages to fix a character's face, video-claw is the layer built for that, and the README's stage-by-stage WebUI screenshots describe it as the product surface.

Licence, maintenance and what an upgrade costs you

VideoClaw is MIT licensed, and the repository links a LICENSE file from the README badge. MIT permits commercial use and modification with attribution and no warranty, but this is a statement about the project's own code, not about the models it calls. Your generated output is governed by the terms of whichever LLM, image and video APIs you configure. The README names Wan and Kling as video generation examples; their licences and usage terms are separate from VideoClaw's and are not covered by it. That is not legal advice; check the terms of each provider you plug in.

Maintenance: the last push to the repository was on 2026-08-26. The README's news timeline runs from 2026/3/27 through 2026/6/11, so the documented feature set has been stable for a while, and no releases have been cut. Upgrading therefore means pulling main and re-reading the README news entries to see what changed, since there is no changelog file in the repository root and no versioned release to diff against. If you need a pinned dependency you can audit, this repository does not currently give you one.

Editorial conclusion

VideoClaw fits teams that already have API keys for LLM, image and video models and want a staged, inspectable pipeline they can stop and correct between stages. It does not fit anyone expecting a single prompt-to-MP4 box, or anyone unwilling to pay per-call model costs across six stages. Before adopting, verify which video model endpoints your keys actually reach, because the README names Wan and Kling as examples without pinning versions, and check the video-claw directory for the install path that matches your setup. The repository had its last push on 2026-08-26 and carries no releases, so treat the main branch as the only distribution channel.

Frequently asked questions

How can I create automated videos with VideoClaw?

Install the project, open the WebUI, and set your API keys and default models on the global configuration page. Then create a project by entering a creative title and a synopsis; the system generates a structured multi-scene script, and you continue through character and scene design, storyboard planning, reference image generation, video generation and post-production export. The README also documents driving it conversationally through OpenClaw.

What Python version does VideoClaw need?

The README badge states Python 3.9 or newer. The README does not document a pip package name, so installation goes through the repository's own instructions rather than a published package.

Does VideoClaw work with OpenClaw?

Yes. The README marks the project as OpenClaw compatible and describes talking to OpenClaw directly with an instruction to generate a video, which it calls the third configuration method. A ClawHub listing is linked at clawhub.ai/hit-cxf/video-claw.

Which video generation models does VideoClaw support?

The README names Wan and Kling as examples of high-performance video generation models it calls. Since the 2026/6/11 update, the main flow lets you switch between first-frame-to-video, first-and-last-frame-to-video and reference-image-to-video, with a separate model configuration for each mode.

What licence is VideoClaw released under?

The repository is MIT licensed and links a LICENSE file from the README. That covers the project's own code; the terms of the model APIs you configure are separate.

Official sources

  1. HITsz-TMG/VideoClaw on GitHub
  2. Issues
  3. License: MIT
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/hitsz-tmg-videoclaw.svg)](https://hysenlabs.com/projects/hitsz-tmg-videoclaw)