HKUDS/ViMax: an agentic pipeline that turns a concept into a finished video
"ViMax: Agentic Video Generation (Director, Screenwriter, Producer, and Video Generator All-in-One)"
At a glance
- What is it?
- ViMax chains screenwriting, storyboarding, image generation and video generation into one Python workflow, with a TUI and a Web UI on top. It is a good fit for people who want a scripted pipeline they can inspect; it is the wrong tool if you want a one-line prompt and a clip.
- Who is it for?
- Adopt ViMax if you already have a provider key, are comfortable with Python 3.12 and uv, and want the intermediate artifacts (story, characters, script, storyboard, shots) as files you can inspect and revise. Do not adopt it if you need a single prompt-to-clip tool, or if you cannot supply your own image and video generation credentials.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 10 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap ViMax targets: short clips with no story behind them
The README opens with three complaints about current video generation tools: output limited to short clips, characters and scenes that change unpredictably across frames, and a visual-only focus that leaves out script, audio and narrative structure. Those three complaints map directly onto the four roles ViMax claims to fill at once: director, screenwriter, producer and video generator. The pitch is that you supply a concept and the system handles scriptwriting, storyboarding, character creation and final assembly end to end.
The intended user is not someone who wants a five-second clip from a text prompt. It is someone producing a multi-scene piece who cares about whether the same character looks the same in shot three as in shot one, and who wants the script to exist as a reviewable artifact before any pixels are generated. The repository structure reflects that: there are separate entry points for `main_idea2video.py`, `main_script2video.py` and `main_agent.py`, which suggests the workflows are distinct enough to be launched independently rather than being one monolithic command.
Four workflows, one pipeline: Idea2Video, Script2Video, Novel2Video, AutoCameo
The feature list names four input shapes. Idea2Video takes a short concept and expands it into structured stories, characters, scripts, storyboards, shots and a finished video. Script2Video starts from an explicit screenplay and preserves its creative intent while producing multi-scene, multi-shot video. Novel2Video adapts long-form fiction into episodic visual narratives, and the README specifically mentions narrative compression, character tracking and scene planning as the mechanisms for that. AutoCameo places a person or pet from a reference photo into generated stories while holding their appearance consistent.
The interesting design decision is that all four converge on the same downstream stages. The README describes a consistent production pipeline that coordinates references, first frames, camera continuity and final assembly. That means the hard part (keeping a character stable across shots) is solved once, at the reference and first-frame level, rather than being re-solved per workflow. It also means a weakness in that layer propagates to every workflow at once. The repository layout supports this reading: there is a single `pipelines/` directory and a single `agents/` directory, not one per input type.
How the agent loop and the Web UI sit on top of the pipeline
Version 1.2.0, released on 2026-07-20, added a Web UI described as supporting named projects, Agent Loop conversations, artifact and storyboard previews, render checkpoints, file uploads, provider settings and dark mode. The Agent Loop itself arrived earlier, on 2026-06-08, integrated with a TUI for interactive planning, revision, rendering control, session reuse and context compaction.
Read the two together and the architecture becomes clearer. The pipeline stages produce artifacts (stories, characters, scripts, storyboards, shots). The agent loop is the layer where a human argues with those artifacts before committing to a render. Render checkpoints and persistent render status, both mentioned in release notes, exist because a full render is long enough that losing it to an interruption is expensive. The 2026-06-28 note lists stronger LLM retries, landscape image guards and Script2Video resume fixes, which tells you where the failure modes actually were: flaky model calls, aspect-ratio mismatches, and interrupted renders.
Generation is parallelized where shots are compatible, according to the feature list. That is a scheduling claim, not a benchmark, and the README does not quantify the speedup.
Installing ViMax and running a first Idea2Video job
The project targets Python 3.12 and ships a `uv.lock`, so uv is the expected installer. The pyproject file defines a `pytorch-cu128` index that is applied on Linux and Windows only, via a platform marker, which means macOS users resolve torch from the default index instead. Clone the repository and sync:
git clone https://github.com/HKUDS/ViMax.git
cd ViMax
uv syncThe dependency list is worth reading before you start, because it tells you what the machine has to do. It includes `moviepy` for assembly, `opencv-python` and `scenedetect[opencv]` for shot handling, `faiss-cpu` for retrieval, and `langchain-openai` plus `google-genai` for model access. There is no local diffusion model in the dependency list: image and video generation are delegated to external providers. That is why the Web UI has a provider settings panel.
Once synced, the entry points are scripts at the repository root. The README's quick start points to the Idea2Video path for a concept-to-video run:
uv run main_idea2video.pyThe README does not reproduce the flags or config keys for that script in the text available here. Before running it, inspect `configs/` and the script's own argument parser, and set provider credentials for whichever image and video backends you intend to use. The news entries name several: OpenRouter GPT Image 2 for images, Seedance 2.0 Fast and Google Omni for video, and MiniMax as a chat model provider. You should expect the first run to be a configuration exercise rather than a one-command demo.
Where ViMax breaks down, and when it is simply the wrong tool
The most concrete limitation is that ViMax does not generate video itself. Nothing in the dependency list is a video generation model. Every frame that reaches the final assembly came from a provider you configured and paid for, which means quality, latency and cost are set by those providers, not by ViMax. The framework's contribution is orchestration and consistency, and if the underlying generator cannot hold a character's face, no amount of storyboard coordination fixes it.
Second, this is a long-running, multi-stage, network-dependent process. The release notes themselves document LLM retries and resume fixes, which is an admission that calls fail and renders get interrupted. If your workflow cannot tolerate a job that takes many model calls and may need to be resumed, this is the wrong shape of tool.
Third, if you want a prompt box and a clip, ViMax is overkill. It will ask you to think about characters, scripts and shots. That is the point, but it is also friction. The README's own framing, that it orchestrates scriptwriting and storyboarding before generation, is a description of a slower path than a direct text-to-video call.
Finally, the documentation available here does not cover rollback of a partially completed project, nor does it state resource requirements for the assembly stage. Treat those as unknowns to check in `tests/` and the pipeline code rather than assumptions.
How ViMax differs from a direct text-to-video API
The obvious alternative is calling a video generation API directly, for example the Seedance or Google Omni endpoints that ViMax itself can use as backends. The difference is not the generator; it is everything before the generator. A direct API call takes a prompt and returns a clip. ViMax inserts four stages between your idea and that call: story and character definition, script, storyboard, and shot breakdown, each of which is an artifact you can read and edit.
That matters in exactly one situation: when you need the same subject to appear across multiple shots and scenes. A single API call has no memory of the previous shot. ViMax's consistent production pipeline, per the README, coordinates references and first frames so that continuity is a property of the pipeline rather than of the prompt. If your output is one shot, the extra stages buy you nothing and cost you time and tokens.
A second comparison point is scope of input. Tools built around a prompt box do not accept a full screenplay or a novel. ViMax has dedicated paths for both, and the Novel2Video path adds narrative compression and character tracking specifically because long fiction does not fit into a single context window. If your source material is a finished script, Script2Video is the path that preserves it rather than regenerating it.
Licence, maintenance and what an upgrade costs you
ViMax is MIT licensed, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are included. That is the least restrictive common option and it does not impose copyleft obligations on your own code. It says nothing about the terms of the image and video providers you connect, and those are separate agreements you enter into directly. This is not legal advice; read the LICENSE file in the repository for the exact terms.
The repository is not archived, and the last push was on 2026-09-20, one day before the date used for this assessment. The release history shows a steady cadence through 2026: MiniMax chat provider in March, Google Omni in June, Novel2Video in June, the Agent Loop and TUI in June, OpenRouter GPT Image 2 and Seedance 2.0 Fast in July, and the v1.2.0 Web Workspace on 2026-07-20.
The upgrade cost is real, though. The package is named `autolongvideogeneration` in pyproject while the repository and entry points are ViMax, and the version there is 1.2.0. Because image and video providers are pluggable, provider additions land as code changes you inherit on upgrade. The 2026-06-28 note about Script2Video resume fixes suggests that sessions created before a fix may not resume cleanly after it. Pin a version for anything you need to reproduce.
Editorial conclusion
Adopt ViMax if you already have a provider key, are comfortable with Python 3.12 and uv, and want the intermediate artifacts (story, characters, script, storyboard, shots) as files you can inspect and revise. Do not adopt it if you need a single prompt-to-clip tool, or if you cannot supply your own image and video generation credentials. Before committing, verify three things in your own environment: that `uv sync` resolves the PyTorch CUDA index on your platform, that your chosen providers are among those the configs directory actually wires up, and that the Script2Video resume path behaves as described in the 2026-06-28 release note when a render is interrupted.
Frequently asked questions
What is HKUDS/ViMax used for?
It is an agentic video creation framework that connects narrative planning, visual consistency, image generation, video generation and final assembly in one workflow. It offers Idea2Video, Script2Video, Novel2Video and AutoCameo paths.
How do you install and use ViMax?
The project targets Python 3.12 and ships a uv.lock, so the expected route is cloning the repository and running uv sync, then launching one of the root entry points such as main_idea2video.py with uv run. You also need to configure image and video generation providers, since the dependency list contains no local generation model.
Does ViMax generate the video itself?
No. The dependency list contains no video generation model, and the release notes describe support for external providers such as Seedance 2.0 Fast and Google Omni. ViMax orchestrates scriptwriting, storyboarding, character creation and assembly around those providers.
What does the ViMax Web UI add over the command line?
Version 1.2.0 introduced a Web UI with named projects, Agent Loop conversations, artifact and storyboard previews, render checkpoints, file uploads, provider settings and dark mode. The Agent Loop and TUI workflow had been integrated earlier, on 2026-06-08.
What licence is ViMax released under?
The repository is MIT licensed according to the README badge and the LICENSE file. That permits commercial use and modification provided the copyright and permission notices are retained.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/hkuds-vimax)