Library / SDK
huangserva/ComfyUI_MiniMaxH3_Director avatar
huangserva/ComfyUI_MiniMaxH3_Director

Cloning the MiniMax H3 Director repository installs nothing, and three of its five workflows lack their weights

ComfyUI MiniMax H3 Director workflow

1,140 stars123 forksUnknownApache-2.0

At a glance

What is it?
Five ComfyUI workflow files kept as a download copy of someone else's plugin, with no install step of its own, no documented weight download, and an honest limitations section that says the character swap mode has no identity lock. The readme is in Chinese.
Who is it for?
Treat this as a set of five parameter files rather than as software, because that is what it is: a convenient copy of workflows that the upstream node package generates, kept so they can be downloaded, diffed and tested. Two things to establish before anything renders.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 60 days ago.
What is it written in?
GitHub does not report a main language for this repository.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Cloning this repository installs nothing

The repository holds three entries: a licence, a readme, and a directory of example workflows. There is no code, no install script, and no configuration.

The readme says what it is in plain terms. It is a copy of a workflow set that came from another author's plugin, kept here so that the files are convenient to download, reproduce and test.

Then the installation instructions go and clone the other project:

bash
cd ComfyUI/custom_nodes
git clone https://github.com/AIMixer/ComfyUI_MiniMaxH3_Director.git
python -m pip install -r ComfyUI_MiniMaxH3_Director/requirements.txt

So the node package you need is installed from upstream, and the requirement file you install from is also upstream. Nothing in these three steps touches the repository you are reading. After restarting the editor, you drag the JSON files from this repository's example directory onto the page.

Two small things about that block. The version floor is stated as a specific editor release or newer, and the requirement install is unqualified by a virtual environment, so it goes into whichever Python the editor is running under. For an application that is usually its own environment and usually fine, but it is the kind of line that bites when the editor runs from a system Python.

The value of the copy is therefore narrow and real: you get five known-good parameter sets without opening the upstream repository, and you can diff them against a newer upstream.

Three of the five workflows need a weight the author did not have

The workflow table has five rows, and the third column is the one that decides what you can run.

Two rows point at one diffusion weight. The text-to-video workflow and the first-and-last-frame workflow, which doubles as image-to-video when you supply only a leading frame, both use what the table calls the fl2va model.

Three rows point at a different one. The reference-driven generation workflow, the timeline editing workflow and the source-video-plus-reference workflow all use what the table calls ref2va.

Then the usage notes say which of those the author actually had. The verified card only has the ref2va weight, so the three reference workflows can run directly, while the text, image and frame workflows need the other weight to be obtained first.

So three of five workflows are reproducible from the documented environment and two are not. That is stated plainly rather than glossed, which is more than most workflow repositories do.

The verified environment itself is worth quoting exactly, because it bounds the claim: an RTX 4090 with 48 GB of memory, the editor at one specific release, PyTorch at a specific version with a specific CUDA version, the reference weights in an eight-bit integer build, and the director node registering successfully.

The standard configuration of that card carries 24 GB, so the stated 48 GB is either a different variant or a mistake. Either way it does not match what most readers have on the desk, and the shared text encoder alone is a 32-billion-parameter model in a four-bit build.

No step tells you where the weights come from

The model requirements section names five files and their expected locations, split into the two diffusion weights and three shared components: a text encoder, a video autoencoder in half precision, and an audio autoencoder in single precision.

What it does not contain is a download step. There is no command, no script, and no model manager entry. The sources section at the bottom points at a Hugging Face organisation page for the weights and at the editor's own documentation, and that is where the fetching is left.

So the path to a first render is: install the upstream node package, find and place five files yourself in the directories the workflow expects, set one dropdown in the loader, and restart. For a repository whose entire purpose is reproducibility, that is the largest gap in the chain.

Two details from the file names are worth passing on, because they tell you what you are downloading. Both diffusion weights are pruned and eight-bit integer builds with the same suffix, and the text encoder is a 32-billion-parameter model in a low-bit format. This is a large download and a memory-hungry one, which is why the verified card is a big one.

The audio half is separate and easy to miss: a video decoder in half precision and an audio decoder in single precision, and the text-to-video workflow is described as generating audio alongside the picture.

The character swap mode has no identity lock, by the author's own account

One of the five workflows takes a source video plus a reference image, and the readme recommends it for testing a character swap. That is the one with the most moving parts, so it is worth reading alongside what the readme says the tool does not guarantee.

The workflow is five steps. Import the file. Upload a source video and either split it by shot or use an automatic segmentation. Upload a reference image of the person, plus a reference audio clip if you want the voice constrained. Write a prompt per segment, where the source clip is bound to a numbered placeholder and reference assets are bound to their own numbered placeholders. Then render a short clip first, inspect the face, the clothing, the motion and the shot boundaries, and only then extend.

Two design details help. The director allows segments to be selected and cached individually, so changing one segment does not mean re-running the timeline. And audio can be generated by the model, kept from the original, or muted, per segment.

The placeholder scheme is the fiddly part: because references are bound by index, the prompt text has to stay in step with the order things were uploaded, and inserting a new reference renumbers everything after it.

Now the limitations, which are the author's own and are stated without hedging. Identity is constrained jointly by the model and the reference material, and the director has no hard identity lock. The handoff between segments passes the previous segment's last frame to the next segment's first frame, and that cannot substitute for a character consistency check.

Read together: this is a tool for looking at whether a shot works, not a tool for holding one person's face consistent across a finished sequence.

The A/B paragraph is the most useful thing in the file

Five of the usage notes are about how to test rather than how to run, and they are worth isolating.

The first two are the weight situation already described: what the verified card can run, and what it cannot.

The next two are the identity caveats above.

The last one is a testing discipline: when comparing two runs, fix the source asset, the prompt, the seed, the resolution, the frame count and the step count, and then change one thing.

Six variables pinned before a comparison is more than most people hold constant, and each of them is one that silently changes output on its own. A seed and a step count are the obvious two; frame count and resolution change what the motion model does rather than just resampling it.

The optional acceleration section follows the same spirit. An attention patch can be installed and placed between the model loader and the director node's model input, and the instruction is to confirm the output quality is identical first and record the speed change afterwards. In order.

That ordering is the whole point. Verify the substitution changed nothing you can see before you record that it changed something you care about, which is the opposite of the usual advice to benchmark first.

Apache-2.0 covers the workflow files, not the weights

The licence section draws a line that is easy to miss.

It states that the workflows and the upstream plugin are released under Apache-2.0, pointing at a licence file in the repository. It then states separately that the model weights follow their own licence terms.

So what is permissively licensed here is five JSON parameter files. The thing that makes them work, a set of model weights several times larger than the repository, carries terms that are not stated in this repository at all, and the sources section points at a model hosting page rather than reproducing the terms.

That is the normal arrangement for a workflow repository and it is the right thing to say, since a reader who assumed Apache-2.0 covered the output would be wrong.

The rest of the credit chain is also explicit: the upstream plugin with its repository, the weights with their hosting page, and the editor's own documentation for the workflow. A three-file repository with a four-line tree is not much, but the attribution is complete, which is the part that usually goes missing.

The readme itself is in Chinese and there is no translated version in the tree, so the limitations notes above are read from a single-language document.

Editorial conclusion

Treat this as a set of five parameter files rather than as software, because that is what it is: a convenient copy of workflows that the upstream node package generates, kept so they can be downloaded, diffed and tested. Two things to establish before anything renders. Which weight you actually have decides everything, since three of the five workflows need a diffusion file the author did not have, and the readme names the expected filenames without documenting where to fetch them, so budget for finding and placing several very large files by hand. And if you are drawn to the source-video plus reference variant, read the author's own limitations first: identity is jointly constrained rather than locked, and the segment handoff is a last-frame-to-first-frame join that cannot substitute for a consistency check.

Frequently asked questions

What is the ComfyUI MiniMax H3 Director workflow repository?

A copy of five workflow files for a director node from another author's plugin, kept for convenient download, reproduction and testing. It contains no code and no installer: the documented setup clones the upstream plugin into the editor's custom nodes directory and installs that project's requirements, then the JSON files from this repository are dragged into the page.

What hardware and software were the MiniMax H3 workflows verified on?

An RTX 4090 with 48 GB of memory, the editor at release 0.30.0, PyTorch 2.11.0 with CUDA 12.8, and the reference-to-video-and-audio weights in an eight-bit integer build, with the director node registering successfully. The readme notes that only the reference weight path was verified on that machine.

Which model weights does the MiniMax H3 Director workflow need?

Two diffusion weights, one for the text, image and first-and-last-frame workflows and one for the reference workflows, plus a shared text encoder, a video autoencoder in half precision and an audio autoencoder in single precision. The readme names the expected file names and locations but documents no download step, and states that the weights follow their own licence terms.

Does the MiniMax H3 Director keep a character's identity across a video?

Not as a hard guarantee. The readme states that identity is constrained jointly by the model and the reference material, that the director has no hard identity lock, and that the last-frame-to-first-frame handoff between segments cannot replace a character consistency check. It recommends rendering a short clip and checking face, clothing, motion and shot boundaries before extending.

How do I compare two MiniMax H3 Director runs?

Pin the source asset, the prompt, the seed, the resolution, the frame count and the step count, then change one thing. An optional attention patch can be placed between the model loader and the director node, and the readme's instruction is to confirm output quality is identical before recording any speed change.

Official sources

  1. huangserva/ComfyUI_MiniMaxH3_Director on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/huangserva-comfyui-minimaxh3-director.svg)](https://hysenlabs.com/projects/huangserva-comfyui-minimaxh3-director)