# The MiniMax H3 Director declares an NVIDIA CUDA environment and keeps its NVIDIA package out on purpose, and one bundled workflow's filename disagrees with its own task type

> AIMixer/ComfyUI_MiniMaxH3_Director is a single ComfyUI node that turns long multi-segment audio-video generation into a timeline, where one node carries four required inputs, five optional ports and seven outputs, the requirements file installs three packages the page calls optional while leaving the fourth commented out with a note about breaking non-AMD machines, and the model weights come from outside the repository.

**AIMixer/ComfyUI_MiniMaxH3_Director** — Multi-segment MiniMax H3 Director for official ComfyUI MiniMax-H3

- Repository: https://github.com/AIMixer/ComfyUI_MiniMaxH3_Director
- Stars: 2,154 · Forks: 256
- Language: JavaScript
- License: Apache-2.0
- Published: 2026-09-16 · Updated: 2026-09-16 · Language: en
- Canonical page: https://hysenlabs.com/projects/aimixer-comfyui-minimaxh3-director

## The manifest asks for CUDA and the requirements file refuses it

The two files that describe this package disagree about hardware, in writing, on purpose.

The manifest carries a classifier declaring an NVIDIA CUDA GPU environment. The requirements file carries the opposite decision, as a comment explaining why:

```text
# Optional: Refine upscale_method=nvidia_rtx_vsr (NVIDIA RTX GPU only).
# Do not add as a hard dep — AMD / cloud / no-VSR installs would break.
# pip install nvidia-vfx --extra-index-url https://pypi.nvidia.com
```

That reasoning is sound. The RTX video super-resolution path is one of three upscale methods, alongside pixel interpolation and upscaling the H3 latent, and hard-installing an NVIDIA package would break the machine it is meant to protect. But the classifier in the same package still tells a registry or an index that this is a CUDA project.

The node itself needs no NVIDIA GPU, since the NVIDIA package is not a requirement. So the classifier overstates the base requirement and the requirements file understates the one optional path that does need one.

There is a second, quieter disagreement in the same place. The page describes three packages as optional, but the file installs all three unconditionally: `scenedetect>=0.6.4,<0.8`, `opencv-python-headless>=4.8` and `imageio-ffmpeg>=0.4`. Only the NVIDIA one is commented out.

## Four required inputs, five optional ports and seven outputs

The whole thing is one node, and its socket list is the clearest statement of scope on the page.

Required inputs are `model`, `video_vae`, `audio_vae` and `clip`. Optional inputs are `i2v_groups`, `r2v_groups`, `semantic_bridge`, `selflift` and `refine`, each of which takes an external companion node.

Outputs are seven, in order: `images`, `audio`, `fps`, `frame_count`, `source_images`, `report` and `images_pre_refine`. Two of those seven exist only under a condition. `images_pre_refine` is the first pass before upscaling, so it is populated only when a Refine node is connected. And `source_images`, which is meant to show the source frames used for a video-to-video edit, fills only its own socket, does not change the main `images`, and has to be wired to a preview or a compositing node to be seen at all.

The page is specific about the failure case for that one: when decoding fails, the run report says so explicitly and a grey placeholder is emitted rather than something that looks like a generated frame.

Two model constraints come with the inputs. The CLIP loader's type must be set to `minimax`, which is Qwen3-VL, and the UNET is chosen by mode.

## One bundled workflow's filename disagrees with its own task type

The repository ships eight example workflows, and each row of the table that lists them gives a filename, a `task_type` and which UNET the workflow expects. The mapping is consistent except for one.

Six modes exist in total: `t2v` for text to video, `i2v` for image to video, `fl2v` for first and last frame, `r2v` for a reference subject or asset group, `v2v` for video to video, and `rv2v` for editing a video with a reference. The first three run on the `fl2va` UNET and the last three on `ref2va`, which the table bolds for the reference modes.

The file named for external image-to-video groups, `minimax_h3_director_external_groups_i2v.json`, is listed with a `task_type` of `fl2v`. So the name and the row disagree, and getting it wrong costs you the wrong UNET rather than a subtle quality difference, since the two UNETs are different files.

The other external-groups workflow, the one for reference groups, is consistent, and the last row covers a refine pass wired to an external Refine node, where both the refined and pre-refine outputs are used.

## Ultralytics is a hard dependency for a feature whose weights you supply by hand

The fourth package in the requirements file is the only one that needs something the installer cannot provide, and the comment above it says so in three lines.

```text
# Director FaceRefine (track + crop + stitch). Unconnected Director does not
# import this at execute time. Weights are not pip-installable: put
# face_yolov8m.pt (or similar) in models/ultralytics/bbox/
ultralytics>=8.0
```

So a face detection and tracking library is a mandatory install for every user of the node, in exchange for a feature that only runs when you have connected something that triggers it, and whose weights you have to find yourself and place in a directory inside the model folder by hand.

That is the opposite arrangement from the NVIDIA package two entries earlier, and the reasoning is inconsistent in the same file. The NVIDIA one is left out of the install because a hard dependency would break machines. The Ultralytics one is left in the install even though it cannot function without a file you supply, so a user with no intention of using face refinement still gets the dependency.

The page's feature table, for its part, does not list face refinement among the capabilities at all, so the dependency has no corresponding entry a user could find in the documentation.

## The weights are not in the repository and the first source is a blog

Everything this node needs to run is a model file, and none of them is here. The page points at three places: an article on a third-party site described as the complete resource pack of weights and example workflows, the organisation's Hugging Face repository, and the ComfyUI documentation's own MiniMax H3 tutorial.

The recommended files are specific. Two UNETs, a pruned int8 build with convolution rotations for the `fl2va` modes and another for the `ref2va` modes, both under the diffusion models directory. A CLIP named for Qwen3-VL at 32B in an `nvfp4_awq` format, which is a four-bit quantisation. And two VAEs, one for video in fp16 and one for audio in fp32.

Two things follow from those names. The recommended UNETs are int8 and the text encoder is four-bit, so the suggested setup is a quantised one rather than the full precision release. And the model supply is decoupled from the plugin: the plugin is Apache-2.0 and installs from its repository, while the weights come from a separate official integration and, first on the list, from an article hosted elsewhere.

The example workflows themselves are included, in an `example_workflows` directory, so the graph side is reproducible even though the weights are not.

## The pack format is ASCII only and one bundled filename is not

The director pack is the feature that makes a segment plan portable. It is exported as a zip with ASCII-only paths so that it does not depend on the interface language, and the layout is fixed:

```
pack.json
shared_params/shared_params.json
shared_params/Picture1.png
asset_groups/01/group.json
asset_groups/01/Picture4.png
timeline.json
```

The naming rule that keeps it unambiguous is that slot numbers continue rather than restart. When shared parameters occupy the first three picture slots, a group's files continue from `Picture4`, and the page is explicit that you should not rename a group's first picture to `Picture1`.

Packs exclude the UNET, CLIP and VAE, so they carry no weights. Import overwrites the current node timeline, with a confirmation, and media lands in a dedicated input subfolder.

Against that ASCII discipline, one of the eight bundled example workflows is named in Chinese characters, `minimax_h3_director_二采_加速.json`, describing a second-pass speedup. The pack format warns about path encoding problems and then the repository ships a workflow filename that would need quoting on some systems.

## Three READMEs sit at the root and the retrieved one is the Chinese page

The repository carries three readme files side by side at the top level: the default one, an English version and a Chinese version. The page this description is based on is the Chinese one, with an English link at the top of it, and the page structure is that of a Chinese technical document: a capability table, a pack path table, a model file table and an example workflow table.

That matters for anyone reading the repository rather than this description. The strings that matter for wiring are all in English regardless, since node names, port names, `task_type` values and pack paths are identifiers, and the page is careful to point that out when it matters, noting that pack paths are ASCII and independent of the current interface language.

The rest of the tree is arranged the way a ComfyUI custom node usually is: an `__init__.py` entry point at the root, `nodes/` for the node definitions, `lib/` and `director/` for the logic, and a `web/` directory for the interface extension. The project is recorded as JavaScript, which is the interface layer; the node itself is Python.

There is also a `.comfyignore` at the root, which is the file ComfyUI's own tooling uses when deciding what to ship to a registry listing. The registry block in the manifest names a publisher and a display name, and leaves its icon field empty.

## Inter-segment guidance is off by default and wants twenty-two context frames

The feature that makes segments join rather than cut is called inter-segment guidance, and it is off by default. When enabled for a multi-segment run in any of the six modes, it pins the trailing motion of the previous segment's output, and the generated audio with it, into the sampling of the next segment, then trims the prefix back off.

Its one exposed parameter is the context frame count, and the page offers four values: 5, 22, 39 and 56, with 22 marked as the recommended default. That is a short window in absolute terms, so the continuity being borrowed is the last fraction of a second of motion rather than a semantic summary of the previous segment.

The page credits another project for the implementation idea, which is worth noting because it tells you where to look if the behaviour is wrong.

Two other continuity paths exist with different rules. The first and last frame mode has its own separate timeline with multiple keyframe groups, where an empty group can borrow the last frames of the previous segment when inter-segment guidance and the reference-previous option are both set. And the SelfLift companion changes the first sampling pass itself into a low-resolution prefix, a 3D lift, and a high-resolution tail. Both are off unless the relevant external node is connected.

## Conclusion

This director is worth reading if you generate long segmented video and want the segment plan to be an object you can inspect and edit rather than a prompt you rewrite, because the timeline, the asset groups, the six task modes and the director pack are all addressed directly instead of through an abstraction. Two things to check before you build on it. The weights are not in the repository and the first source the page names is a third party article, so treat the model supply as a separate problem from the plugin. And the manifest's NVIDIA CUDA classifier does not match the portability decision the requirements file makes in writing, so judge the hardware requirement from the code rather than from the listing.

## FAQ

### What does ComfyUI MiniMax H3 Director need before it loads?

ComfyUI v0.30.0 or newer, including the official MiniMax H3 nodes, and the plugin's own requirements file, which installs opencv-python-headless, imageio-ffmpeg and scenedetect. The manifest records requires-comfyui of 0.30.0 or newer.

### Which model files does the MiniMax H3 Director need?

Two UNETs, a pruned int8 build for the fl2va modes and a pruned int8 build for the ref2va modes, plus a Qwen3-VL CLIP whose loader type must be set to minimax, a video VAE and an audio VAE. None of the weights are in the repository.

### How many task modes does the director support?

Six, selected with task_type: t2v, i2v, fl2v, r2v, v2v and rv2v. The first three use the fl2va UNET and the last three use ref2va.

### What is inside a MiniMax director pack?

A zip holding pack.json, a shared_params folder, numbered asset_groups folders and timeline.json for a lossless round trip. It deliberately excludes UNET, CLIP and VAE weights, and importing overwrites the current node timeline after a confirmation.

### Does the MiniMax H3 Director need an NVIDIA GPU?

Not for the node itself. The requirements file keeps the NVIDIA VSR package out as a hard dependency on purpose, noting that doing so would break AMD, cloud and no-VSR installs. The manifest still classifies it as an NVIDIA CUDA environment, and the RTX VSR upscale path does need an NVIDIA GPU.

## Sources

- [AIMixer/ComfyUI_MiniMaxH3_Director on GitHub](https://github.com/AIMixer/ComfyUI_MiniMaxH3_Director)
- [Issues](https://github.com/AIMixer/ComfyUI_MiniMaxH3_Director/issues)
- [License: Apache-2.0](https://github.com/AIMixer/ComfyUI_MiniMaxH3_Director/blob/main/LICENSE)
- [README](https://github.com/AIMixer/ComfyUI_MiniMaxH3_Director/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/aimixer-comfyui-minimaxh3-director
