# ComfyUI-MiniMaxH3-Easy declares 1.3.2 while its newest tag is 1.3.0

> A ComfyUI node suite that collapses text, image, first and last frame, reference and digital human generation into one node. The packaging runs ahead of the tags, the Manager install is a nightly, and the model loader no longer checks filenames.

**nkxx188/ComfyUI-MiniMaxH3-Easy** — The easiest way to use MiniMax H3. One compact workflow for T2V, I2V, first/last-frame, and reference video generation, with a unified multi-media input, powerful @ references, and inline dialogue blocks. Less wiring. More creative control.

- Repository: https://github.com/nkxx188/ComfyUI-MiniMaxH3-Easy
- Stars: 801 · Forks: 57
- Language: JavaScript
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/nkxx188-comfyui-minimaxh3-easy

## pyproject declares 1.3.2 and the newest tag is 1.3.0

The packaging file is short and specific. The distribution is comfyui-minimaxh3-easy, the version is 1.3.2, the licence is the LICENSE file, and the dependencies list has exactly two entries, requests and psutil. The release tags tell a different story. The three most recent are v1.3.0 on 2026-09-13, v1.2.0 on 2026-09-01 and v1.1.0 on 2026-08-30, and v1.3.0 is the highest of them. So the declared version sits two patch increments above the newest tag, and there is no tag that corresponds to what the package says it is. The ComfyUI publisher block is there too, with a PublisherId of nkxx188 and a DisplayName matching the repository name, and an empty Icon field. For a suite people install by cloning into custom_nodes, the practical version a user reports is whichever of the three axes they happened to land on.

## The Manager install is a nightly, and the published versions lag

There are two documented install routes and they are not equivalent. The clone is one line:

```bash
git clone https://github.com/nkxx188/ComfyUI-MiniMaxH3-Easy.git
```

meant to be run inside ComfyUI/custom_nodes, with no further setup step shown. The other route is ComfyUI Manager, where the instruction is to search for ComfyUI-MiniMaxH3-Easy and then, in bold, select the Nightly version during installation, because Nightly is the current up-to-date release and other published versions may lag behind this repository. That is an unusual thing to have to tell a Manager user, because it inverts the normal assumption that the listed release is the current one. Put together with the version gap above, there are three moving points: the Manager nightly, the tags, and the version string in pyproject.toml. The page also asks you to update ComfyUI to a recent version that includes the official MiniMax H3 nodes before installing, and to restart ComfyUI after installing or updating Python files.

## The model loader lost its filename whitelist

The Easy Loader can select FL2VA, Ref2VA, the text encoder and both VAEs directly, and it does so by listing files rather than by recognising them. The filename whitelist and the naming filters have been removed, so the node shows every available file from the corresponding ComfyUI model directories and you select the correct file for each role. That is a deliberate simplification with a real consequence: nothing in the node tells you that you handed it a GGUF when the slot wanted a text encoder, and the first sign of a mistake is whatever fails downstream. The file locations are otherwise simple:

```text
ComfyUI/models/diffusion_models/
ComfyUI/models/text_encoders/
ComfyUI/models/vae/
```

with a fourth directory, ComfyUI/models/latent_upscale_models/, for the H3 3D latent upscaler used by the Latent Upscale refinement. The suite also has an escape hatch for the other direction: native, community or GGUF loaders can be connected through MiniMax H3 Easy Model Adapter, so a mixed setup is expected rather than discouraged.

## Text-to-video is the absence of media, not a mode of its own

The generation table has five rows and two of them are conditional on what you connect rather than on a setting. Text-to-video says to select either I2V or First/Last Frame or Reference-to-video and connect no media. Both of those modes automatically become text-to-video when no media is connected, so text generation is not tied to only one mode. Which means the first row of the table describes the same two modes as the rows below it, with an empty input, and the mode selector never rejects an empty graph. The other rows do constrain their inputs: image-to-video is one image, first/last-frame is two images, reference-to-video takes images, videos or standalone audio with at least one image or video required, and Digital Human takes visual references plus exactly one driving audio track. One behaviour is stated for the frame modes specifically: first and last frame images are center-cropped when they need to fit the selected canvas, and are never stretched horizontally or vertically.

## Two transformers and an implicit preference between them

The suite expects two transformer models and picks between them on its own. When both FL2VA and Ref2VA are configured, regular text and image generation prefers FL2VA and full-reference generation prefers Ref2VA. If only one transformer is configured, that model serves every mode. So the choice between the two is a function of which files are present, not something the generation table exposes as a control, and the same workflow file will behave differently on two machines with different model directories. That is a reasonable design when you only want text-to-video and it quietly substitutes for you when you only have one model. It also means the mode you picked in step two of the quick start is not by itself a promise about which weights will run. The quick start itself is five steps: import workflow/1.MiniMax_H3_Easy.json, pick the models in the loader, pick a mode, type a prompt and connect media, then choose resolution, aspect ratio and duration and queue.

## Digital Human falls back to reference-to-video instead of failing

The digital human mode has one behaviour worth planning around. When it is selected and a single Media audio item is connected, that item is locked into the generated result as its driving track and is no longer treated as ordinary reference audio; the visual references alongside it may be images or videos. If no audio is supplied, the node automatically falls back to ordinary reference-to-video instead of failing. A silent substitution means a graph wired the wrong way produces output of a different kind rather than an error, and the only way to notice is to look at what came out. The mode also differs from the others in input count: exactly one driving audio track, not the at-least-one rule that reference-to-video applies. Elsewhere in the suite the same preference for degrading gracefully shows up in the prompt layer, where typing @ inserts a reference and typing # opens a dialogue block, both without the user writing the underlying tags by hand.

## Only one of four media paths caches decoded video

There are four ways to get media into a node: Media Loader, direct Media-port links, Media Bridge for API and headless work, and Media Splitter, which is the reverse utility and expands one Media Bundle into standard IMAGE, VIDEO and AUDIO outputs with a configurable number of ports. Only Media Loader caches decoded video references, and the page is careful about what that buys you: after a video is decoded once, later generations reuse it through ComfyUI's node cache, which reduces loading and decoding time when the same reference video is reused, and replacing or modifying the file invalidates the cache automatically. It saves video preparation time, not the model's sampling time. Direct links, Media Bridge and Media Splitter do not get that cache. Media Loader can hold a large shared library, with the actual media limits applied by whichever consuming node runs, and it accepts images, videos and audio dragged from the system file manager or pasted with Ctrl+V on the selected node.

## The dialogue block section stops mid-word

The prompt editing section covers media references in full and then stops. Typing @ in a reference or context prompt opens a picker for connected media, references can be shown by index or by filename, and they are converted into the H3 tags the runtime needs; images, videos and audio clips are numbered independently, so @Image1, @Video1 and @Audio1 can all exist at once. The next subsection is headed Dialogue blocks and raw view, and its only sentence begins Type # to create a dialogue block. It is convert. So the feature the heading promises, dialogue blocks and a raw view, is named once and not explained, and the conversion the sentence starts describing is where the page stops. The repository does hold the material behind the prompt layer, with a prompt_guides/ directory and a sampling_strategies.py alongside nodes.py, h3_latent_upscaler.py and a tests/ directory, so the behaviour is documented somewhere other than here.

## Conclusion

This suite suits somebody who already runs ComfyUI with the official MiniMax H3 nodes and wants one node instead of a wall of wiring, especially for segmented long video where per-segment prompts and context carryover are the point. Before installing, settle which build you want, because there are three version axes and they disagree: pyproject says 1.3.2, the newest tag is v1.3.0, and the Manager channel that is described as current is a nightly. Then decide how much validation you want on model files, because the loader's filename whitelist and naming filters were removed on purpose and the wrong pick for a role is only reported by whatever fails downstream. And read the mode table against the fallback rules rather than against the labels, since text-to-video is not its own mode and Digital Human quietly becomes reference-to-video when its audio track is missing.

## FAQ

### Where do I put the MiniMax H3 model files for ComfyUI-MiniMaxH3-Easy?

Regular models go in ComfyUI/models/diffusion_models/, ComfyUI/models/text_encoders/ and ComfyUI/models/vae/. The H3 3D latent upscaler for the Latent Upscale refinement goes in ComfyUI/models/latent_upscale_models/.

### How do I install ComfyUI-MiniMaxH3-Easy?

Either clone it into ComfyUI/custom_nodes, or search for ComfyUI-MiniMaxH3-Easy in ComfyUI Manager and select the Nightly version, which the page says is the current up-to-date release. Update ComfyUI to a recent version with the official MiniMax H3 nodes first, and restart after installing.

### Does ComfyUI-MiniMaxH3-Easy check which model file I picked?

No. The Easy Loader's filename whitelist and naming filters were removed, so it shows every file from the matching ComfyUI model directory and you have to select the correct file for each role yourself.

### What happens if I use ComfyUI-MiniMaxH3-Easy Digital Human mode with no audio?

The node falls back to ordinary reference-to-video instead of failing. With audio connected, that single audio item is locked into the result as the driving track rather than treated as ordinary reference audio.

### Does ComfyUI-MiniMaxH3-Easy speed up reuse of the same reference video?

Only through Media Loader, which is the one media path that caches decoded video references through ComfyUI's node cache and invalidates the cache when the file is replaced. It saves preparation time, not sampling time.

## Sources

- [Issues](https://github.com/nkxx188/ComfyUI-MiniMaxH3-Easy/issues)
- [License: MIT](https://github.com/nkxx188/ComfyUI-MiniMaxH3-Easy/blob/main/LICENSE)
- [nkxx188/ComfyUI-MiniMaxH3-Easy on GitHub](https://github.com/nkxx188/ComfyUI-MiniMaxH3-Easy)
- [README](https://github.com/nkxx188/ComfyUI-MiniMaxH3-Easy/blob/main/README.md)
- [Releases](https://github.com/nkxx188/ComfyUI-MiniMaxH3-Easy/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nkxx188-comfyui-minimaxh3-easy
