Open-source project
Rimagination/h3lite avatar
Rimagination/h3lite

h3lite routes a MiniMax H3 video request through Windows, NVIDIA and ComfyUI

A hardware-aware Codex skill for local MiniMax H3 video generation through ComfyUI

454 stars49 forksPythonMIT

At a glance

What is it?
H3 Lite is a Codex and WorkBuddy skill that prepares a local MiniMax H3 video pipeline inside ComfyUI. Its own two timings put an 8 GB laptop card roughly 7.7 times behind a 16 GB desktop card, and its case list ends partway through its last line.
Who is it for?
H3 Lite earns a trial if you already run ComfyUI on Windows with an NVIDIA card and can absorb a first install measured in model downloads: on the project's own numbers a 16 GB card finished the same clip in about 77 seconds where an 8 GB laptop card needed about 591.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 39 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The verified path is one matrix row, and the Mac route belongs to another project

The support table in the Chinese document lists four platforms and marks one as the main supported route: Windows with NVIDIA, using the ComfyUI, doctor, planner and fastpath pieces that ship in the repository. macOS on Apple Silicon is a community route that points at the MLX and Metal tool `mmh3turbo`. macOS on Intel is unverified, and the advice there is to fall back on a hosted or API backend. Linux with NVIDIA is labelled experimental, and the reader is told to adapt paths, nodes and run parameters on their own.

The Apple Silicon row is where the only literal command in the whole document appears, `uvx mmh3turbo`, paired with community weight bundles hosted at huggingface.co/yunfengwang/mmh3turbo-bundles. Those parts are stated to be used apart from the ComfyUI components of this skill. So an owner of a Mac ends up with a different codebase: none of the four step flow, no `anchors.json`, no `anchor_qa` and no component manifest come with it. The repository description calls the skill hardware aware, and what that resolves to in practice is one Windows and NVIDIA configuration plus a pointer elsewhere for everybody else.

Two cards, identical settings, a gap of about 7.7 times

The document prints two timings and no others. Under the same Set B, the same compatible workflow, the same prompt, the same seed and the same 640x352 canvas at 4 steps, an RTX 4060 Ti 16 GB finishes in about 77.08 seconds while an RTX 4070 Laptop 8 GB needs about 591.22 seconds. Dividing the second figure by the first gives roughly 7.7, and both machines are listed as verified starting points.

The difference between them is the route label, not the number in the product name. The desktop machine is tagged `NORMAL_VRAM` with Set B for text to video and image to video. The laptop is tagged `LOW_VRAM`, where Set A handles both and Set B is a compatible route for text to video only. So the set that runs on the slow card is also the one treated as a compatibility path there, and memory capacity decides which of the two paths you land on. The same section adds that system memory, pagefile, disk speed and laptop power draw all move the result.

The default preset is `fast`, which is 4 steps with native audio at 640x352. `balanced` is 6 steps and `quality` is 8. The shipped image to video sample is not on the default: it runs 8 steps at 864x480 on the Set A compatible route.

A four reference request can return as a single image job

The routing table maps four intents onto four routes. Text only goes to `T2VA`. A named first frame goes to `I2VA`, with the reference pinned at 0.00s and only later change described. First and last frames go to `FL2VA`. Several image, video or audio references go to `Ref2VA`, where each asset first needs a role, a set of things to keep and a set of things allowed to change, followed by a storyboard.

That last route is the fragile one, because the document states that when the Ref2VA components are incomplete the agent switches to I2VA or marks the run as experimental. Nothing in that sentence surfaces a failure to the person who asked for four references, and the substitution is the described behavior rather than an error path. The thing to check after a run is therefore which route executed, not only whether a video came back.

The bookkeeping exists for exactly that question. Anchor cards are written to `anchors.json` and their paths are recorded in `manifest.json`, and reference assets get stable names such as `Subject A` and `Picture 1`. The check called `anchor_qa` then samples the first, middle and last frames for continuity, and the same sentence concedes that identity, clothing and composition still need human review. Three sampled frames is the whole automated pass.

Multiples of 32 bend the resolution labels

Every listed resolution divides by 32, and the table pays for that rule in its own labels. The row called about 1080p is 1920x1056, twenty four lines short of 1080, because 1080 is not a multiple of 32 and 1056 is. The rule also puts two rows at the same height: 1216x704 is labelled about 704p and 1280x704 is labelled about 720p, though both are 704 pixels tall and differ only in width.

The 16:9 label is looser still. 864x480 is a ratio of 1.80, 1376x768 is 1.79, both 1280x704 and 1920x1056 are 1.82, and 1216x704 is 1.73. Rows filed under the same aspect ratio span about 0.09 of ratio, and the default canvas is at the wide end of that span.

When nothing is specified the default is 640x352 at 0.23 MP, with 864x480 offered as the common 16:9 quality step at about 0.4 MP. The 8 second starship case asks for a 16:9 video without naming a resolution, so it lands on the default 640x352 canvas rather than on 1.78. For expensive work the guidance is to preview at a low resolution first and raise the canvas only afterwards.

Nothing to run, a sentence to send, and two extraction codes

The quick start is a Chinese sentence to hand to Codex or WorkBuddy, not a command:

text
请帮我安装 H3 Lite,并根据我的电脑配置准备本地 MiniMax H3 视频生成环境:
https://github.com/Rimagination/h3lite

Keeping the environment off the system drive means adding a second sentence to the same message:

text
请把 MiniMax H3 和 ComfyUI 安装到 F:\MiniMax-H3;如果那里已经有健康环境就直接复用。

First install time is left to model size, network and disk speed. The manual path is the repository page, pick Code and Download ZIP, unpack it, drop the `h3lite` folder into the Codex skills folder, then reopen Codex. Models come from two Baidu Netdisk links with extraction codes `4hri` for Set A and `1hjx` for Set B, and what gets merged into `<ComfyUI>` is the `models` and `custom_nodes` directories plus an imported workflow JSON, with `component-manifest.json` kept on the side.

The top level of the repository holds `scripts/`, `tests/`, `agents/` and `references/`, and the Windows row names doctor, planner and fastpath, yet no invocation for any of them appears anywhere in the document. There is no version pin, no checksum and no upgrade path either.

Six acceleration variants and not one timing

The Set A comparison holds the rolling red ball scene, the prompt, the Set A components and 640x352 at 4 steps with native audio constant, and changes only the acceleration node combination. The six variants are the compatibility baseline, Sage only, FFN only, Block Cache only, Sol only, and all four nodes together. Each clip is 5 seconds at 640x352.

All six results are presented as posters that open a GitHub Pages player fed by the MP4 files in `docs/videos/`, with the `assets-rolling-redball` release offered as a download backup. No duration, no memory reading and no quality note is attached to any variant, so the one comparison that changes something measurable ships as pictures. The two timings in the document belong to Set B, the other component set, which means the node comparison and the measured comparison never touch the same build.

Both releases were published on 2026-08-20 and are asset drops rather than packages: `assets-rolling-redball` and `assets`. No versioned software release exists, and the last push on the default branch is dated 2026-08-23.

The document is Chinese and its last case stops inside a line

The copy shipped as README.md is written in Chinese, with README.en.md linked from its header and both files present in the repository tree. Every worked prompt is Chinese, and the writing template for vague requests is references/prompt-assist.md, linked twice.

One wording rule is worth knowing before a first run, because it decides whether you get sound. A request for no dialogue keeps ambient and action sound; only a request for complete silence turns the audio off. Short prompts are meant to be organized in three parts, frame and atmosphere, action and camera in playback order, then sound.

Four cases are worked. A 5 second bouncing red ball serves as the first check of the whole chain:

text
请使用 H3 Lite,生成一个 5 秒横屏视频:一颗小型哑光红色橡胶球,在灰色混凝土地面上弹跳两次,然后向右滚出画面。低机位固定镜头,阴冷的多云日光,浅景深、35mm 电影质感;保留两次撞击地面的声音和滚动声,不配音乐。

A 5 second golden retriever waking up is split into a 0s-2s and a 2s-5s block. An 8 second starship jump is adapted from a MiniMax H3 official reproducible case. The 5 second image to video living room channel change uses a sample first frame from assets/examples, at 864x480, 8 steps, Set A compatible route, with original characters in an American adult animation style. After that player link the text ends partway through a line, at 点击封面打, so FL2VA and Ref2VA, the two routes the routing table promises, have no worked case in this document.

Editorial conclusion

H3 Lite earns a trial if you already run ComfyUI on Windows with an NVIDIA card and can absorb a first install measured in model downloads: on the project's own numbers a 16 GB card finished the same clip in about 77 seconds where an 8 GB laptop card needed about 591. Check manifest.json after any run to see which route actually executed, because a four reference request can finish as a single image job, and read the six way node comparison as a visual demo rather than a benchmark. Mac and Linux owners get a community route in another project and an experimental note as the real boundary of what is verified. Anyone who needs a runnable command, a version to pin, or timings for the six acceleration variants will not find them in the repository yet.

Frequently asked questions

What does h3lite need before it can generate a video?

Windows with an NVIDIA card and ComfyUI. macOS on Apple Silicon is a community route through mmh3turbo with separate community weight bundles, macOS on Intel is unverified, and Linux with NVIDIA is an experimental route that needs manual adaptation of paths, nodes and run parameters. Set A pairs a W4A8 diffusion model with a 4B INT4 text encoder, Set B uses a 4B FP8 text encoder, and only one component set is needed.

Why does an RTX 4070 Laptop need about 591 seconds where an RTX 4060 Ti needs about 77?

Both figures come from the same Set B, the same compatible workflow, the same prompt, the same seed and the same 640x352 canvas at 4 steps. The 8 GB laptop card runs the LOW_VRAM route while the 16 GB desktop card runs NORMAL_VRAM, and the same section notes that system memory, pagefile, disk and laptop power draw all affect speed.

What happens when the Ref2VA components are incomplete?

The agent falls back to I2VA or marks the run as an experimental route instead of reporting a failure. Anchor cards are saved to anchors.json and their paths recorded in manifest.json, and anchor_qa only samples the first, middle and last frames, so identity, clothing and composition still need a person to look at them.

What is the default output size and step count in h3lite?

The fast preset is 4 steps, native audio and 640x352, and all listed resolutions are multiples of 32. Balanced uses 6 steps and quality uses 8, and the shipped image to video sample is further off the default at 864x480 with 8 steps on the Set A compatible route.

How are the model components obtained?

Two Baidu Netdisk links with extraction codes, 4hri for Set A and 1hjx for Set B. You merge the models and custom_nodes directories into your ComfyUI folder, import the workflow JSON, and keep component-manifest.json. The manual route for the skill itself is Code and Download ZIP, unpack the h3lite folder into the Codex skills folder, then reopen Codex.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. Releases
  5. Rimagination/h3lite on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/rimagination-h3lite.svg)](https://hysenlabs.com/projects/rimagination-h3lite)