Open-source project
Songssx/ComfyUI-MiniMaxH3-TimelineDirector avatar
Songssx/ComfyUI-MiniMaxH3-TimelineDirector

The H3 Timeline Director manifest is at 0.8.0 and its only tag is v0.6.0

Editable reference-media timeline director for ComfyUI MiniMax H3 Reference to Video

560 stars59 forksPythonGPL-3.0

At a glance

What is it?
A ComfyUI plugin that puts reference video, soundtracks, guides, images and audio on one editable timeline, and then generates a long clip by splitting it into segments inside a single execution. The interesting parts are the rules it states about itself: a two-stage sampling split with a validated range, a step where locking the original audio fixes the master to silence, and a creator benchmark whose saver may re-encode the soundtrack.
Who is it for?
This is worth reading if you generate long clips on a local video model, because the segment continuation approach is the part that decides whether a two minute output has visible seams. Three things to check before you trust a result.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 15 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 6, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The manifest and the only tag disagree by two minor versions

The project manifest declares version 0.8.0.

The repository has exactly one release, tagged 0.6.0 and dated 2026-09-01, with a release note that names the headline feature it shipped, lightweight unlimited-length video generation. There are no other tags and no later patch releases.

So a user who installs from the release page gets the 0.6.0 build, and a user who clones the branch gets whatever 0.8.0 contains, with no tag to point at. There is a changelog in the repository, so the intervening work is documented, but nothing in the version numbers tells you which features are in which copy.

The packaging is otherwise minimal and worth reading for what it does not declare. The dependency list is empty, and so is the development extra group. That is unusual for a plugin that loads a latent upscaler model from disk and drives a video model, though it is consistent with the design: the upstream model and the host application provide everything, and the plugin's job is graph wiring.

The registry metadata is also local. The publisher identifier is a placeholder rather than an account, and the icon field is an empty string, so this plugin is installed from a clone rather than from a package index.

Long video is segmented inside one execution, not in a loop

The central claim is about how the work is divided rather than how fast it is.

Any target duration is split into continuous segments, and all of them complete in a single ComfyUI execution. The pieces that make that work are named: per-segment generation sharing one seed, direct continuation from the previous audio-visual latent tail, adaptive drift-control video masking, soft audio continuity across the join, overlap removal, and a final synchronized assembly.

Extending the clip means increasing the segment count. The documentation is explicit about what it is not doing: no generic loop nodes and no duplicated sampler chains. That distinction matters in a node graph, where the usual way to repeat a step is to duplicate it, and duplication is how you end up with nine copies of a sampler that have to be kept in sync by hand.

The length limits are named too, and they are all local: video memory, system memory, disk space, and whatever ceiling the host application puts on one execution. No model-side cap is claimed.

The path into that machinery is also the plain one. With no uploaded media and no segment windows, the finite segment sampler creates a standard empty latent for the model from the global prompt and the current duration, which means text to video goes through exactly the same code as the long case. Adding manual windows extends that same path to long form, so there is one route rather than two.

Two-stage sampling splits the step budget, and the range is checked

Progressive resolution sampling is the second idea, adapted for this model and built into the planner rather than bolted on as a separate custom node.

Turning it on means selecting an upscaler model from the host's latent upscaler directory and setting a high-resolution step count. The recommended starting point is given as a quarter of the scheduler's total steps, and the worked example is concrete: on an eight-step schedule, set two high-resolution steps, which runs six low-resolution steps, then the latent lift, then two high-resolution steps.

There is a validation rule, which is the kind of thing that saves an hour. The high-resolution value must be greater than zero and lower than the scheduler's total step count.

What the mechanism is not is worth as much as what it is. It is not a per-segment upscale shortcut. Every later segment carries two things: the preceding native low-resolution latent tail and the final high-resolution context, with drift control applied at both resolution stages.

The stated reason is accumulated error. Repeated resizing and repeated encoder round trips across dozens of segments compound, and the artefacts named are blur, white flashes, and visible seams. Carrying both resolutions forward is the claim that the quality loss is avoided rather than traded for speed.

The benchmark is one clip on one machine, and the saver may re-encode

One performance figure is given, and it is attributed to the creator rather than to the project.

A video at 1536 by 832 and 29 seconds, rendered directly in approximately ten minutes on the recommended schedule of 75 percent low-resolution and 25 percent high-resolution. The documentation immediately qualifies it: actual speed depends on the graphics card, its memory, the model, the step count, and how complex the references are.

The audio claims are qualified the same way. When standalone audio is locked, or when source audio from a reference video is enabled, a zero-denoise path restores one continuous source waveform, and the creator's own tests retain over 99 percent content and timing consistency. Then the sentence that earns the rest of the paragraph's credibility: the final MP4 saver may still re-encode the audio.

That last caveat is the one to carry away. A saved file can differ from what the graph produced, and the documentation says so rather than letting someone discover it by comparing waveforms.

The two demonstration clips are described with the same precision. Both were produced in one plugin execution, both are 52.625 seconds or 1263 frames at 24 frames per second, and both are hosted as release assets so that they add nothing to a clone or an installation. One is finite direct-latent continuation and the other uses references with a 48-frame overlap, which is the variable under test. The work is credited to a named third party with their own channels, rather than presented as the project's own output.

Two audio modes that contradict each other by design

The audio handling has a branch point, and the two branches do opposite things to your soundtrack.

One mode locks the original audio. For long-form digital humans, the source waveform is sliced along the timeline and injected through the native audio-visual mask and sigma path, so lip motion stays driven by the audio while the finished soundtrack preserves the uploaded recording unchanged. In both one-stage and two-stage reference-video generation, an enabled original-audio setting is automatically locked through that native path and restored as one continuous original waveform.

The other branch is the opposite. When that setting is disabled, the audio-visual stream and the final master are fixed to silence. Not degraded, not attenuated: silent.

And there is a priority rule on top: an explicitly uploaded locked audio asset takes precedence over the setting.

So a user who turns off original audio expecting the reference audio to remain gets silence, and a user who leaves it on expects their upload to be preserved byte for byte and may still find the saver re-encoded it. Neither behaviour is a bug, and both are stated. What is missing is a sentence telling you which mode you are in, so the practical step is to check the setting before a long render rather than after.

Output separation is handled separately: the graph produces merged outputs for timeline soundtracks and for standalone reference audio as distinct results.

Typing one segment prompt disables the global prompt for the whole graph

The prompt model has a rule that is easy to trip over, and it is stated precisely.

The global prompt is reused only when every segment prompt is empty. Entering a segment prompt requires you to complete every segment, and doing so disables the global prompt entirely. There is no per-segment override that falls back to the global text; it is all or nothing.

Selection is spatial rather than manual. Only media that intersects the cyan generation range takes part in the current reference or guide plan, so clips outside the range are ignored rather than uploaded and ignored. Each clip carries one of three modes: a fixed guide, an editable reference, or boundary only.

The media tokens are ordered stably from the interface through to the model inputs, with separate counters for pictures, videos and audio, so a reference that moved position in the timeline keeps its slot.

Two capacity rules sit on top of that. The global image and audio libraries have no upload-count limit at all, while every generated segment is still bounded by the model's own limits of at most nine reference images and three reference audio clips. So you can stage a hundred assets and still only get nine into any one segment.

The rendering compromises are also stated. Silent low-resolution monitoring proxies up to 480 by 270 at 12 frames per second are available for checking a timeline without paying for full quality, and decoding resizes to the node's configured dimensions specifically to protect video memory. The timeline itself is serialised into the host's workflow file, so a graph carries its own editing state.

Editorial conclusion

This is worth reading if you generate long clips on a local video model, because the segment continuation approach is the part that decides whether a two minute output has visible seams. Three things to check before you trust a result. The version you install from a tag is not the version the manifest declares. The two-stage sampling needs an upscaler model you supply from a specific directory, and the recommended split is a quarter of your step budget. And the audio behaviour has two branches that contradict each other on purpose, so read which mode you are in before assuming your soundtrack survived.

Frequently asked questions

How long a video can ComfyUI MiniMax H3 Timeline Director generate?

Any duration the local machine can support, split into continuous segments completed in a single execution. The stated limits are video memory, system memory, disk space and the host application's execution limits, with no model-side cap claimed.

What is the recommended two-stage sampling split?

A quarter of the scheduler's total steps at high resolution. On an eight-step schedule that means six low-resolution steps, then the latent lift, then two high-resolution steps, and the high-resolution value must be above zero and below the total step count.

What happens to the audio when original audio is disabled?

The audio-visual stream and the final master are fixed to silence. When the setting is enabled instead, the source waveform is sliced on the timeline and restored as one continuous original recording, and an explicitly uploaded locked audio asset takes priority over the setting.

What does the creator benchmark for this plugin measure?

A 1536 by 832, 29-second video rendered in approximately ten minutes on a 75 percent low-resolution and 25 percent high-resolution schedule, with speed dependent on the card, memory, model, step count and reference complexity.

How many reference images can one generated segment use?

At most nine reference images and three reference audio clips per segment, following the model's own limits. The global image and audio libraries themselves have no upload-count limit.

Does the H3 Timeline Director need any Python dependencies?

The manifest declares an empty dependency list and an empty development extra. The upstream model and the host application supply the rest, and the two-stage sampling needs a latent upscaler model you select from the host's own directory.

Official sources

  1. Issues
  2. License: GPL-3.0
  3. README
  4. Releases
  5. Songssx/ComfyUI-MiniMaxH3-TimelineDirector on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/songssx-comfyui-minimaxh3-timelinedirector.svg)](https://hysenlabs.com/projects/songssx-comfyui-minimaxh3-timelinedirector)