# Minimax-H3-Turbo: Distilled LoRA Checkpoints for Fast Video Generation

> Minimax-H3-Turbo provides distilled LoRA checkpoints for MiniMax-H3 video generation that reduce inference to 4 or 8 steps, with Diffusers and ComfyUI formats, an online studio, and a reference image resizing policy tuned to match the distillation training.

**ModelTC/Minimax-H3-Turbo** — Distill Minimax-H3 into 4 steps

- Repository: https://github.com/ModelTC/Minimax-H3-Turbo
- Stars: 372 · Forks: 15
- Language: Python
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/modeltc-minimax-h3-turbo

## What Minimax-H3-Turbo Does

MiniMax-H3 is a video generation model that supports three task types: first-last frame to video with audio (FL2VA), text to video with audio (T2VA), and reference image to video with audio (Ref2VA). The standard model uses a full denoising schedule at inference time, which requires many forward passes. Minimax-H3-Turbo reduces that cost through distillation: LoRA adapters checkpoint the denoised state at earlier steps, allowing the model to produce outputs in 4 or 8 steps instead of the full schedule.

This repository, maintained by ModelTC, provides the distilled LoRA checkpoints, Diffusers-compatible batch inference code, ComfyUI workflows, and documentation covering the inference settings that match the distillation training. The LoRA weights are hosted on Hugging Face at lightx2v/Minimax-h3-Turbo. The repository itself does not contain the base MiniMax-H3 model weights, which must be downloaded separately.

An online studio for testing the Turbo LoRA without local hardware is available at x2v.light-ai.top. The studio currently deploys the FL2VA Turbo 8-step v1.0 768p checkpoint with recommended settings of 8 NFE, video shift 6, and audio shift 3. An API at x2v.light-ai.top/api-docs provides programmatic access for integrating MiniMax-H3 Turbo into external applications without running local inference.

## Five Available Checkpoints Across Two Resolutions

The repository documents five checkpoint variants, each with a Diffusers safetensors file and a ComfyUI safetensors file on Hugging Face:

FL2VA Turbo 4-step v0.1 runs at 544p with a mixed aspect ratio. Training used video shift 12 and audio shift 3. Inference runs in 4 steps.

FL2VA Turbo 8-step v1.0 also runs at 544p with mixed aspect ratio. Training used video shift 12 and audio shift 3. Recommended inference is 8 steps, though 4 steps is also listed as a viable option.

FL2VA Turbo 4-step v1.0 768p runs at 768p resolution, fixed at 1344x768. Training used video shift 6 and audio shift 3. Inference runs in 4 steps.

FL2VA Turbo 8-step v1.0 768p also runs at 768p. Training used video shift 6 and audio shift 3. Recommended inference is 8 steps. This is the checkpoint currently deployed in the online studio.

Ref2VA Turbo 4-step v0.1 handles the reference-image-to-video task at 544p with mixed aspect ratio. Training used video shift 12 and audio shift 3. Inference runs in 4 steps.

All FL2VA checkpoints also support T2VA (text-to-video-with-audio), since the text conditioning path shares the same architecture. There is no 768p variant listed for Ref2VA.

## Understanding the Inference Step Budget: NFE and Shift Parameters

The number of function evaluations (NFE) and the shift parameters for video and audio together define the specific denoising path each checkpoint was trained to follow. The README provides the full calculation for NFE=4 with video shift=12 and audio shift=3.

For NFE=4, the N transformer evaluation points on the unshifted grid are q_i = (N - i) / N where i = 0, 1, ..., N - 1. With N=4 this gives q = [1, 0.75, 0.5, 0.25]. Applying video shift 12 transforms these to sigma values [1, 0.9730, 0.9231, 0.8000] -> 0. Applying audio shift 3 transforms them to sigma values [1, 0.9000, 0.7500, 0.5000] -> 0. Each list uses exactly four function evaluations.

For the 768p checkpoints, the video shift is 6 rather than 12, giving a different sigma sequence that better matches the training distribution at that resolution. Using the wrong shift value at inference time means the evaluation points will not match where the distillation was trained to denoise, which degrades output quality without an obvious error message.

The audio shift of 3 is consistent across all five checkpoint variants listed in the README, regardless of resolution or step count. Keeping this parameter fixed when switching between checkpoints is important for audio quality in the generated outputs.

## Three Reference Image Resizing Policies for Ref2VA

For the Ref2VA task, the repository documents three reference image resizing policies that control how the reference image is scaled to fit the target video canvas:

The match mode scales the reference so its pixel area matches the target canvas area, preserving the reference aspect ratio and never upscaling a smaller reference. The scale factor is min(1, sqrt(target_area / ref_area)). This is the mode used during distillation training.

The max mode preserves the reference aspect ratio and only scales down references whose short edge exceeds 2048 pixels. The scale factor is min(1, 2048 / ref_short_edge). It makes a conservative choice for large reference images.

The diffusers mode preserves the reference aspect ratio and forces the short edge to exactly 2048 pixels, matching the original Diffusers behavior. The scale factor is 2048 / ref_short_edge, which will upscale small reference images.

All three policies keep the reference aspect ratio, use the H3 resolution grid where dimensions are rounded to multiples of 32, and avoid cropping the reference content. The Ref2VA inference entry point exposes the mode through --reference-resize-mode and defaults to match. The README recommends using match for inference with the distilled models because the distillation training used that policy. Switching to the diffusers mode restores the original Diffusers 2048-pixel short-edge behavior and produces different pixel-area scaling than the training distribution used.

## Resolution Grid and the 32-Pixel Rounding Constraint

All three resizing policies apply the H3 resolution grid, which requires that both dimensions be rounded to multiples of 32 pixels. The resolution_util.py file in the repository root implements this rounding. The 544p checkpoints support a mixed aspect ratio during training, which means the model was exposed to frames with varying width-to-height ratios that all resolve to pixel counts divisible by 32 at approximately 544 vertical pixels. The 768p checkpoint uses a fixed 1344x768 resolution, giving a 1.75:1 aspect ratio.

For FL2VA tasks, the first and last frame must share the same dimensions. The aspect ratio of these frames determines what the intermediate video frames will look like. For Ref2VA tasks, the reference image is resized to fit the target canvas using the chosen policy before it is passed to the model. The round-to-32 step ensures that the pixel grid aligns with the transformer's patch size, which the H3 architecture requires to avoid padding artifacts.

The resolution_util.py file is referenced alongside the inference script in the top-level repository entries, suggesting it is a dependency for the inference code rather than a standalone utility.

## Running Inference with Diffusers and ComfyUI

The repository provides two inference paths. For Diffusers-based inference, the inference_minimax_h3.py script handles batch processing, and the minimax_h3_ref2va_pipeline.py script is specifically for the Ref2VA pipeline. Full setup instructions for the Diffusers path are in DIFFUSERS_SETUP_AND_INFERENCE.md in the repository root. This file covers dependency installation, LoRA loading, and the configuration options that match each checkpoint's training settings.

For ComfyUI users, the example_workflows/ directory contains workflow files that connect the base model with the Turbo LoRA. Full setup instructions for ComfyUI are in COMFYUI_SETUP_AND_INFERENCE.md. ComfyUI's MiniMax H3 nodes handle the LoRA loading natively, and the reference image resizing guidance in ComfyUI's MiniMax H3 documentation describes the same three resizing policies documented here.

The examples/ directory includes test image sets (i2va_testset/ and ref2va_testset/) and three prompt files: prompts_i2va_test.json, prompts_ref2va_test.json, and prompts_t2va_test.json. These serve as starting points for verifying that your setup produces reasonable outputs before running on your own content.

The README notes that the model version deployed in the online studio may change over time, so the studio is not a stable reference for evaluating a specific checkpoint's output quality. Use the offline inference path for reproducible evaluation.

## Limitations and Comparisons

Minimax-H3-Turbo provides LoRA adapters and inference scripts only. Running inference requires downloading the base MiniMax-H3 model weights separately before the LoRA can be applied. The repository does not include training code, distillation code, or instructions for creating new distilled variants. Researchers who want to distill their own step count cannot do so from this repository alone.

The 544p checkpoints use a mixed aspect ratio training setup, which gives some flexibility in input dimensions. The 768p checkpoint uses a fixed 1344x768 resolution, so inputs or reference images with a significantly different aspect ratio will be resized in ways that may not match the scene composition you intend.

For Ref2VA specifically, only a 4-step 544p variant is available. There is no 768p Ref2VA checkpoint listed in the README at the time this documentation was written. Users who need higher-resolution reference image conditioning have no Turbo option for that task yet.

A comparable alternative in the video generation distillation space is consistency distillation applied to other video diffusion models, such as the Wan2.1 family with accelerated samplers. Those models produce video-only output without a joint audio stream. MiniMax-H3's distinction is that it generates video and audio together in a single model pass, with the audio shift parameter built into the distillation schedule. The last push to this repository was on 2026-08-27. The project is Apache-2.0 licensed.

## Conclusion

Minimax-H3-Turbo is a direct tool for practitioners who need faster inference from MiniMax-H3 and are comfortable with Diffusers or ComfyUI workflows. The repository does not include training code or full model weights; it provides LoRA checkpoints and inference scripts only. Use the match reference image resizing mode for Ref2VA tasks, since the distillation training used that mode, and switching to the diffusers mode will produce images with different pixel-area scaling than the training distribution.

## FAQ

### How many inference steps does Minimax-H3-Turbo require?

The distilled checkpoints support 4-step and 8-step inference depending on the variant. The 4-step v0.1 and v1.0 checkpoints require 4 NFEs. The 8-step v1.0 checkpoint recommends 8 steps but can also run at 4.

### Which reference image resizing mode should I use with Minimax-H3-Turbo?

The README recommends the match mode for inference with the distilled models because the distillation training used that resizing policy. Passing --reference-resize-mode diffusers restores the original Diffusers 2048-pixel short-edge behavior, which uses different pixel-area scaling than the training distribution.

### Where can I try Minimax-H3-Turbo without a local setup?

The online studio at x2v.light-ai.top lets you test the FL2VA Turbo 8-step v1.0 768p checkpoint directly in a browser. An API is available at x2v.light-ai.top/api-docs for integration into external applications.

## Sources

- [Issues](https://github.com/ModelTC/Minimax-H3-Turbo/issues)
- [License: Apache-2.0](https://github.com/ModelTC/Minimax-H3-Turbo/blob/main/LICENSE)
- [ModelTC/Minimax-H3-Turbo on GitHub](https://github.com/ModelTC/Minimax-H3-Turbo)
- [README](https://github.com/ModelTC/Minimax-H3-Turbo/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/modeltc-minimax-h3-turbo
