# h3.c: Native MiniMax-H3 Video and Audio Generation for Apple Silicon

> h3.c is a C implementation of the MiniMax-H3 inference engine for Apple Silicon, using Metal shaders and the Accelerate framework to run text-to-video and text-to-audio generation locally. It supports first and last frame conditioning, multimodal reference images, and an SSD streaming mode that reduces the DiT model's GPU memory footprint from around 36 gigabytes to about 2 gigabytes.

**antirez/h3.c** — MiniMax H3 inference engine for Mac computers

- Repository: https://github.com/antirez/h3.c
- Stars: 2,799 · Forks: 220
- Language: C
- License: MIT
- Published: 2026-09-16 · Updated: 2026-09-16 · Language: en
- Canonical page: https://hysenlabs.com/projects/antirez-h3-c

## What h3.c Does and What Hardware It Requires

h3.c is a native inference engine for the MiniMax-H3 model, written in C with Objective-C Metal shaders. It runs text-to-video generation, text-to-audio generation, and combined video-plus-audio generation directly on Apple Silicon without Python, PyTorch, or a cloud API. The Metal shaders and the Accelerate framework handle GPU execution on the unified memory architecture of Apple Silicon chips.

The README describes development and performance testing on M3 Max and M5 Max. It does not document support for Intel Macs or other platforms. The build system uses Clang with C11, and the Makefile links against the Foundation, Metal, MetalPerformanceShaders, MetalPerformanceShadersGraph, and Accelerate frameworks alongside libicucore.

The project requires FFmpeg and FFprobe to be available on PATH for media muxing and analysis. The MiniMax-H3 model weights must be downloaded separately from Hugging Face; the README assumes the snapshot is in ./MiniMax-H3 relative to where you run the binary.

## Building the Binary and Inspecting the Model

The build is a single make command:

```sh
make -j8
mkdir -p outputs
./h3 --info -d ./MiniMax-H3
```

The --info flag checks the model layout and prints the selected Metal device without mapping all weights or generating any media. It is the correct first step to confirm the model snapshot is intact and the binary can locate the device. The --help flag prints the complete CLI reference.

Without a -p prompt, the same binary starts an interactive session:

```sh
./h3 -d ./MiniMax-H3 --width 512 --height 512 --steps 6
```

The interactive session keeps the BF16 prompt conditioning, the prepared DiT model, and the video decoder in memory between generations. Repeating a prompt with a different seed avoids reloading those components. The session recognises commands prefixed with ! including !status, !seed random, !seconds 2, !show, !save output.mp4, and !cache.

## Generating a First Video with the Balanced Preset

The README recommends starting with a balanced preset that uses 20 denoising steps, 45 of the 50 transformer blocks, and a reuse factor of 2:

```sh
./h3 --profile \
  -d ./MiniMax-H3 \
  -p "A red fox walks through fresh snow in a pine forest. Medium tracking shot, natural winter light, realistic fur, soft footsteps and wind." \
  --width 512 --height 512 \
  --frames 22 --steps 20 \
  --layers 45 --reuse 2 \
  --show \
  -o outputs/fox-fast.mp4
```

The --layers 45 flag skips 5 of the 50 transformer blocks, reducing both generation time and unified-memory use. The --reuse 2 flag computes 11 fresh denoiser velocities instead of all 20 and extrapolates the rest, halving the number of full model passes. The --show flag loads a resident preview VAE and displays one middle-video frame after each Euler transition in terminals that support Kitty, Ghostty, iTerm2, WezTerm, or Konsole graphical protocols; the README notes this adds roughly 10 gigabytes of temporary model residency.

For the fastest iteration, the README suggests four denoising steps with --reuse 1:

```sh
./h3 --profile \
  -d ./MiniMax-H3 \
  -p "A red fox walks through fresh snow in a pine forest." \
  --width 512 --height 512 --frames 22 \
  --steps 4 --layers 50 --reuse 1 \
  --show \
  -o outputs/fox-four-step.mp4
```

The README states that the selected four-pass result measured 0.556 full-video SSIM against a 29-pass reference, and that a four-pass denoise took about 3.5 seconds on M5 Max versus 26.4 seconds for the reference run.

## SSD Streaming: Running with 2 GB Instead of 36 GB

The most significant memory-reduction option is --ssd-streaming, which keeps only two DiT blocks in unified memory at a time and reads the next block from SSD while the GPU runs the current one:

```sh
./h3 --profile \
  -d ./MiniMax-H3 \
  -p "A red fox walks through fresh snow in a pine forest." \
  --width 512 --height 512 --frames 22 --steps 20 \
  --layers 50 --reuse 1 --ssd-streaming \
  -o outputs/fox-ssd.mp4
```

The README reports that tracked DiT storage fell from about 36.5 gigabytes to 2.0 gigabytes at 512-square resolution and to 2.1 gigabytes at 864x480. The throughput cost is significant: a warm 50-block forward pass measured 1.35 seconds at full residency versus 2.49 seconds with SSD streaming at 512-square (84% slower), and 2.14 seconds versus 2.68 seconds at 864x480 (26% slower). The README states the results were byte-identical between the two paths.

The 2.0 to 2.1 gigabyte figure is the tracked tensor storage for the DiT only. Prompt encoding and the two VAEs run in separate phases; the OS, media buffers, and output resolution add additional overhead. Using --show adds roughly 10 gigabytes for the preview VAE, so the README advises omitting it for the lowest-memory run. SSD streaming cannot be combined with --use-int8-row-fc2. In an interactive session it is enabled with !ssd-streaming on.

## First/Last-Frame Conditioning and Ref2VA References

The interactive session supports persistent first and last frame conditioning for constrained video generation:

```
h3> !first opening.png
h3> !last ending.png
h3> The camera moves slowly around the subject.
```

The !first and !last anchors stay active for subsequent prompts in the same session. They are cleared with !first clear or !last clear.

For general multimodal conditioning, !ref-image PATH appends an image to an ordered reference list. The model receives the images as Picture 1, Picture 2, and so on; filenames have no meaning to the model. !refs lists the current order, !ref-remove N removes one entry by number, and !refs clear removes them all. The README states that Ref2VA references cannot be mixed with !first and !last anchors in the same generation.

These conditioning modes work end-to-end according to the README. The current development focus is incremental H3-specific Metal performance and memory optimisation on M3 Max and M5 Max.

## Repository Layout and the C Source Structure

The entire project is a flat collection of C and Objective-C source files with no subdirectories except tests/. The core library object files are h3.c, h3_host.c, h3_safetensors.c, h3_weights.c, h3_text_encoder.c, h3_dit_schedule.c, h3_dit.c, h3_video_vae.c, h3_video_encoder.c, h3_audio_vae.c, h3_ffmpeg.c, h3_terminal.c, h3_vision_encoder.c, and h3_multimodal.c. The Metal-specific Objective-C files are h3_metal.m, h3_gpu.m, and h3_tokenizer.m. The shaders are in h3_shaders.metal.

The CLI entry point is main.c with h3_cli.c providing the command-line parser. The Makefile produces both an h3 executable and a libh3.a static library, and it defines separate test targets for each subsystem: the tokenizer, the BF16 layer, the text encoder, the audio GPU path, the audio VAE, the video encoder, the vision encoder, and the AV muxer. Third-party linenoise.c provides the interactive session readline functionality.

The licence is MIT. Third-party notices are in THIRD_PARTY_NOTICES.md.

## Limitations and What h3.c Does Not Cover

h3.c runs only on Apple Silicon. There is no support for NVIDIA CUDA, AMD ROCm, or Intel Arc as documented in the README. Developers who need to run MiniMax-H3 on a Linux GPU workstation must use a different runtime, such as the official Python-based inference code from MiniMax.

The project has no GitHub releases and no versioned binary distributions. Every installation is a build from source at the current commit. The README describes the project as a work in progress built as a sequence of vertical slices, and the current focus is performance and memory optimisation rather than API stability.

The model weights are not included and must be obtained separately from Hugging Face. The README does not document the download size of the MiniMax-H3 snapshot, the VRAM or unified memory required beyond what the SSD streaming mode documents for the DiT layer, or the disk space needed for the full model checkpoint.

Generation quality degrades at very low step counts. The README states that steps from 4 to 7 progressively improve detail and motion, and it documents specific failure modes of aggressive tail-heavy denoising schedules that were evaluated and rejected: they produced woven texture, weak motion, or clipped colours.

## Conclusion

h3.c is the right tool for developers and researchers with Apple Silicon machines who want to run MiniMax-H3 video generation locally without a Python environment or cloud dependency. The SSD streaming mode makes it practical on machines with limited unified memory, at the cost of throughput: the README reports an 84% slowdown at 512-square resolution compared to full residency. Anyone who needs higher throughput than the hardware can sustain, or who needs a multi-platform deployment, should use a cloud-based inference endpoint instead. The project has no GitHub releases and no versioned binary packages; you build from source at the current commit. The last push was on 2026-08-11.

## FAQ

### Does h3.c work on Intel Macs or other platforms?

The README does not document support for Intel Macs, Linux, or Windows. It describes development and performance measurements on M3 Max and M5 Max. The build system links against Metal, MetalPerformanceShaders, and Accelerate, all of which are Apple-specific frameworks.

### How much GPU memory does h3.c require?

At full residency the DiT layer uses around 36.5 gigabytes of unified memory. With --ssd-streaming, tracked DiT storage drops to about 2.0 gigabytes at 512-square and 2.1 gigabytes at 864x480. The README notes that prompt encoding, VAEs, and OS overhead add to those figures, and that --show adds roughly 10 gigabytes for the preview VAE.

### Can h3.c generate audio as well as video?

Yes. The README states that prompt-to-video/audio and first/last-frame conditioning work end to end. The source tree includes dedicated h3_audio_vae.c and h3_video_vae.c files. The CLI can produce combined video-plus-audio output as an MP4 file, with FFmpeg handling the mux.

## Sources

- [antirez/h3.c on GitHub](https://github.com/antirez/h3.c)
- [Issues](https://github.com/antirez/h3.c/issues)
- [License: MIT](https://github.com/antirez/h3.c/blob/main/LICENSE)
- [README](https://github.com/antirez/h3.c/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/antirez-h3-c
