Open-source project
FlashML-org/FreeVideo avatar
FlashML-org/FreeVideo

FreeVideo: local MiniMax H3 video generation on consumer GPUs with 8 GB VRAM

Make videos on the computer you already own. FreeVideo runs MiniMax H3 in as little as 8 GB of VRAM and 16 GB of RAM, and adapts its acceleration path to your hardware.

1,077 stars100 forksPythonApache-2.0

At a glance

What is it?
FreeVideo is a local inference engine for MiniMax H3 that adapts its memory strategy and compute path to the GPU you already have, running with as little as 8 GB of VRAM and 16 GB of RAM. It integrates into ComfyUI and supports Windows launchers, a macOS Apple silicon preview, and a Linux CLI.
Who is it for?
FreeVideo suits a developer or creative who wants to generate AI video locally without cloud costs, has a GPU with at least 8 GB of VRAM, and is comfortable staying on Python 3.12. Skip it if you need 3.13 compatibility or a production-grade deployment without a macOS notarization step.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 6, 2026, and from our analysis. They are not legal advice.

Editorial analysis

FreeVideo runs MiniMax H3 locally by coordinating VRAM, system memory and disk

FreeVideo addresses the cost and latency of cloud-based AI video generation by running the MiniMax H3 model on hardware you already own. The minimum hardware floor is 8 GB of VRAM and 16 GB of RAM, achieved through weight streaming, asynchronous prefetching and chunked computation that keep peak memory low.

The model behind it is VDN-H3, an 8-step checkpoint from OpenVDN built on Video DeltaNet's hybrid attention architecture. FreeVideo does not expose a raw model loader; it wraps that model in an inference engine that coordinates placement across VRAM, system RAM and disk depending on what the GPU has available.

Hardware adaptation goes deeper than memory management. FreeVideo chooses the FP8 compute path based on GPU architecture: either native FP8 or FP8 storage with BF16 compute. It also probes the available attention kernels automatically at startup rather than requiring the user to pick one.

Input types include text prompts, first and last frame images, and image, video and audio references. MiniMax H3 LoRAs are supported through ComfyUI's node view, and the documentation points to a docs/LoRA.md file with examples. ComfyUI is the primary interface for all of these input paths, with a dedicated creative workspace that includes a history of past creations, two-pass sampling and batch generation.

Python 3.12 is the only supported version; 3.13 is explicitly excluded by pyproject.toml

A constraint that does not appear in the README headline but is visible in pyproject.toml is the Python version ceiling. The requires-python field reads >=3.12,<3.13, which means exactly Python 3.12. Python 3.13 is excluded by the upper bound, and the lower bound rejects earlier versions.

This matters for anyone who manages multiple Python projects. An environment that standardizes on 3.13 for other tools cannot run FreeVideo in the same interpreter. A virtual environment for exactly 3.12 is required, and the setup scripts on Linux and Windows assume this.

The package name on PyPI is freevideo-engine, not freevideo. The internal version in pyproject.toml is 2026.9.16.41444, a date-based build identifier, while the GitHub repository uses semantic-style tags (v0.2.4 is the current tag). Those are two different version surfaces that do not match each other, which can create confusion when checking what is installed.

Runtime dependencies are pinned tightly: transformers==5.15.0, accelerate==1.14.0, peft==0.20.0, triton==3.7.1 on Linux and triton-windows==3.7.1.post27 on Windows. That level of pinning avoids breakage from upstream changes but means other tools in the same environment may conflict if they pull different versions of transformers or accelerate.

Windows and macOS launchers handle the full setup; Linux uses two shell commands

Installation splits into three paths depending on the operating system.

For Windows, download FreeVideo.exe from GitHub and run it. The launcher asks for an existing ComfyUI folder or offers to install a new one, reuses model folders you already have, downloads any missing models automatically, and then opens ComfyUI in the browser. An offline path exists for machines without internet access: download the packages from Quark (the README names pan.quark.cn) and drag the ZIP files into the launcher without extracting them. The common models and the model pack for your GPU (30/40 series or 50 series) are required, plus the environment package for a fresh ComfyUI install.

For macOS with Apple silicon, download FreeVideo-Mac-arm64.dmg, open it, drag FreeVideo.app into Applications, and follow the same install-and-launch flow. The macOS path is discussed in its own section below.

For Linux or an existing ComfyUI installation on any platform, two commands cover the setup:

bash
git clone https://github.com/FlashML-org/FreeVideo.git && cd FreeVideo
./setup.sh

Generating a video on Linux uses the freevideo CLI:

bash
./freevideo generate --prompt-file prompt.txt --out video.mp4

For an existing ComfyUI installation, clone into the custom_nodes directory instead:

bash
cd ComfyUI/custom_nodes
git clone https://github.com/FlashML-org/FreeVideo.git

After restarting ComfyUI, open Workflow, then Browse Templates, select FreeVideo, and open FreeVideo-All-in-One to complete the setup in FreeVideo Settings.

Four quality levels set the trade-off between generation time and output quality

FreeVideo offers four quality levels for each video: Light, Medium, High and Max. Higher levels produce higher-quality output but take longer to generate. The choice is made per video rather than as a global setting, so a workflow can mix levels depending on the stage of the project.

The README and news section do not publish specific generation times for each level in the main README text, but the Mac guide at docs/Mac.md is named as the place to find generation times and memory figures for Apple silicon. The gallery at freevideo-community.pages.dev shows the four quality levels side by side for comparison.

Results from any quality level can be exported as sharing images or as videos that embed the generation time and GPU. That export format appears alongside the workflow embedding added in v0.2.3 and is separate from it.

Two-pass sampling is available inside the ComfyUI workspace for finer control over the output. The node view in ComfyUI gives access to the full workflow, where LoRAs and custom steps can be added. This two-surface design means casual users get the creative workspace while pipeline builders get the node graph, both running the same underlying engine.

v0.2.3 embeds the prompt, seed and settings into each video file

One specific feature from v0.2.3 is worth understanding before committing to a workflow: FreeVideo video files carry the prompt, seed and generation settings as embedded metadata. Dropping a FreeVideo video back onto the ComfyUI canvas restores those settings automatically, which means a video file is also a workflow checkpoint.

The news entry for v0.2.3 also notes that earlier videos can receive this metadata retroactively, though the README does not detail the mechanism for doing that.

This only works when the receiving ComfyUI instance has FreeVideo installed and knows how to read the embedded data. A standard ComfyUI setup without the FreeVideo plugin will not restore the workflow from a dropped video, and a non-ComfyUI viewer will display the video normally without surfacing the metadata.

The practical consequence for any team sharing output files is that the workflow is portable only within the FreeVideo ecosystem. If you share a video with someone who lacks the plugin, the generation parameters are not recoverable from the file without FreeVideo. Check what version your collaborators are running before relying on this for parameter handoffs.

The macOS preview is not notarized, requiring a manual approval step on first launch

The macOS path is described as a preview, tested on an M5 Mac with 24 GB of unified memory. Generation times and memory figures for Apple silicon are in docs/Mac.md rather than the main README.

One step that will catch users off guard is the notarization status. FreeVideo-Mac-arm64.dmg is not notarized by Apple as of the README's current state. macOS blocks unsigned applications on the first open. The README says to download the package only from the official repository rather than third-party mirrors, then check the file and approve it as described in docs/Mac.md. That guide covers what to click to approve FreeVideo without changing other security settings.

For Windows and Linux users, this is not a concern. The issue is specific to macOS Gatekeeper's behaviour with unsigned distributable packages. Apple's notarization process is a cost and workflow overhead for small open source projects, and the README acknowledges the status rather than hiding it.

FreeVideo is distributed under Apache-2.0, and the repository includes a THIRD_PARTY_NOTICES.md and a NOTICE file at the top level that list dependency attributions.

Anyone bundling FreeVideo inside an enterprise distribution should read THIRD_PARTY_NOTICES.md carefully, since the Windows and Mac launchers bundle the runtime environment and models in a form that may complicate standard dependency review workflows.

Weight streaming and adaptive FP8 are the technical constraints to understand before deployment

FreeVideo's low-memory claim rests on three mechanisms: weight streaming, asynchronous prefetching and chunked computation. Weight streaming means not all model weights are in VRAM at once; parts are moved from system RAM or disk as needed. This reduces peak VRAM but increases generation time relative to a full-VRAM scenario, and the amount of slowdown depends on the speed difference between VRAM, system RAM and storage.

The FP8 path is hardware-dependent. FreeVideo selects either native FP8 or FP8 storage with BF16 compute based on the GPU architecture. An NVIDIA 30-series or 40-series GPU may take one path and a 50-series a different one. The execution planner is documented in docs/execution-planning.md, which is the place to look if generation quality or speed does not match expectations.

For a direct comparison: Wan2.1 and HunyuanVideo are other open source video generation models that can run locally on consumer GPUs. Both expose their own memory management and quantization options independently of ComfyUI. FreeVideo's distinction is its tight integration with ComfyUI as the primary interface and the launcher-based setup, which reduces friction for users already in that ecosystem. A user who wants to evaluate multiple models side by side without a dedicated launcher will find the pure Python paths of those alternatives easier to swap.

Bug reports go to GitHub Issues; community questions go to the Discord, QQ or WeChat groups linked in the README.

Editorial conclusion

FreeVideo suits a developer or creative who wants to generate AI video locally without cloud costs, has a GPU with at least 8 GB of VRAM, and is comfortable staying on Python 3.12. Skip it if you need 3.13 compatibility or a production-grade deployment without a macOS notarization step. Before adopting it for any workflow that saves output files, confirm that your ComfyUI version accepts the workflow metadata embedded in video files added in v0.2.3, since older installations will not restore the prompt and settings automatically.

Frequently asked questions

What is FreeVideo?

FreeVideo is a local inference engine for MiniMax H3 video generation, built on OpenVDN's VDN-H3 model. It runs as a ComfyUI plugin and adapts its memory and compute strategy to the GPU available, with a minimum of 8 GB of VRAM and 16 GB of RAM.

What GPU do I need to run FreeVideo?

FreeVideo runs with as little as 8 GB of VRAM and 16 GB of RAM through weight streaming and chunked computation. The FP8 compute path is selected automatically based on GPU architecture, with different paths for 30/40-series and 50-series NVIDIA GPUs.

Does FreeVideo work on macOS?

FreeVideo offers an Apple silicon preview, tested on an M5 Mac with 24 GB of unified memory. The macOS package is not notarized by Apple yet, so the first launch requires a manual approval step described in docs/Mac.md.

How do I install FreeVideo on Linux?

Clone the repository, run ./setup.sh to set up the environment, then use ./freevideo generate --prompt-file prompt.txt --out video.mp4 to generate a video. For an existing ComfyUI installation, clone into the custom_nodes directory instead.

What Python version does FreeVideo require?

FreeVideo requires Python 3.12 exactly. The pyproject.toml sets requires-python to >=3.12,<3.13, so Python 3.13 and higher are excluded, and versions below 3.12 do not meet the minimum.

Official sources

  1. FlashML-org/FreeVideo on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/flashml-org-freevideo.svg)](https://hysenlabs.com/projects/flashml-org-freevideo)