Framework
aigc-apps/VideoX-Fun avatar
aigc-apps/VideoX-Fun

VideoX-Fun's Docker recipe runs with host networking and no seccomp profile

πŸ“Ή A more flexible framework that can generate videos at any resolution and creates videos from images.

2,282 stars194 forksPythonApache-2.0

At a glance

What is it?
A video and image generation pipeline for diffusion transformer models, with training support, three inference front ends of different granularity, and three GPU offload modes. The installation instructions are worth reading closely, because the container flags and the two weight directory layouts decide how much trouble your first run is.
Who is it for?
VideoX-Fun fits someone with a consumer NVIDIA card and 60 GB of disk who wants to run video diffusion models and fine tune them, rather than call a hosted endpoint. Five things to check before you start.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 7 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The container runs with host networking and an unconfined seccomp profile

The Docker section is three commands, and the run line is the interesting one because of what it disables.

code
# pull image
docker pull mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:cogvideox_fun

# enter image
docker run -it -p 7860:7860 --network host --gpus all --security-opt seccomp:unconfined --shm-size 200g mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:cogvideox_fun

Four flags in that line are worth naming individually, because each removes a boundary.

--network host puts the container directly on the host's network interfaces instead of a bridge, so every port the process listens on is a host port and there is no network namespace between it and the machine. --security-opt seccomp:unconfined disables the syscall filter Docker normally applies. --shm-size 200g raises shared memory far above the default, which video pipelines need for large frame batches. And --gpus all exposes every GPU rather than one.

None of those is unusual for a research container that has to reach CUDA devices, and the combined effect is predictable: this container has the host's network, no syscall filtering, and all of the GPUs. It is a development image and the page says so, requiring that you have already installed the driver and CUDA environment correctly on the host first.

Two more things to note. The image comes from a personal looking Aliyun container registry rather than a vendor namespace, and the path segment is easycv with a torch_cuda base, so this is a derivative image rather than an official CUDA one. And the repository itself is cloned inside the container after it starts, rather than being mounted, which means the code you run is whatever the default branch contained at the moment you ran the command.

There is also a cloud path that skips all of this. The page points at Aliyun's notebook service, which offers free GPU time applied once per user and valid for three months, and claims you can start CogVideoX-Fun within five minutes from it.

Sixty gigabytes of weights and two directory layouts to get right

The page asks for about 60 GB of free disk for saving weights before anything else happens, and then explains where they have to go, which is where most first-run failures live.

There are two documented layouts and they are not interchangeable. If you are running through ComfyUI, the models go into that project's weights folder under a dedicated subfolder:

code
πŸ“¦ ComfyUI/
β”œβ”€β”€ πŸ“‚ models/
β”‚   └── πŸ“‚ Fun_Models/
β”‚       β”œβ”€β”€ πŸ“‚ CogVideoX-Fun-V1.1-2b-InP/
β”‚       β”œβ”€β”€ πŸ“‚ CogVideoX-Fun-V1.1-5b-InP/
β”‚       β”œβ”€β”€ πŸ“‚ Wan2.1-Fun-14B-InP
β”‚       └── πŸ“‚ Wan2.1-Fun-1.3B-InP/

If you are running the Python files or the project's own interface, they go elsewhere:

code
πŸ“¦ models/
β”œβ”€β”€ πŸ“‚ Diffusion_Transformer/
β”‚   β”œβ”€β”€ πŸ“‚ CogVideoX-Fun-V1.1-2b-InP/
β”‚   β”œβ”€β”€ πŸ“‚ CogVideoX-Fun-V1.1-5b-InP/
β”‚   β”œβ”€β”€ πŸ“‚ Wan2.1-Fun-14B-InP
β”‚   └── πŸ“‚ Wan2.1-Fun-1.3B-InP/
β”œβ”€β”€ πŸ“‚ Personalized_Model/
β”‚   └── your trained trainformer model / your trained lora model (for UI load)

Three observations. First, the same four model names appear in both, so choosing a model is separate from choosing a front end. Second, the base weights directory and the trained models directory are different, and a fine tune you train yourself is loaded from the second one by the interface. Third, and this is the trap, the four names are not consistent with themselves: three end in a trailing slash in the listing and one does not, which is a small sign that the folder names were assembled by hand at different times rather than generated.

The downloads come either from a Hugging Face link or from ModelScope, which matters in regions where one of the two is unreliable.

The verified hardware list is short enough to be useful. On Windows, Python 3.10 and 3.11 with PyTorch 2.2.0 on CUDA 11.8 or 12.1 and cuDNN 8 or newer, tested on a 12 GB 3060 and a 24 GB 3090. On Linux, Ubuntu 20.04 or CentOS with the same software, tested from a 16 GB V100 up to an 80 GB A100. The manifest asks for PyTorch 2.1.2 or newer, so the verified version is above the floor rather than at it.

The web UI exposes fewer parameters than the Python file

There are three ways in, and the project publishes the trade between them rather than pretending they are equivalent.

The table names each entry, the scenario it suits, and its configuration granularity. A Python file is for batch generation and parameter debugging, with full parameters. The web interface is for interactive experience, with common parameters only. ComfyUI is for people who already have a workflow, with node parameters.

So the interface you choose determines which knobs you can turn. The Python file exposes everything; the web interface exposes a subset described as the common parameters; ComfyUI exposes whatever its node schema defines. There is no claim that a generation produced in the web interface can be reproduced exactly from its settings, and given the difference in granularity that is probably true.

That is a reasonable design and worth understanding as a consequence rather than a criticism. Parameter count and usability trade against each other, and the honest way to present three interfaces over one engine is to say which one has all of them.

The common thread across the three is that video and image models share the same inference entry point, provided by scripts or interfaces under a directory named after the model. So adding a model means adding a directory, not writing a new service.

The tree shows the same three shapes. There is a directory for the ComfyUI integration, a configuration directory, a scripts directory, the package itself, and a models directory at the root. There is also a file named Dockerfile.ds rather than a plain Dockerfile, so the container recipe in the readme and the build file in the tree are not necessarily the same artefact, and a root level init file suggests the repository doubles as something importable from a checkout.

You configure a generation by editing the script, not by passing flags

The inference walkthrough is built around modifying files, and that is the single most consequential thing about how this project is used.

For text to video, you modify four values in one script: the prompt, the negative prompt, the guidance scale, and the seed. Then you run the script and wait. The output lands in a directory named for the model and the task.

Image to video adds two more, an image start and an image end, where the start is the first frame of the video and the end is the last. Video to video replaces the start with a reference video and keeps the end image. So each task type has its own script, its own parameter set, and its own output folder, with the image and video results going to separate directories.

Three things follow. There is no command line interface in the documented workflow, so a parameter sweep means editing a file repeatedly or writing your own loop, which is what the Python-file entry is for. A configuration is not recorded anywhere except in the file you edited, so unless you version the example directory you have no record of what produced a given video except the seed. And because seed, guidance scale and the two prompts are the documented surface, anything else the model exposes is undocumented here.

The same shape applies to the other models: different models are distinguished by directory name under the examples folder, and the page says their supported features vary, so you pick the directory that matches your weights and your goal rather than one universal script.

For anyone automating this, the practical consequence is that the reproducible unit is a script file plus a seed, not a command line, and you will want to capture both.

Three offload modes, and the page recommends the mildest one

The memory section exists because the models are large, and it is one of the more clearly written parts of the readme.

The reason given is stated plainly: one of the model families has a very large number of parameters, and memory optimisation strategies are needed to fit it on consumer grade cards. The setting is a memory mode in each prediction file, with three choices.

model_cpu_offload moves the entire model to the CPU after use, which saves some GPU memory. model_cpu_offload_and_qfloat8 does the same and additionally quantises the transformer model to float8, saving more. sequential_cpu_offload moves each layer to the CPU after use, which is slower but saves a significant amount.

The guidance at the end is the part to act on: the float8 quantisation may slightly reduce model performance but saves more memory, and if you have sufficient GPU memory the page recommends the plain offload mode. That is the opposite of the instinct when reading a list of memory options, and it is the right recommendation, because sequential offload trades a large amount of wall clock time for VRAM you may not need.

Two operational details follow from the design. The mode is set per prediction file, which means it is part of the file you edit, not a runtime flag, so it inherits the reproducibility problem of the previous section. And the setting is described as also applicable to the other model family, so it is a general mechanism rather than a per model hack.

Twenty five example families, one of them documented

The examples directory is much larger than the walkthrough, and the naming pattern inside it explains how the project stays current.

There are twenty five subdirectories under examples, and they are named for model families rather than for tasks. Some names appear twice in different forms: a base name, a numbered follow up, and a variant with a suffix. The video families include a CogVideoX variant, two LTX variants at different versions, a Hunyuan Video directory, a Long Cat Video directory, and several others. The image families are the same story with Flux, Flux 2, Qwen Image and Qwen Image 2.1, plus instant and fun variants.

The suffix is the interesting part. Several directories exist in pairs, one plain and one carrying an extra word, which by the naming convention reads as the upstream model and this project's own fine tune of it. So the repository is not only a runner for other people's checkpoints; it also publishes its own tuned variants, which is why the weights layout documentation includes those directories among the four names.

Against that, the visible quick start walks through exactly one family, the CogVideoX one, with three scripts. Everything else has a directory and no prose in the part of the page visible here.

That ratio is worth holding in mind when judging the project. Twenty five families in the tree is a maintenance surface, and each one is a place where an upstream release can change behaviour. The three readmes in the root, in English, Chinese and Japanese, suggest the audience is broad; the section list in the table of contents also promises supported models, video works, references, a citation section, and a section on limitations and risks, which is where a project like this would put the honest caveats.

The manifest and the requirements file list different dependencies

The packaging has two dependency lists and they do not match, which is a small thing with practical consequences.

The manifest lists Pillow, einops, safetensors, the vision library, a token merging library, PyTorch with a floor, several differential equation libraries, a video decoding library, the datasets library, numpy, scikit-image, OpenCV, a configuration library, a sentence piece library, an augmentation library, two imageio variants, tensorboard, beautifulsoup4, ftfy, a timeout helper, accelerate, gradio, diffusers and transformers.

The requirements file lists the same set plus two entries the manifest omits: an audio library and an ONNX runtime. So someone who installs from the manifest gets a smaller dependency set than someone who installs from the requirements file, and the audio and ONNX pieces are in the second group only.

The manifest also has no build system table, so it relies on the packaging tool's default, and it declares a single package directory rather than the repository's full contents. There is also a repository URL written in a non-canonical form with mixed case, which GitHub resolves by redirect but which will not match a literal string comparison.

Two entries are about ComfyUI rather than about Python. A tool section names the project as the display name in the Comfy registry and gives a publisher identifier, with a comment saying the section exists for the registry. So the package is published to a node registry under a third party publisher, which is how a ComfyUI user installs it without pip.

The rest of the tree is training and data oriented: a datasets directory, a configuration directory, a reports directory, an install script, a scripts directory, and a skills directory at the root. An install script plus a requirements file plus a manifest is three ways to set up the same environment, and which one a given instruction refers to matters.

Editorial conclusion

VideoX-Fun fits someone with a consumer NVIDIA card and 60 GB of disk who wants to run video diffusion models and fine tune them, rather than call a hosted endpoint. Five things to check before you start. Disk and memory, because the page asks for about 60 GB for weights and the offload modes exist because these models do not fit on a 24 GB card without help. Which front end you use, since the web UI deliberately exposes fewer parameters than the Python files and a generation you cannot reproduce from the UI is a generation you cannot reproduce. That your configuration lives in the script, because the documented workflow is to edit a file and run it rather than to pass flags. Whether you take the Docker path at all, since its flags trade isolation for convenience. And which of the 28 example families you actually want, because the quick start only walks through one of them.

Frequently asked questions

What is VideoX-Fun?

It is an Apache licensed video generation pipeline written in Python that generates AI images and videos, and can train baseline and LoRA models for diffusion transformers. It supports direct prediction from pre-trained baseline models at different resolutions, durations and frame rates, plus training your own models for style transformation.

How do I configure a VideoX-Fun generation?

By editing the example script rather than by passing flags. For text to video you modify the prompt, negative prompt, guidance scale and seed inside predict_t2v.py and then run that file; image to video adds start and end images, and video to video adds a reference video. Each task has its own script and its own output folder under samples.

What are the VideoX-Fun GPU offload modes?

Three, set through a memory mode in each prediction file: model_cpu_offload moves the whole model to CPU after use, model_cpu_offload_and_qfloat8 does that and also quantises the transformer to float8, and sequential_cpu_offload moves each layer after use, which is slower and saves considerably more. If you have enough VRAM the page recommends the plain offload mode.

Where do VideoX-Fun model weights go?

It depends on the front end. For ComfyUI they go under the project's models folder in a dedicated Fun Models subdirectory. For the Python files or the project's own interface they go under a Diffusion Transformer directory, with anything you trained yourself in a separate personalized model directory so the interface can load it. Downloads come from either Hugging Face or ModelScope, and the page asks for about 60 GB of disk.

Official sources

  1. aigc-apps/VideoX-Fun on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/aigc-apps-videox-fun.svg)](https://hysenlabs.com/projects/aigc-apps-videox-fun)