ComfyUI-MultiGPU moves model weights off your compute card, and the repository is archived with its own notice still in place
This custom_node for ComfyUI adds one-click "Virtual VRAM" for any UNet and CLIP loader as well MultiGPU integration in WanVideoWrapper, managing the offload/Block Swap of layers to DRAM *or* VRAM to maximize the latent space of your card. Also includes nodes for directly loading entire components (UNet, CLIP, VAE) onto the device you choose
At a glance
- What is it?
- ComfyUI-MultiGPU is a ComfyUI custom node that adds Virtual VRAM to UNet, CLIP, and VAE loaders and layers MultiGPU placement into WanVideoWrapper. Its DisTorch modes split a model across devices by bytes, ratio, or fraction, expressed in one text string. The repository is archived, so nothing here will be fixed.
- Who is it for?
- ComfyUI-MultiGPU is worth reading if you already run ComfyUI and are hitting VRAM ceilings on long videos, because the allocation grammar is documented with real examples and the node clones keep the original loader behaviour. It is not a place to build on: the repository is archived, the maintainer states there will be no further code updates or support, and no fork is endorsed.
- Can I use it commercially?
- Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- No. The owners have archived the repository on GitHub, so it is read-only and no longer receives changes.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The archive notice sets a date and refuses a successor
The repository metadata already records this repository as archived. The pinned notice at the top of the README sets the date as 30 September 2026 and states that there will be no further code updates or support from the maintainer. The last commit on the default branch landed on 2026-09-20. There are also no GitHub releases, so the only version marker anywhere is `version = "2.6.4"` in pyproject.toml. Fork maintainers are invited to announce themselves in the pinned coordination issue, issue 223, while the same notice says the repository is not being transferred and that no fork is endorsed. Those two sentences sit together on purpose: continuity is expected to come from forks, and none of them is named as authoritative. The licence is GPL-3.0, declared in pyproject.toml as a file reference rather than as an SPDX string.
DisTorch moves weights and does not run steps in parallel
The footnote under the core description is the most important paragraph in the file. It says this enhances memory management, not parallel processing. Workflow steps still execute sequentially, with components, in full or in part, loaded across the specified devices. Performance gains come from avoiding repeated model loading and unloading when VRAM is constrained. Capability gains come from offloading as much of the model, meaning VAE, CLIP, and UNet, off the main compute device as possible, so that latent space is left for actual computation. DisTorch stands for distributed torch, and the donor device is either the CPU's DRAM or another cuda device's VRAM. The repository description calls the same mechanism a block swap of layers, so two different names describe one mechanism in this project's own text.
One text string expresses three different allocation units
Normal mode is the simple case: a `virtual_vram_gb` slider picks one donor device and the more you add, the more of the model moves there. Expert mode replaces the slider with a single flexible text string, and three units can go in that string. Bytes names absolute sizes per device with `*` as the wildcard for the remainder, and the CPU is the default wildcard if none is given. Ratio names percentages, which the documentation compares to llama.cpp's `tensor_split`. Fraction is the original expert mode and is a different measurement: it splits by the fraction of each device's total VRAM rather than by the size of the model. Three examples make the grammar concrete:
cuda:0,2.5gb;cpu,*
cuda:0,25%;cpu,75%
cuda:0,0.1;cpu,0.5The first puts the first 2.50GB of the model on `cuda:0` and the rest on the CPU. The second splits the model one to three. The third uses 10% of `cuda:0`'s VRAM and 50% of the CPU's RAM.
The wildcard works as a suffix, not only as a whole term
The bytes examples show the grammar has more than one shape, and the second one is the interesting case. In `cuda:0,2.5gb;cpu,*` the wildcard stands alone as the CPU's entire allocation. In `cuda:0,500mb;cuda:1,3.0g;cpu,5gb*` the wildcard is attached to a size, and the documentation describes the result as 5.00GB or the remainder on the CPU. So `*` can be a whole term or a modifier on a term, and the difference decides whether a device gets a floor or a cap. The ratio mode has no wildcard at all in its examples, and the percentages have to add up on their own. The three-device ratio example, `cuda:0,8%;cuda:1,8%;cpu,4%`, is described as an 8:8:4 ratio resolving to 40%, 40%, and 20%, which is what dividing each term by twenty gives.
The flag name, the caption, and the device family are misspelled
Four spelling problems are baked into the documentation, and each one costs a search. The expert allocation flag is written `expert_mode_alloaction`, with the two a's transposed, and it is the only name given for the setting that expert users are pointed at. A caption describing the DisTorch nodes reads Vitual VRAM rather than Virtual VRAM, while the same feature is spelled correctly everywhere else. The two new model-driven Expert Modes are introduced as inutuitive. And the donor devices are described as a cuda/xps device, where xps does not name a CUDA device family. None of this changes behaviour, but if you are grepping the workflow JSON or the node code for the setting name, the misspelling is the string that will match.
DynamicVRAM survives only on devices comfy-aimdo initialised
Compatibility is claimed for all `.safetensors` and GGUF-quantized models, with a paragraph that gets specific about recent ComfyUI builds. On builds with DynamicVRAM and comfy-aimdo enabled, MultiGPU keeps DynamicVRAM active on the CUDA devices that comfy-aimdo has initialised, and falls back to legacy model patching for off-grid MultiGPU CUDA devices. The stated effect is that MultiGPU placement is preserved for devices such as `cuda:1` even when comfy-aimdo has only initialised the primary device. So there are two code paths in play on a multi-GPU box with a recent ComfyUI, and which one you get depends on whether your device index is one the other layer touched. That paragraph is the one to read first if you see placement ignored on a second GPU.
Loader nodes are cloned and given one extra parameter
The extension works by generating MultiGPU versions of existing loader nodes rather than by adding a new place to load models. Each MultiGPU node keeps the functionality of its original counterpart and adds a `device` parameter for picking the GPU. The nodes listed are automatically detected if available, and the naming carries both tracks: `CheckpointLoaderAdvancedMultiGPU` sits next to `CheckpointLoaderAdvancedDisTorch2MultiGPU`, and `CheckpointLoaderSimpleMultiGPU` next to `CheckpointLoaderSimpleDisTorch2MultiGPU`, followed by `UNETLoaderMultiGPU` and its DisTorch2 counterpart. Per-node documentation lives under `web/docs/`, one file per node, and the tree also carries `example_workflows/` and a `ci/` directory.
Install through the manager, and there is no dependency list
Installation through ComfyUI-Manager is the preferred route: search for `ComfyUI-MultiGPU` in the node list and follow the instructions there. The manual route is to clone the repository inside `ComfyUI/custom_nodes/`. Neither route tells you what to install first, because the pyproject.toml `[project]` table has no `dependencies` key at all, only a name, a description, a version, and a licence file reference. For a custom node that patches ComfyUI's model management and reads GGUF weights, that omission is worth checking against your own ComfyUI environment. The same file does pin down its own tooling, including a ruff rule set that selects S307 for suspicious eval usage and S102 for exec, alongside the usual E, W, F, and T series, with E501, E722, E731, E712, E402, and E741 ignored.
Editorial conclusion
ComfyUI-MultiGPU is worth reading if you already run ComfyUI and are hitting VRAM ceilings on long videos, because the allocation grammar is documented with real examples and the node clones keep the original loader behaviour. It is not a place to build on: the repository is archived, the maintainer states there will be no further code updates or support, and no fork is endorsed. If you install it, pin version 2.6.4, read the DynamicVRAM compatibility paragraph before combining it with comfy-aimdo builds, and check the pinned coordination issue before assuming a fork will be maintained.
Frequently asked questions
What does ComfyUI-MultiGPU actually do?
It is a ComfyUI custom node that adds one-click Virtual VRAM to UNet and CLIP loaders and MultiGPU integration inside WanVideoWrapper, moving model layers to DRAM or to another device's VRAM. Its DisTorch nodes shift the static parts of the UNet off the main compute card onto a slower donor device, and there are also nodes for loading a whole UNet, CLIP, or VAE onto a device you choose.
Does ComfyUI-MultiGPU run ComfyUI workflow steps in parallel on several GPUs?
No. The documentation says it enhances memory management, not parallel processing, and that workflow steps still execute sequentially with components loaded across the specified devices. The performance gain comes from avoiding repeated model loading and unloading when VRAM is constrained, and the capability gain comes from freeing latent space by offloading.
How do I install ComfyUI-MultiGPU into ComfyUI?
Through ComfyUI-Manager, which is described as the preferred route: search for `ComfyUI-MultiGPU` in the node list and follow the installation instructions. The manual route is to clone the repository inside the `ComfyUI/custom_nodes/` directory.
What is the difference between the DisTorch bytes, ratio, and fraction modes?
Bytes names exact gigabytes or megabytes per device and uses the wildcard `*` for the remainder, with the CPU as the default wildcard. Ratio distributes the model by percentage, in the spirit of llama.cpp's `tensor_split`. Fraction is the original expert mode and splits by the fraction of each device's total VRAM instead of by the size of the model. All three are written into a single flexible text string.
Is ComfyUI-MultiGPU still maintained?
No. The repository is archived, and the pinned notice states it will be archived on 30 September 2026, that there will be no further code updates or support from the maintainer, and that fork maintainers can announce themselves in the pinned coordination issue. The same notice says the repository is not being transferred and no fork is endorsed. There are no GitHub releases; the version is 2.6.4 in pyproject.toml.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/pollockjj-comfyui-multigpu)