Open-source project
Acly/comfyui-tooling-nodes avatar
Acly/comfyui-tooling-nodes

comfyui-tooling-nodes: Direct Image I/O for External ComfyUI Backends

Nodes for using ComfyUI as a backend for external tools. Send and receive images directly without filesystem upload/download.

674 stars87 forksPythonGPL-3.0

At a glance

What is it?
comfyui-tooling-nodes is a GPL-3.0 Python package that lets external applications send images to ComfyUI via Base64 or an in-RAM HTTP cache and receive results over WebSocket, while adding region attention masking and tile-based processing for large-image workflows.
Who is it for?
comfyui-tooling-nodes is the right choice for engineers building external tools or applications on top of ComfyUI where the default filesystem-based upload/download cycle is too slow or too fragile. The package is licensed under GPL-3.0, which requires any derivative that incorporates the code to be distributed under the same terms.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The Filesystem Problem comfyui-tooling-nodes Bypasses

ComfyUI's standard workflow exchanges images through the filesystem: an external caller uploads a file, submits a prompt that references that file, waits for execution, then downloads the result. The package's README notes that this multi-step process introduces a class of issues around file cleanup and timing that do not arise with direct data transfer. The files are written to disk even when ComfyUI is used purely as a processing backend and the results will never be displayed on its own web interface.

comfyui-tooling-nodes provides three alternative transfer paths: Base64-encoded images embedded directly in the prompt JSON, a WebSocket channel that streams output images as binary messages, and an in-RAM HTTP cache that stores images temporarily between the completion notification and the download request. All three paths eliminate the filesystem round-trip. They also ensure every message in a long conversation exports completely, in order, because the data comes from the API rather than the rendered page.

Base64 Input and WebSocket Output Nodes

The Load Image (Base64) node accepts a PNG image encoded as a base64 string embedded directly in the prompt payload. It outputs an RGB image tensor and, when the PNG has an alpha channel, a separate mask tensor. The Load Mask (Base64) node does the same for single-channel masks.

On the output side, the Send Image (WebSocket) node streams completed images to the client over the existing WebSocket connection as PNG binary data. The README specifies the wire format: the server sends one binary message per image in the batch, each consisting of two 32-bit big-endian integers with values 1 and 2 followed by the raw PNG bytes. After sending all images in the batch, the server sends a JSON notification over the same connection. The node supports RGB and RGBA images and handles batches natively, so a workflow that generates multiple images in one pass transmits all of them before the notification arrives.

The WebSocket path is well suited to interactive tools and streaming previews. For large images where bandwidth is a concern, the HTTP cache is faster.

HTTP Cache for Large Image Transfers

Load Image from Cache and Save Image to Cache provide a two-step HTTP transfer path that the README describes as faster than WebSocket for large images. When Save Image to Cache processes a batch, it stores the result images in RAM temporarily and emits a WebSocket notification containing image IDs:

json
{
  "type": "executed",
  "data": {
    "node": "<node ID>",
    "output": {
      "images": [
        {"source": "http", "id": "<image ID>", "content-type": "image/png", "type": "output"}
      ]
    },
    "prompt_id": "prompt ID"
  }
}

The external caller reads those IDs from the notification and fetches each image with a GET request to /api/etn/image/{id}. The README states the images are cached for a few minutes, so the caller must fetch them before that window closes.

For the input side, Load Image from Cache accepts an ID of an image that was previously uploaded via a PUT request to /api/etn/image/{id}. The server returns 201 on a new upload and 200 if that ID was already cached. This avoids re-embedding the same image as base64 in every prompt when the same source image is used across multiple workflow runs.

Region Attention Masking for Multi-Area Prompting

The region nodes implement attention masking so that different text prompts apply to different spatial areas of a generated image. This is distinct from conditioning masking: the README describes the method as less forceful but producing more natural image compositions.

A region list starts with a Background Region node that takes a text prompt and no mask. Its prompt applies to all image areas not covered by any subsequent mask. Each Define Region node appends a new entry to the list by taking a prompt and a mask that defines the spatial area that prompt governs. Masks must match the image size or the latent size, which is one-eighth the image resolution.

The Regions Attention Mask node patches the model to use the assembled list, replacing the standard positive text conditioning passed to the sampler. ControlNet and other conditioning sources can still be passed alongside it. A workflow file named region_attention_mask.json is included in the workflows/ directory as a working example.

The List Region Masks node outputs all masks from a region list, which is useful for debugging whether mask shapes are correct before committing to a full generation run.

Tiled Processing for Large-Resolution Outputs

The tile nodes split a large image into overlapping sub-images, process each tile independently through a diffusion workflow, and merge the results with smooth blending at tile edges. The README notes that many existing tiling implementations encode a fixed pipeline, while these nodes only handle the split and merge, leaving the per-tile workflow entirely up to the caller.

Create Tile Layout defines the parameters: min_tile_size sets the minimum resolution in pixels (tiles may be larger to fit the image evenly), padding sets the overlap with neighboring tiles in pixels, and blending controls how much of the padding area is used for smooth edge transitions. The total tile count is computed as image_size divided by the sum of min_tile_size and twice the padding.

Extract Image Tile and Extract Mask Tile split out individual tiles using a column-major index (tile 1 is typically below tile 0). Merge Image Tile reassembles a processed tile into the full image using the overlap and blending values from the layout. A workflow file named image_tiles.json is included as an example of the full pipeline.

This flexibility comes at a cost: implementing a tiling workflow is more complex than using a fixed-pipeline node, because the caller is responsible for iterating over tile indices and managing the merge order.

Installation and Optional Dependencies

The package is named comfyui-tooling-nodes and is at version 3.3.0 according to the pyproject.toml. The README has an Installation section at the repository page that describes how to add the package to a ComfyUI instance.

The only external dependency beyond ComfyUI itself is argostranslate, which is listed in requirements.txt as optional and required only for the Translate Text node:

bash
# Optional, only required for Translate node:
pip install argostranslate

The Translate Text node accepts language directives in the form lang:xx (where xx is a two-letter language code) to identify the source language. Multiple directives within a single input string are allowed and take effect at the point where each appears. lang:en passes fragments through untranslated. This node is designed for use with keyword-heavy prompts that mix multiple languages and need consistent English output to the model.

The nsfw.py module in the repository suggests a safety filter node is also present, but the README material does not document its interface.

Limitations and the GPL-3.0 License Constraint

The base64 and WebSocket transfer paths trade filesystem overhead for payload size. Embedding a large PNG as a base64 string in a prompt JSON increases the prompt size and the memory used to parse it. For very large source images, the HTTP cache upload path is more efficient because it separates the image transfer from the prompt submission.

A comparable alternative is the ComfyUI API directly, which supports WebSocket connections for the standard output path but still requires filesystem uploads for inputs. ComfyUI-tooling-nodes is specifically designed to remove that asymmetry.

The package is released under GPL-3.0, which requires any software that incorporates or links against it to be distributed under the same license terms if distributed publicly. An application that calls ComfyUI as a separate process over its API is not subject to this requirement, but a tool that packages comfyui-tooling-nodes as part of a closed binary is. This is a meaningful constraint for commercial tool developers.

Editorial conclusion

comfyui-tooling-nodes is the right choice for engineers building external tools or applications on top of ComfyUI where the default filesystem-based upload/download cycle is too slow or too fragile. The package is licensed under GPL-3.0, which requires any derivative that incorporates the code to be distributed under the same terms. Anyone evaluating it should check whether the GPL-3.0 terms are compatible with their distribution model before integrating the library into a closed product. The repository had its last push on 2026-09-26 and is not archived.

Frequently asked questions

What are nodes in ComfyUI?

Nodes are the building blocks of ComfyUI workflows: each node performs one operation (loading an image, running a sampler, saving output) and passes data to the next node through typed connections. comfyui-tooling-nodes adds specialized nodes for Base64 image input, WebSocket and HTTP cache output, region attention masking, tiled processing, and text translation.

Is there a custom node manager for ComfyUI?

ComfyUI has a community-maintained node manager that allows installing custom node packages without manual cloning. The comfyui-tooling-nodes package is available as a custom node and can be installed through that manager; the README's Installation section at the repository page describes the available installation methods.

How to make a node for ComfyUI?

A ComfyUI custom node is a Python module placed in the custom_nodes directory that defines node classes with INPUT_TYPES, RETURN_TYPES, FUNCTION, and CATEGORY attributes. comfyui-tooling-nodes shows a complete implementation: each node class in nodes.py declares its inputs and outputs, and the package registers all nodes through the __init__.py file that ComfyUI reads on startup.

Official sources

  1. Acly/comfyui-tooling-nodes on GitHub
  2. Issues
  3. License: GPL-3.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/acly-comfyui-tooling-nodes.svg)](https://hysenlabs.com/projects/acly-comfyui-tooling-nodes)