These nodes exist because the stock model drops your reference image
Comfy UI Nodes for Krea 2 LoRAs trained with AI Toolkit Experimental Edit
At a glance
- What is it?
- A two-node pack that encodes reference images the way an edit adapter was trained and patches the model to actually consume them, since the built-in path ignores reference latents. Scope is narrow by design and the caching option has a documented trap.
- Who is it for?
- This node pack is the right answer if you trained an edit adapter with the specific toolkit and settings it targets and want to run it in this interface, where the built-in path would drop your reference conditioning without saying so. It has nothing to offer anyone outside that combination, and a different editing workflow with native reference support avoids the custom nodes entirely if you have not already trained an adapter.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 75 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Two nodes that make reference images actually reach the model
This is a pair of custom nodes for a popular node-based image generation interface, written to run edit adapters trained with a particular training toolkit against a particular image model.
The problem is stated plainly in the middle of the README and it is the reason the project exists: the stock version of this model inside the interface ignores reference latents. So an adapter trained to edit an image based on a reference will load, run, and quietly not use the reference, which is the worst category of failure because it produces plausible output that is not what you asked for.
One node handles the input side, encoding the prompt together with reference images. The other patches the model so it consumes what that encoding produced. Both are needed, and the second exists purely because the stock path drops the information.
The audience is narrow to the point of being self-selecting: people training edit adapters with that specific toolkit who want to run them in that specific interface. Anyone outside that intersection has nothing to gain here.
Matching the training conditions exactly is the whole design
What makes these nodes more than glue is that they reproduce the conditions the adapter was trained under, and the README is specific about each one.
Prompt and reference images are encoded together through the model's vision-language text encoder using the conditioning template the model expects, with numbered picture placeholders described as the same layout used during training. Conditioning that differs from training is the usual reason an adapter underperforms, and matching the template removes that variable.
Image sizing is handled the same way. Images fed to the encoder are downscaled to fit a small pixel budget and, the README notes, never upscaled. Reference latents are fitted to a larger budget separately. The refusal to upscale matters: enlarging a small reference to hit a target size introduces interpolation artefacts the model never saw in training, and declining to do it is the correct choice even though it means a small reference stays small.
On the model side, each reference is appended to the image token sequence and conditioned at the first timestep, with the denoising prediction covering only the target image tokens. That last clause is the part that keeps the references as context rather than as things being generated.
Installing it, and the two failure modes to know about
Installation is a clone into the interface's custom node directory followed by a restart.
cd ComfyUI/custom_nodes
git clone https://github.com/ostris/ComfyUI-Krea2-Ostris-Edit.gitNo additional dependencies are required, which the project manifest confirms with an empty dependency list, and the nodes appear under their own category in the node menu after the restart. For an ecosystem where custom nodes routinely drag in conflicting package versions and break an installation, adding none at all is worth noting.
Two conditions will produce confusing results if you miss them, and both are documented.
The first concerns the text encoder checkpoint you load. It must include the vision weights, or the reference images simply cannot be encoded. A checkpoint missing them is a plausible thing to have on disk, and the resulting failure points at the images rather than at the checkpoint.
The second concerns the caching option on the patch node, which is off by default. It precomputes the attention keys and values for the reference tokens in a single pass at the first timestep and reuses them on every denoising step, so the references stop travelling through the per-step sequence. That is faster, particularly at high step counts. The README states in bold that the adapter must have been trained with the matching option in the training toolkit for this to work properly, and to leave it off otherwise. Enabling it against a normally trained adapter is a quiet correctness problem rather than an error.
The fallback behaviour is the detail that shows care
One sentence in the patch node's description is worth singling out. If the conditioning carries no reference latents, the patched model behaves exactly like the stock model, so the node is safe to leave in the graph.
That is a small guarantee with a real effect on how people work. Node graphs get reused, duplicated and adapted, and a patch node that broke ordinary generation when no reference was connected would force users to maintain two versions of every workflow. Degrading to the unmodified behaviour means one graph serves both cases.
It also reflects the right instinct about what a patch should do, which is to add a capability when its inputs are present and otherwise stay out of the way. The alternative, failing loudly when references are absent, would be defensible and would be worse here, because the absence of a reference is a legitimate state rather than a mistake.
The input surface is correspondingly small. The encoding node takes the text encoder, a prompt, and optionally an autoencoder and up to three reference images. The patch node takes a model and the caching flag. Three references is a fixed ceiling rather than a list, which suits a node interface where each input is a socket.
Where this does not apply
The scope conditions are the main limitation and they are severe by design.
This works with adapters trained using one toolkit, with one architecture setting, and one edit flag enabled. An adapter trained any other way is not addressed, and the caching option narrows it further to adapters trained with that additional setting. That is three separate compatibility conditions before anything runs.
The repository carries no tagged releases, with the version recorded in the project manifest at 1.0.1 and the last push on 2026-07-17. For a node pack tracking a fast-moving interface and a recent model, that combination means pinning a commit rather than tracking the default branch.
There is also an upstream dependency worth naming. These nodes exist because the stock implementation ignores reference latents. If that changes upstream, the patch node's purpose narrows to whatever the stock path still lacks, and its behaviour against a newer stock model is not something the repository can promise.
Finally, nothing here reports quality. There are no example outputs or comparisons in the README, so whether an adapter run through these nodes produces what its training implied is left for you to establish.
Using the stock nodes is the alternative, and it silently loses the reference
The alternative is to skip this pack and build the graph from the interface's built-in nodes for the same model.
The difference is not preference and it is not performance. The README states the stock model ignores reference latents, so the alternative path runs your edit adapter without the reference conditioning it was trained to use. You get an image. It is produced by an adapter deprived of half its input.
That makes this less a competing implementation than a missing piece. The honest framing is that the built-in nodes are correct for ordinary generation with this model and incomplete for reference-based editing with these adapters, and this pack fills exactly that gap while deliberately not changing anything else.
The wider alternative is a different editing approach altogether, using a model and workflow where reference conditioning is supported natively by the interface. That avoids a custom node pack, a compatibility matrix and a pinned commit, at the cost of abandoning the adapter you trained. For anyone who has already invested in training with this toolkit, that is not a real option, which is precisely the position this project is written for.
MIT terms, and what to check before a long run
The project is MIT licensed with the file present, which for a small node pack is the expected and unobtrusive choice. The model and adapter weights you load carry their own terms and are a separate question from the licence on these nodes. This is not legal advice.
The repository is as small as its purpose: an initialisation file, a single node module, the project manifest, a workflow directory holding example graphs, and an ignore file for the interface's packaging. Shipping example workflows matters more than usual in this ecosystem, because a node graph is the real documentation and a wiring description in prose is harder to follow than a file you can open.
Before committing to a long generation run, three checks are worth making in order. Confirm the text encoder checkpoint you are loading includes the vision weights, since that is the documented condition for references to encode at all. Confirm how your adapter was trained before enabling the caching option, because the README's warning is explicit and the failure is silent. Then run one image with the patch node bypassed and one with it active, since the difference between those two outputs is the entire reason to install this.
Editorial conclusion
This node pack is the right answer if you trained an edit adapter with the specific toolkit and settings it targets and want to run it in this interface, where the built-in path would drop your reference conditioning without saying so. It has nothing to offer anyone outside that combination, and a different editing workflow with native reference support avoids the custom nodes entirely if you have not already trained an adapter. Check that your text encoder checkpoint carries the vision weights before anything else, leave the caching option off unless the adapter was trained with the matching setting, and pin a commit, since the repository has no tagged releases.
Frequently asked questions
Why do I need a model patch node for reference images?
Because the stock model in the interface ignores reference latents, as the README states. Without the patch the adapter runs but never receives the reference conditioning it was trained to use, producing output that looks valid and is not what the reference asked for.
Is the patch node safe to leave in a workflow?
Yes. The README states that when the conditioning carries no reference latents, the patched model behaves exactly like the stock model, so one graph can serve both reference-based and ordinary generation.
When should I enable the KV cache option?
Only when the adapter was trained with the matching option in the training toolkit, which the README states in bold. It precomputes the reference tokens' attention keys and values in a single first-timestep pass and reuses them, which is faster at high step counts but should be left off for normally trained edit adapters.
Does this node pack need extra dependencies?
No. The README states no extra dependencies are required and the project manifest declares an empty dependency list. Installation is a clone into the custom node folder followed by a restart, after which the nodes appear under their own category.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ostris-comfyui-krea2-ostris-edit)