# See-through: single-image layer decomposition for anime characters, reviewed

> See-through turns one anime illustration into a layered PSD with up to 23 semantically separated parts. It is a research codebase with a two-stage head and body pipeline, a heavy PyTorch stack, and no packaged release.

**shitagaki-lab/see-through** — "Single-image Layer Decomposition for Anime Characters" (SIGGRAPH 2026 Conference Paper)

- Repository: https://github.com/shitagaki-lab/see-through
- Stars: 4,285 · Forks: 393
- Language: Python
- License: Apache-2.0
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/shitagaki-lab-see-through

## What See-through solves for anime illustrators and riggers

A flat anime illustration is a single raster image. To animate or rig it, an artist normally repaints the character as separate parts: hair, face, eyes, clothing, accessories, each with the pixels behind it filled in. See-through automates that decomposition. The README states the framework "decomposes a single image into fully inpainted, semantically distinct layers with inferred drawing orders, up to 23 layers including hair, face, eyes, clothing, accessories, and more." Inpainting is the part that matters. A cut-out layer leaves transparent holes where the hair covered the face; See-through generates the missing pixels so each layer stands alone. The target user is the person doing rigging work, not someone who wants a quick filter. The paper is published in ACM SIGGRAPH 2026 Conference Papers, so the code is a research artifact with the expectations that come with one.

## How the pipeline stratifies one image into 23 layers

The main script is inference/scripts/inference_psd.py. According to the README it "applies the LayerDiff 3D model for transparent layer generation and the fine-tuned Marigold model for pseudo-depth inference, then stratifies the character into up to 23 semantic layers and exports a layered PSD file." Three models sit behind that: layerdiff3d, a diffusion-based transparent layer generator built on SDXL; a Marigold depth model fine-tuned for anime; and l2d_sam_iter2, a semantic body part segmentation model. Depth gives the ordering cue, segmentation gives the semantic labels, and the diffusion model produces the inpainted pixels. The README is explicit that head and body separation run as "two continuous stages, which may lead to a longer time than the original model mentioned in the paper." That is a real cost, not a rounding error: two passes over the same image. Output lands in workspace/layerdiff_output/ by default and includes the PSD plus intermediate depth maps and segmentation masks. A separate script, inference/scripts/heuristic_partseg.py, does depth-based or left-right stratification after the fact.

## Installing See-through and running a first decomposition

There is no package on PyPI. Setup is a conda environment plus a pinned PyTorch wheel set. The README gives Python 3.12 and CUDA 12.8:

```bash
conda create -n see_through python=3.12 -y
conda activate see_through
pip install torch==2.8.0+cu128 torchvision==0.23.0+cu128 torchaudio==2.8.0+cu128 \
  --index-url https://download.pytorch.org/whl/cu128
```

Two hardware caveats are in the README itself. On aarch64 the pinned versions may not be available and it says to use torch>=2.9.0 instead. AMD ROCm users are told to install the ROCm wheel set matching their local runtime, with ROCm 7.2 shown as an example. After the torch install, the dependencies and an assets symlink:

```bash
pip install -r requirements.txt
ln -sf common/assets assets
```

The requirements file installs ./common and ./annotators as editable local packages, so the repository layout is part of the install. Optional annotator tiers add detectron2, SAM2 and mmcv/mmdet for body attribute tagging, language-guided segmentation and anime instance segmentation respectively. The README warns to "always run scripts from the repository root as the working directory", which is easy to miss because the script lives two directories down. The first real run:

```bash
python inference/scripts/inference_psd.py \
  --srcp assets/test_image.png \
  --save_to_psd
```

Passing a directory to --srcp processes every image in it. The result appears under workspace/layerdiff_output/ as a layered .psd with the intermediate masks alongside. Expect the first run to spend most of its time downloading the three model repositories from HuggingFace.

## Where See-through breaks down or is the wrong tool

The README does not document rollback, so there is no stated way to undo a bad decomposition other than discarding the output directory. There are no retrieved releases, which means no versioned artifact to pin against and no changelog to read when behaviour shifts. The stack is heavy: transformers 5.0.0, diffusers 0.37.0, PyQt6 6.9.1 and a timm build pulled from a specific git commit. Several dependencies are pinned to git revisions rather than published versions, so reproducibility depends on those repositories staying reachable. The two-stage head and body separation makes this slower than the paper's model, by the README's own admission. Layer count is a ceiling, not a guarantee: "up to 23 layers" means simple images may produce fewer, and the README does not say what happens when the segmenter mislabels a part. If you need a CPU-only pipeline, or a stable library with a semantic version, this is the wrong project. If your source images are photographs or 3D renders rather than anime line art, the fine-tuned depth and body parsing models were not built for that.

## See-through versus manual layer separation in an editor

The realistic alternative is not another repository. It is an artist cutting and repainting layers by hand in Clip Studio Paint or Photoshop, which is what the pipeline is trying to replace. The difference is in what each produces. Manual separation gives the artist full control over layer boundaries and drawing order, and the result matches the intended rig exactly, at the cost of hours per character. See-through gives a PSD in minutes but the boundaries come from a segmentation model, and the drawing order is inferred from pseudo-depth. For a quick background character or a prototype rig, the automated output is likely good enough to clean up. For a hero character with a specific rigging plan, the cleanup may cost more than starting from scratch. The heuristic_partseg.py script is the escape hatch: it lets you re-stratify the generated PSD by depth or left-right position instead of accepting the model's grouping.

## Maintenance, licence and the cost of upgrading

The repository is not archived. The last push was on 2026-08-05, which is recent enough that the project is still being touched, though with no retrieved releases there is no upgrade path to speak of beyond pulling main. That matters for a research codebase: the pinned git dependencies in requirements.txt mean an upgrade is a deliberate act of re-resolving several moving parts, not a version bump. The licence is Apache-2.0, which permits commercial use and modification with the usual attribution and notice requirements; the model weights on HuggingFace are separate artifacts and the README does not state their licence terms, so check each model repository before shipping anything. The README also carries a notice that this is an open-source research project with no paid service, and that any site charging for the functionality is not affiliated with the authors. That is worth reading as a provenance warning, not just a disclaimer.

## Conclusion

Adopt See-through if you are an animator, Live2D rigger or researcher who needs editable, inpainted layers from a single anime illustration and can run a CUDA PyTorch environment from the repository root. Skip it if you need a packaged CLI, a pip-installable library, CPU-only inference, or a documented rollback path; there is no release, no versioned wheel and the README does not document rollback. Verify first that layerdiff3d, marigold and l2d_sam_iter2 load on your GPU, then run inference_psd.py on one test image and open the resulting PSD in your editor before committing to a batch.

## FAQ

### How do I install See-through?

Create a Python 3.12 conda environment, install the pinned CUDA 12.8 PyTorch wheels from the PyTorch index, then run pip install -r requirements.txt and create the assets symlink with ln -sf common/assets assets. The README notes that aarch64 users may need torch>=2.9.0 and ROCm users should match their local runtime.

### How many layers does See-through produce from one image?

The README states the pipeline stratifies a character into up to 23 semantic layers including hair, face, eyes, clothing and accessories, and exports a layered PSD. The count is a ceiling rather than a fixed output.

### Can I try See-through without installing it?

Yes. The README links a HuggingFace Space with ZeroGPU where a registered user can run roughly 1-2 PSD extractions per day at 1280 resolution, taking about 2-3 minutes each, plus a ModelScope demo for users in Mainland China that the README describes as completely free with slightly higher resolution.

### Why does See-through take longer than the model described in the paper?

The README explains that separation for the head and the body runs in two continuous stages, which it says may lead to a longer runtime than the original model mentioned in the paper.

## Sources

- [Issues](https://github.com/shitagaki-lab/see-through/issues)
- [License: Apache-2.0](https://github.com/shitagaki-lab/see-through/blob/main/LICENSE)
- [README](https://github.com/shitagaki-lab/see-through/blob/main/README.md)
- [shitagaki-lab/see-through on GitHub](https://github.com/shitagaki-lab/see-through)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/shitagaki-lab-see-through
