# drawio-diagram-builder trades first-draft XML for a screenshot loop

> Will-hxw's agent skill wraps draw.io diagram generation in a render, screenshot, self-review cycle because the first XML draft usually has overlapping text, bad arrow routes and icons that do not match. Its install commands, however, name a differently spelled repository than the one hosting it.

**Will-hxw/drawio-diagram-builder** — Portable agent skill for research-style editable draw.io diagrams and screenshot-driven refinement

- Repository: https://github.com/Will-hxw/drawio-diagram-builder
- Stars: 413 · Forks: 13
- Language: Python
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/will-hxw-drawio-diagram-builder

## Without browser automation the skill degrades to plain XML

The prerequisites table is where the real design decision sits. Python 3, 3.7 or newer, covers the preview and validation scripts. Browser automation, whether that is Playwright MCP, Puppeteer or whatever browser tools the agent already has, covers the screenshot feedback loop. Internet access is needed because the preview loads `https://embed.diagrams.net/`. Strip the middle row out and the skill still works, in a much weaker way: the agent can still generate `.drawio` XML, but it cannot visually verify the result. The README is direct about which half matters, calling the iterative refinement loop the skill's main value. The workflow it enforces is a cycle: synthesise the user's text and image inputs into a diagram brief, create editable XML, preview through a local URL rather than a giant encoded one, screenshot, self-review both visible and semantic defects, fix, repeat, then validate. The local URL detail is not cosmetic, since one of the documented failure modes is large diagrams crashing on Windows with long URL failures.

## Six named ways the first draft goes wrong

The argument for a skill rather than a prompt is a list of specific, reproducible defects that a language model produces when it writes draw.io XML unaided. Text overlaps other text or escapes the box it belongs to. Arrows route incorrectly. Loop arrows look wrong. Icons go missing or turn inconsistent across a single figure. Reference figures get embedded as images instead of redrawn as editable objects, which quietly destroys the thing you asked for. And large diagrams crash on Windows with long URL failures. Each of those is a thing the screenshot pass is supposed to catch, and each is a thing XML validation cannot catch, because all six of them render without error. That distinction is the whole argument for the browser automation row in the prerequisites: a valid file is a weak guarantee, and the failure list above is a catalogue of files that are valid and still wrong. The inputs are equally broad, which is why the workflow starts with synthesis rather than drawing. A figure can come from a paper, a prompt, a codebase, project context, a screenshot, or one or more visual references, and the skill's job at the first step is to turn that pile into a single diagram brief that states what the figure has to say before anyone decides where anything goes. Two of the six defects on the list are about fidelity to a source rather than tidiness, and they are the two that punish an agent working from memory: an embedded reference image and an icon that does not match the paper's own symbol.

## Reference images force a spec, a grid and a defect log

When you hand the skill a figure to copy, the protocol gets stricter than XML authoring. Before drawing anything, the agent has to write a visual spec, set out a coordinate layout grid, maintain an asset ledger, and keep a defect log. Those four artifacts are the difference between copying and improvising: the grid removes guesswork about where things sit, the ledger records every icon decision, and the defect log is what stops the same mistake being reintroduced on pass three. The handoff rule is equally strict, because the final deliverable has to include a screenshot reviewed defect log rather than just valid XML. That is a deliberate inversion of the usual contract. Most diagram tooling treats a file that opens as finished, and this skill treats a file that opens as the starting point for the review that has not happened yet. A screenshot reviewed defect log is something you can read and argue with before you accept the figure.

## Mixed inputs are classified before any arrow is drawn

For prompts, papers, codebases or a mix of all three, a separate self supervision protocol applies. Every input gets classified as content, structure, style, layout or asset evidence, which stops a colour sample from being treated as a layout instruction. Connector semantics get defined before arrows are drawn, not debugged afterwards, and the list the skill audits afterwards is specific: fan-in, fan-out, feedback loops, grouped routes and arrowhead direction. The screenshot inspection then looks for six named categories of problem, requirement mismatches, arrow meaning, text overlap, icon coherence, style drift and regressions. Style drift and regressions are the two that only appear after the second pass, when a fix for one part of the figure quietly breaks another part that had already been corrected. The stated aim is to turn your inputs into observable requirements and visual constraints, render the editable result, compare it against the brief and the references, and then fix concrete mismatches rather than reworking the whole figure from a fresh prompt. The auditing list is what makes the connector rules checkable, because each term describes an arrangement a reviewer can look for and name, rather than a feeling that the arrows look busy. Fan-out is the branch point where one box feeds several, fan-in is several feeding one, feedback loops are the arrows that close back on their origin, grouped routes are several lines travelling together as one relationship, and arrowhead direction is the cheapest defect to fix and the easiest to get backwards. Defining them before drawing means a screenshot reviewer can say which category a problem belongs to instead of describing a symptom.

## Every install line spells the repository with a -skill suffix

Here is the discrepancy worth knowing before you copy anything. The project is published as `Will-hxw/drawio-diagram-builder`, but the one line installer points somewhere else:

```bash
npx skills add Will-hxw/drawio-diagram-builder-skill
```

The manual routes use the same suffixed name in their clone URL. For Claude Code:

```bash
git clone https://github.com/Will-hxw/drawio-diagram-builder-skill.git
cp -R drawio-diagram-builder-skill/skills/drawio-diagram-builder ~/.claude/skills/
```

The directory you copy into, `skills/drawio-diagram-builder`, carries the unsuffixed name, so the skill identifier and the repository path disagree while the inner folder does not. If a script you copy by hand fails at the clone step, that mismatch is the first thing to check. The agent prompt given for a quick install carries the same suffixed URL, and it also asks for something the manual routes do not: run the bundled smoke test afterwards and report the exact skill path, so you learn where it landed rather than assuming it landed in the default directory.

## Three copy routes, one per agent home directory

Without `npx skills`, the skill is a folder you copy into whichever directory your agent reads skills from. Codex on Windows uses PowerShell:

```powershell
git clone https://github.com/Will-hxw/drawio-diagram-builder-skill.git
New-Item -ItemType Directory -Force "$env:USERPROFILE\.codex\skills" | Out-Null
Copy-Item -Recurse -Force .\drawio-diagram-builder-skill\skills\drawio-diagram-builder "$env:USERPROFILE\.codex\skills\"
```

Codex on macOS or Linux uses the shell equivalent:

```bash
git clone https://github.com/Will-hxw/drawio-diagram-builder-skill.git
mkdir -p "$HOME/.codex/skills"
cp -R drawio-diagram-builder-skill/skills/drawio-diagram-builder "$HOME/.codex/skills/"
```

Restart the agent once the copy is done, then verify rather than assume: ask it to report its active skill path and run `python <installed-skill-dir>\scripts\check_skill_update.py`. If that script reports `OUTDATED` or `UNKNOWN`, the instruction is to reinstall from the canonical repository rather than to go looking for specific files or text snippets, which is a deliberate refusal to let you pattern match your way to a verdict.

## The bundled icons are vectors, not decomposed cells

The asset story has a caveat attached to it, and the caveat is the useful part. A MIT licensed subset of Tabler outline icons ships under `skills/drawio-diagram-builder/assets/icons/tabler/outline/`, with the full inventory and licence notes in `assets/icons/ICON-MANIFEST.md`, covering generic document, media, storage, model, routing, status, metric and tool symbols. What they are not is a set of native draw.io objects: the SVG internals are not decomposed into draw.io primitive cells, so when full object level editability matters you are told to use the primitive recipes instead. That is the difference between an icon that looks right and an icon whose every part you can move. Where a paper uses a branded or paper specific icon that the subset does not cover, there are two acceptable routes, and both are honest: supply the exact asset, or record the approximation in `asset-ledger.md`. The same ledger that the reference replication protocol requires is where that note belongs, which is one of the few places in this workflow where the bookkeeping and the drawing are the same document.

## Six reference files, plain Python scripts, no pip

The skill folder is small enough to read end to end. `SKILL.md` holds the main workflow, `VERSION` is the installed version marker, `agents/openai.yaml` carries the agent specific config, and `references/` holds six documents that each take on one job: `drawio-workflow.md`, `primitive-icons.md`, `reference-replication-protocol.md`, `self-supervision-and-intake.md`, `topconf-paper-style.md` and `xml-authoring.md`. Every helper script is plain Python 3 with no pip packages needed, which is what keeps the prerequisite list down to a Python version, a browser and a network connection. The wider repository adds `tests/`, a `picture/` directory, a `.codegraph/` directory, a Chinese README in `README-cn.md`, and two worked examples you can open directly: `examples/minimal.drawio` and `examples/hierarchical-memory-routing-replica.drawio`. Two gaps are worth naming. The repository layout block in the README is cut off partway through the tree, and the README itself stops mid path at `skills/drawio-diagram-builder/assets/refer`, so the bundled top conference style references that section introduces cannot be enumerated here. The last push to main landed on 2 July 2026 and there are no published releases.

## Conclusion

Adopt it when you need figures a reader can edit rather than images, and when your agent has browser automation available, because without a screenshot step it degenerates into the exact one-shot XML it was written to fix. Do not expect a single pass: the skill states outright that high fidelity reproduction takes several screenshot feedback rounds, and that the bundled SVGs are not decomposed into draw.io primitive cells, so full object level editability means using the primitive recipes instead. Before you install anything, check the repository name in the command you copy, since every install line points at `drawio-diagram-builder-skill` rather than the project name itself, and after copying confirm the version script does not report `OUTDATED` or `UNKNOWN`.

## FAQ

### Can ChatGPT create a draw.io diagram?

It can write draw.io XML, but the first result is usually not right. The README names the recurring defects: text that overlaps or escapes boxes, arrows that route incorrectly, loop arrows that look wrong, icons that are missing or inconsistent, reference figures embedded as images instead of redrawn, and large diagrams crashing on Windows with long URL failures.

### Is DrawIO completely free?

The README does not address draw.io licensing. What it does state is that this skill is not a draw.io replacement and is not affiliated with diagrams.net or JGraph. The bundled Tabler icon subset it does ship is MIT licensed.

### What is draw.io now called?

The skill's own text uses both names, referring to editable diagrams.net and draw.io figures throughout. The preview step is the concrete answer: it loads https://embed.diagrams.net/ through a local URL rather than a giant encoded one.

### What is the best free diagram maker?

The README does not rank diagram makers, and it places this tool on top of draw.io rather than beside it. Its scope is producing editable draw.io figures with a screenshot review loop; it is explicitly not a draw.io replacement and carries no affiliation with diagrams.net or JGraph.

## Sources

- [Issues](https://github.com/Will-hxw/drawio-diagram-builder/issues)
- [License: MIT](https://github.com/Will-hxw/drawio-diagram-builder/blob/main/LICENSE)
- [README](https://github.com/Will-hxw/drawio-diagram-builder/blob/main/README.md)
- [Will-hxw/drawio-diagram-builder on GitHub](https://github.com/Will-hxw/drawio-diagram-builder)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/will-hxw-drawio-diagram-builder
