# One agent draws, another argues, and the loop is bounded

> FigMirror takes a reference figure from a paper, pastes your data over it, and runs a Drawer and Reviewer agent loop inside Claude Code or Codex until the plot matches the reference, handing back an editable matplotlib script and a camera-ready PDF.

**VILA-Lab/FigMirror** — An Automated AI Agent Tool for Plotting Your Data in Any Paper's Figure Style.

- Repository: https://github.com/VILA-Lab/FigMirror
- Website: https://arxiv.org/abs/2608.28814
- Stars: 521 · Forks: 39
- Language: Python
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/vila-lab-figmirror

## The orchestrator delegates drawing and auditing to two named agents

The core mechanism is an agentic Drawer and Reviewer loop rather than a single prompt. A top-level Orchestrator hands drawing to the named figmirror-drawer agent and the visual audit to a reviewer agent, and the reviewer's job is to look at the rendered result and say what is wrong. A June 17, 2026 milestone shows how that loop was sharpened: role-separated Drawer and Reviewer agents, far-near visual review views, reviewer bounding boxes, and annotated feedback passed into the next iteration. Those three details are the difference between a reviewer that comments on a plot and one that points at the axis it means. A bounding box turns a vague remark into a coordinate, and passing the annotation into the next round means the drawer is not re-reading the whole history to guess what was meant. On August 17, 2026 the same production loop was ported to Claude Code with the same decision state machine and the same bounded iteration contract, so the two harnesses behave alike rather than one of them being the better path.

## The paper calls the second half Grounded Measurement

The diagram that explains the loop has two panels, and the second one introduces something called Grounded Measurement. The title of the August 28, 2026 preprint spells the same idea as a sequence: FigMirror: Ground It, Code It, Plot It. Read together, that phrase describes the order the system works in. Ground it means establishing what the reference figure actually encodes before drawing anything, code it means producing script rather than a picture so the result stays editable, and plot it means rendering your data through the style you just established. The illustration is labelled as an illustration rather than a screenshot of a run, so the mechanics of that step are not spelled out in the README. What is documented is the output contract: pick a reference figure, paste your data, and get an editable matplotlib script plus a camera-ready PDF. The script is the deliverable that matters, since it can be rerun when the data changes.

## Installation is a skill plus two custom agents, per harness

There are two supported runtimes and the installer branches on them explicitly, because FigMirror does not run itself. It installs as a skill named figmirror together with two custom agents, figmirror-drawer and figmirror-reviewer, which are the ones the current role-separated algorithm calls. If you are already inside Claude Code or Codex you can paste one line and let the agent do the setup. If you would rather do it yourself, the script auto-detects which harnesses are present, and you can also target one explicitly with the --codex, --claude or --all arguments passed after the installer. Cloning and installing locally runs the same installer from the checkout. What you get is not a binary on your PATH: the figures are produced by Python helpers invoked through uv run --project, which is why the runtime dependencies rather than the development group decide what an agent-generated script is allowed to import. After setup you attach a paper-figure screenshot, paste your data, and ask for the style to be mirrored. Targeting a single harness is a flag on the same installer:

```bash
curl -fsSL https://raw.githubusercontent.com/VILA-Lab/FigMirror/main/scripts/install.sh | bash -s -- --codex
curl -fsSL https://raw.githubusercontent.com/VILA-Lab/FigMirror/main/scripts/install.sh | bash -s -- --claude
curl -fsSL https://raw.githubusercontent.com/VILA-Lab/FigMirror/main/scripts/install.sh | bash -s -- --all
```

## The web UI serves on 127.0.0.1:8765 and takes a backend flag

The browser path exists for people who want to see the iteration history rather than read a transcript, and it is a local service rather than a hosted site. The serve script takes a workspace directory and a backend, and the backend flag is the switch between harnesses: codex runs it through Codex, and the same interface can be run through Claude Code by switching the value. The workspace argument points at a directory under .artifacts, so uploads, previews and iteration history are written inside your checkout rather than in a temporary directory you cannot find afterwards. The interface is then opened at 127.0.0.1:8765, bound to the loopback address, which is the right default for a tool that renders untrusted data but does mean you cannot reach it from another machine without changing the binding. The command needs uv, and the README's only preparation step for that is installing uv through pip if it is missing.

```bash
git clone https://github.com/VILA-Lab/FigMirror.git && cd FigMirror
bash scripts/install.sh
uv run python scripts/figcopy_serve.py --workspace .artifacts/figmirror-workspace --backend codex
```

## Small-root machines need UV_CACHE_DIR moved before the first run

This is the most concrete piece of operational advice in the project and it is easy to miss. On shared machines or anything with a small root filesystem, the instruction is to run df -h and export UV_CACHE_DIR to a directory you own on the largest writable non-root filesystem before the first invocation. The reason is visible in the packaging notes. uv run --project does not install development-group packages when it is pointed at a fresh cache, so anything a generated script imports has to be declared as a runtime dependency. That is why matplotlib and numpy, which every stage-one Drawer script imports, sit in the main dependency list next to pillow, and why pytest alone lives in the development group. Python 3.10 or newer is required. Tests run against the scripts directory directly, with the conftest adding it to the import path, so the helpers are exercised as scripts rather than as an installed library.

## 72.7 on curated references, 76.4 on augmented ones, all on GPT-5.5

The preprint reports results on a 150-instance evaluation subset drawn from the benchmark, made up of all 50 hand-curated references and 100 randomly sampled augmented references. FigMirror takes the highest combined score on both splits, 72.7 on the hand-curated set and 76.4 on the augmented set, and the two splits are reported separately rather than merged. One detail governs how the numbers read: every method in the comparison uses GPT-5.5, so the comparison holds the model fixed and varies the method. The scores are combined scores across code, vision and code-plus-vision judging, not a single axis, and the underlying scorer was described in the July 1, 2026 milestone as an internal hybrid style scorer built for repeatable method comparison. The benchmark is published on the Hugging Face Hub as figmirror/PlotTwin-Bench for reference-conditioned scientific-figure style transfer, with the default bench configuration holding the 150 evaluation tasks.

## 399 references in the dataset, 139 figures in the fallback gallery

Two reference collections exist and they are not the same size, which matters if you are judging coverage. The full configuration of the benchmark dataset holds all 399 references, while the bench configuration is the 150-task evaluation set. Separately, there is a Hugging Face space called figcopy-taxonomy-gallery, and its pitch is for the case where you have no high-quality reference at hand: start from 139 paper figures spread across 25 chart families. So the gallery is a starting menu rather than the evaluation pool, and it is smaller than the full reference set even though it sounds like the same asset under another name. It is not, and the naming does not help, since the gallery, the serve script and the project itself use three different words for closely related things. Treat the gallery as a way to pick a target, and the bench configuration as the set the reported numbers were measured on.

## A v0.1.0 tag from May, a release milestone from June

Version bookkeeping is where the project's early stage shows. The single release is tagged v0.1.0 with the name FigMirror v0.1.0: Public Preview, published on 2026-05-23, while the milestone log dates the FigMirror release itself to June 1, 2026, so the tag predates the stated release by more than a week. The package version in the project file is also 0.1.0, matching the tag rather than the milestone. Nothing about that is fatal, but it means the tag is not a reliable marker for what was in the June release. Other signs of an in-progress project sit in the tree: a spike-results directory next to tests, an openspec directory that holds specifications, and both .claude and .codex directories committed so each harness's configuration travels with the code. The last push landed on 2026-09-30 and the repository is not archived, so work is continuing, and no licence file appears in the tree, which is worth resolving before you build on it.

## Conclusion

FigMirror earns its place when you are writing a paper, a report or a slide deck where the figure has to sit next to figures from a specific source and look like it came from the same lab, because matching that by hand means reading someone else's plotting code and reverse engineering their defaults. It does not fit a one-off chart, and it does not fit work where the numbers matter more than the look, since the benchmark measures style transfer rather than analytical correctness. Before you rely on it, check three things. Whether your harness is one of the two it supports, since the install path branches on Claude Code, Codex or both, and anything else is unsupported. Whether the reference you have is good enough, since the paper's own fallback for a weak reference is a gallery of 139 figures across 25 chart families. And whether you have room for the uv cache, because on a small-root machine the environment file asks you to move UV_CACHE_DIR before the first run.

## FAQ

### What is FigMirror?

It is a Python agent tool that takes a reference figure from a paper as a style target and renders your data through it. It runs inside Claude Code or Codex as a figmirror skill plus figmirror-drawer and figmirror-reviewer custom agents.

### How do I install FigMirror with Claude Code or Codex?

From an agent session, paste the install request with the repository URL and let the agent set it up. From a shell, run the bundled install script, which auto-detects both harnesses, and pass --codex, --claude or --all to target one explicitly.

### What files does FigMirror produce?

An editable matplotlib script and a camera-ready PDF. The script is the part that matters, since it can be rerun when the data changes, and the Python helpers it uses are invoked through uv run --project.

### What is PlotTwin-Bench?

It is the benchmark introduced with the FigMirror preprint for reference-conditioned scientific-figure style transfer, published on the Hugging Face Hub as figmirror/PlotTwin-Bench. The default bench config holds the 150 evaluation tasks and the full config holds all 399 references.

### How was FigMirror evaluated?

On a 150-instance subset covering all 50 hand-curated references and 100 randomly sampled augmented references, reported separately. FigMirror scored 72.7 combined on hand-curated and 76.4 on augmented, with every method in the comparison using GPT-5.5.

### Can I use FigMirror without a good reference figure?

There is a fallback gallery of 139 paper figures across 25 chart families. You can also serve the tool locally with uv run python scripts/figcopy_serve.py, which opens a web interface on 127.0.0.1:8765 for upload, preview and iteration history.

## Sources

- [Issues](https://github.com/VILA-Lab/FigMirror/issues)
- [Project website](https://arxiv.org/abs/2608.28814)
- [README](https://github.com/VILA-Lab/FigMirror/blob/main/README.md)
- [Releases](https://github.com/VILA-Lab/FigMirror/releases)
- [VILA-Lab/FigMirror on GitHub](https://github.com/VILA-Lab/FigMirror)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/vila-lab-figmirror
