FigMirror: Turn a Reference Paper Figure Into a matplotlib Script
An Automated AI Agent Tool for Plotting Your Data in Any Paper's Figure Style.
At a glance
- What is it?
- FigMirror is a Python project that ships as Claude Code and Codex skills plus a local web app, using a Drawer / Reviewer agent loop to restyle your data after a reference figure. Here is what the repository actually documents, and where it stops.
- Who is it for?
- Adopt FigMirror if you already work inside Claude Code or Codex and your bottleneck is matching a venue's figure conventions across many panels; the role-separated Drawer / Reviewer loop and the bounded iteration contract are the parts worth evaluating. Do not adopt it if you need a stable, documented CLI you can wire into CI, or if you cannot accept an agent writing and revising matplotlib code on your machine.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem FigMirror targets: matching a paper's figure style, not just making a chart
Most plotting libraries answer the question "how do I draw this data". They do not answer "how do I draw this data so it looks like it came from the same paper as Figure 3 of the reference I am imitating". That second question is what consumes time in practice: axis ranges, tick density, panel proportions, colormap choice, font sizes, the amount of whitespace between subplots. Getting a joint hexbin plot with marginal histograms to sit next to a dense multi-panel composition from another paper usually means hand-editing matplotlib code repeatedly until the two look related.
FigMirror is aimed at that gap. The README describes the workflow as picking a reference figure, pasting your data, and getting back an editable matplotlib script plus a camera-ready PDF. The audience is narrow on purpose: researchers and graduate students who already write matplotlib and who care about visual consistency with a specific venue or a specific prior paper. The project ships as skills for Claude Code and Codex, so it assumes the user is already working inside one of those agents rather than at a bare terminal.
The Drawer / Reviewer loop and what the iteration contract actually constrains
The mechanism visible in the repository is a two-role agent loop. A Drawer produces matplotlib code for the current iteration; a Reviewer inspects the rendered output against the reference and returns annotated feedback, which becomes input to the next iteration. The milestones list names the components added on 2026-06-17: role-separated Drawer and Reviewer agents, far-near visual review views, reviewer bounding boxes, and annotated feedback passed into the next iteration. The bounding boxes matter because they localize the criticism: instead of "the figure does not match", the Reviewer can point at a region.
On 2026-08-17 the project ported that same production loop to Claude Code, described in the README as the same decision state machine and bounded iteration contract. "Bounded" is doing real work in that sentence. An unbounded generate-and-critique loop is expensive and can oscillate; the repository states there is a contract limiting it, though the README does not publish the iteration cap itself.
The runtime shape is visible in pyproject.toml. The project requires Python 3.10 or newer, and its three runtime dependencies are pillow, matplotlib and numpy. A comment in that file explains why: the per-iteration Drawer and Reviewer subprocess scripts are invoked through `uv run --project`, and that command does not install dev-group packages into a fresh cache, so anything an agent-generated script imports has to sit in the runtime dependency list. That is a small detail with a large consequence. Every script the Drawer writes can only rely on pillow, matplotlib and numpy unless you change the project file.
Installing FigMirror as a Claude Code or Codex skill
The README's quick start does not give a pip command or a git clone line. Its first instruction is to paste a sentence into the agent and let the agent perform the setup. The README gives this example:
Install FigMirror for me: https://github.com/VILA-Lab/FigMirrorIf you are already inside Claude Code or Codex, that is the entire documented installation path. The agent reads the repository and wires up the skill. There is no published `pip install figmirror` in the README, and the package name in pyproject.toml is figmirror with version 0.1.0.
For anyone who wants to inspect or run the project directly rather than through an agent, the repository layout is the guide: .claude/ and .codex/ hold the skill definitions, scripts/ holds the code the agent subprocesses call, tests/ holds the pytest suite, and uv.lock pins the environment. The declared runtime dependencies are the three listed above, so a manual environment needs those plus Python 3.10 or later:
python -m venv .venv
. .venv/bin/activate
pip install "pillow>=12.2.0" "matplotlib>=3.10.9" "numpy>=2.2.6"
pytestThat last line reflects the pytest configuration in pyproject.toml, which sets testpaths to tests. The file notes that the tests run against scripts/ directly, with conftest.py adding that directory to sys.path, and that matplotlib and numpy are runtime dependencies for the agent subprocess scripts rather than for pytest itself. So a green test run confirms the harness, not that the agent loop will produce a figure you like.
The second documented entry point is the Web UI, described as the option to use when you want upload, preview, iteration history and refinement in a browser. The README's text for that subsection is truncated in the published file, ending at "in a brows". Treat the Web UI as real but under-documented until you read the source.
Where FigMirror is the wrong tool
The first limitation is dependency surface. Because agent-generated Drawer scripts run through `uv run --project` against the runtime dependency list, a figure that needs seaborn, pandas, scipy or plotly is outside the declared environment. You can add packages to pyproject.toml, but then you own that divergence, and the comment in the file makes clear the split between runtime and dev groups is deliberate.
The second is verification. This is an LLM agent writing and rewriting plotting code. The README's own framing is iterative: render, review, revise. Nothing in the repository description suggests the output is deterministic or that the same input yields the same figure twice. If your figures feed a reproducibility artifact or a CI check that diffs rendered output, an agent loop is the wrong shape for that job.
The third is fit with your chart family. The paper reports results on PlotTwin-Bench, a benchmark the project itself introduces, using a 150-instance subset: 50 hand-curated references and 100 randomly sampled augmented references. FigMirror scores 72.7 combined on the hand-curated split and 76.4 on the augmented split, with all methods in that comparison using GPT-5.5. Those are the project's numbers on the project's benchmark. A 3D waterfall plot or a joint hexbin with marginal histograms appears in the showcase, but the showcase is a gallery of outputs, not a per-family accuracy table. If your figure type is unusual, you are extrapolating.
Finally, if you want a command-line tool with flags you can script, FigMirror does not present itself that way. It presents itself as a skill for an agent.
How FigMirror differs from Plot2Code, METAL, ChartGalaxy and ChartIR
The arXiv preprint compares FigMirror against four named baselines on PlotTwin-Bench: Plot2Code, METAL, ChartGalaxy and ChartIR. The qualitative comparison figure in the README shows each group with the reference, FigMirror, and those four baselines on the same target data.
The difference in approach is conditioning. The task FigMirror sets up is reference-conditioned style transfer: a reference figure defines the target style, and the system iterates until the output is judged to belong to the same paper family. The evaluation is split by reference source, hand-curated versus augmented, and reported separately rather than pooled. That split is the honest part of the setup, because augmented references are likely easier to match than hand-picked ones, and the two scores differ by 3.7 points.
If you are choosing between these tools, the practical question is whether you have a reference figure at all. FigMirror's whole loop depends on one. If you do not, the repository points elsewhere: a Hugging Face Space gallery with 139 paper figures across 25 chart families, offered as a starting point when no high-quality reference is at hand. That gallery is a source of references, not a FigMirror feature.
Maintenance, licence and upgrade cost
The repository is not archived, and the last push was on 2026-09-02. The most recent release is v0.1.0, published on 2026-05-23 and labeled a public preview. Between the release and the last push, the milestones record an algorithm update in June, an evaluation update in July, Claude Code parity in August, and the arXiv preprint on 2026-08-28. So the version number understates how much has changed: v0.1.0 is the only tagged release, while the skill definitions have moved several times since.
That is the upgrade cost. If you install through the agent instruction, you are tracking the main branch of a public preview with one tagged release, and the Drawer / Reviewer loop is the part most likely to shift. There is no changelog file in the listed top-level entries, so the milestones section of the README is the closest thing to a release history.
On licensing: the repository metadata does not state a licence, and the README does not carry a licence section in the published text. That is not a detail you can defer. Without a licence file you have no stated grant to redistribute, modify or ship the code inside a commercial product, and academic use is not automatically permitted either. Check the repository for a LICENSE file and, if one exists, read it before you build anything on top. This is a factual gap in the project, not a legal opinion.
Editorial conclusion
Adopt FigMirror if you already work inside Claude Code or Codex and your bottleneck is matching a venue's figure conventions across many panels; the role-separated Drawer / Reviewer loop and the bounded iteration contract are the parts worth evaluating. Do not adopt it if you need a stable, documented CLI you can wire into CI, or if you cannot accept an agent writing and revising matplotlib code on your machine. Before committing, verify three things yourself: the licence, since the repository metadata does not state one; whether the Figure 1 numbers in the arXiv preprint map onto your own chart families, since the evaluation uses a 150-instance subset of PlotTwin-Bench; and whether the Web UI path in the README is complete enough to run, because the quick-start section as published ends mid-sentence at the Web UI heading.
Frequently asked questions
How do I install FigMirror?
The README's documented path is to paste the instruction "Install FigMirror for me: https://github.com/VILA-Lab/FigMirror" into Claude Code or Codex and let the agent do the setup. There is no published pip install command in the README. A manual environment needs Python 3.10 or later plus pillow, matplotlib and numpy.
What does FigMirror output?
The README states that you pick a reference figure, paste your data, and get an editable matplotlib script plus a camera-ready PDF. The script is produced by a Drawer agent and revised across iterations based on Reviewer feedback.
Does FigMirror work with Claude Code and Codex?
Yes. The project ships as both a Claude Code skill and a Codex skill, and the milestones note that the role-separated Drawer / Reviewer loop was ported to Claude Code on 2026-08-17 with the same decision state machine and bounded iteration contract.
What is PlotTwin-Bench in the FigMirror paper?
It is the benchmark introduced in the arXiv preprint "FigMirror: Ground It, Code It, Plot It" for reference-conditioned scientific-figure style transfer. The paper reports a 150-instance evaluation subset made of 50 hand-curated references and 100 randomly sampled augmented references.
What licence does FigMirror use?
The repository metadata does not state a licence, and the published README does not include a licence section. You should check the repository for a LICENSE file before relying on any particular terms.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/vila-lab-figmirror)