Academic Figure Generator: a local paper-to-figure pipeline with an agent skill on the side
AI 驱动的学术论文配图生成平台。上传论文 → AI 分析内容生成 Prompt → 一键生成高质量科研配图,还有配套的skill可在主流agent中使用
At a glance
- What is it?
- LigphiDonk's MIT-licensed tool turns a PDF, DOCX or TXT paper into a figure prompt via Claude, then renders it through a NanoBanana or Gemini endpoint. It runs on your own machine, and it asks you for two API keys before it does anything.
- Who is it for?
- Academic Figure Generator fits a single researcher or a small group who already has an Anthropic key and a NanoBanana or Gemini image key, wants paper-to-figure prompts generated locally, and is willing to run a Python backend and a Vite frontend.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap Academic Figure Generator targets: figures as the last unpaid chore of a paper
Writing a paper and drawing its figures are two different jobs, and the second one usually lands on the author at the end, when the deadline is close. The README states the intent plainly: turn "写完论文还要画图" into an upload, confirm, download sequence. That is a narrow claim, and it is worth taking at face value rather than reading it as a general illustration tool.
The audience is a researcher working alone on a local machine. Everything in the repository points that way. The README calls it a personal local version (个人本地版). The database is SQLite, created automatically on first start with no configuration step. Uploaded files and generated images sit in a local directory. There is no account system, no hosted URL, and no homepage field on the repository, so there is nothing to sign up for.
What it does not do is also clear from the feature table. It generates prompts from paper text and renders images from those prompts. It does not lay out a multi-panel figure with precise coordinates, and it does not produce vector output you could edit in Illustrator. The examples in the README (a prediction network architecture, a time-frequency signal processing flow, a deep learning module detail, an annotated anatomy diagram) are all single-image outputs. If your figure needs exact axis labels matching a data table, this is the wrong layer of the stack.
How the pipeline actually moves: parse, prompt with Claude, render asynchronously with NanoBanana
The architecture is a three-stage chain with a synchronous step in the middle and an asynchronous one at the end. The README's flow diagram shows it directly: the browser sends a paper, the backend parses it and extracts text plus section structure, the Claude Agent SDK analyzes that content and produces a figure prompt, the user confirms or edits the prompt and picks a resolution and aspect ratio, and then the NanoBanana API generates the image.
The split between the two AI calls matters. Prompt generation runs synchronously through the Claude Agent SDK, which the prerequisites describe as requiring Claude Code CLI installed and logged in on the local machine. Image generation runs as a background task. That is why the README lists SSE streaming as a feature: the frontend subscribes to progress events instead of polling, because the render step takes long enough that a blocking request would be awkward.
The service layer in the repository layout matches this: claude_code_service.py handles the SDK integration, document_service.py parses PDF, DOCX and TXT, image_service.py talks to NanoBanana, and local_storage_service.py writes files to disk. Metadata (projects, documents, prompts, images) goes into SQLite through async SQLAlchemy. There is also a shortcut path the README calls 快捷生成, which skips the upload entirely and takes a prompt string straight to the image step. That is the mode to use when you already know what you want drawn.
One thing the README does not describe: what happens to a render that fails midway, or whether a partially written image file is cleaned up. The SSE stream tells you about progress, not about recovery.
Installing Academic Figure Generator locally and generating your first figure
You need Python 3.12 or newer, Node.js 18 or newer, Claude Code CLI installed and logged in, and an image API key. The README's quick start begins by cloning the repository and editing the .env file at the project root. The four variables below are the ones the quick start shows; the first two are marked required in the environment variable table.
ANTHROPIC_API_KEY=your-anthropic-api-key
NANOBANANA_API_KEY=your-nanobanana-api-key
NANOBANANA_API_BASE=https://api.keepgo.icu
NANOBANANA_MODEL=gemini-3-pro-image-previewNANOBANANA_API_BASE already has that default, and NANOBANANA_MODEL defaults to gemini-3-pro-image-preview, so you can leave both out and supply only the two keys. Note that the repository root also contains a .env.example whose contents differ substantially: it lists Postgres, Redis, Celery, MinIO, JWT and rate-limit settings. The quick start does not reference those, and the README describes SQLite with zero configuration, so treat .env.example as a template for something other than the documented local setup.
Next, the backend. The README recommends a virtual environment inside backend and installs the package in editable mode.
cd backend
python -m venv .venv
source .venv/bin/activate
pip install -e .
uvicorn app.main:app --host 0.0.0.0 --port 8000 --reloadOn first start the SQLite database at backend/data/app.db and the data directories are created automatically. If the server comes up, the Swagger UI should be reachable at http://localhost:8000/docs.
The frontend is a separate process.
cd frontend
npm install
npm run devThe Vite dev server listens on localhost:5173 and proxies /api requests to localhost:8000, so you open http://localhost:5173 for the app itself. From there the first real use is: create a project, upload a PDF or DOCX or TXT paper (the default upload cap is 50 MB), wait for the prompt to appear, edit it if the wording is off, choose a resolution and aspect ratio, and start the generation. The image lands in the local data directory and can be downloaded, or fed back in with a text instruction for an image-to-image edit.
The academic-figure-prompt skill: the same idea without the platform
The repository ships a second, independent artifact. academic-figure-prompt is an AI Coding Agent Skill, and the README says it works with Claude Code, Gemini CLI, Cursor and similar assistants without deploying the full platform. Its scope is narrower and its cost is lower: it reads a paper (PDF, LaTeX or Word), offers eight preset academic color schemes including colorblind-friendly options such as Okabe-Ito, and emits a detailed English prompt for an image tool of your choice. It does not generate the image.
The install is one command, and there is a manual path for people who prefer to keep the files visible.
npx skills add LigphiDonk/academic-figure-generatorThe manual route clones the repository and copies the skill directory into the assistant's skill folder, for example .gemini/skills/ for Gemini CLI or .claude/skills/ for Claude Code. After that you trigger it in conversation, ask it to read a paper and produce figure prompts, and pick a color scheme by name when you have a preference.
This is the part of the project I would reach for first. The platform's advantage is the end-to-end loop and the image-to-image editing; the skill's advantage is that it needs no Python environment, no Node, no database and no second API key. If your bottleneck is knowing how to describe the figure rather than producing the pixels, the skill covers it. There is also a second directory in the repository root, academic-figure-prompt-pastel, which the README does not document at all, so its relationship to the main skill is unstated.
Where it breaks: two API keys, a CLI dependency, and a very permissive upload path
The most concrete limitation is that the tool cannot do anything without external paid services. Both ANTHROPIC_API_KEY and NANOBANANA_API_KEY are marked required. There is no offline mode and no bundled model. If either key is missing or the image endpoint is unreachable, the pipeline stalls at that stage, and the README documents no fallback.
The second constraint is the Claude Code CLI. The prerequisites say it must be installed and logged in on the machine. That is a heavier environment dependency than a plain API key, and it means the prompt-generation half of the pipeline is tied to Anthropic's tooling rather than to a model-agnostic interface. The README does not explain what error surfaces when the CLI is present but not authenticated.
Third, the upload path accepts PDF, DOCX and TXT up to 50 MB by default, and the README does not document how parsing handles difficult layouts. Two-column PDFs, figures with embedded text, and equations are the usual failure points for text extraction, and nothing in the repository description addresses them. A paper that extracts badly produces a prompt built on garbled content, and you would have to catch that by reading the generated prompt before spending an image call.
Finally, the deployment story is split in a way the documentation does not reconcile. The README and the quick start describe SQLite and a local filesystem. The .env.example describes Postgres, Redis, Celery and MinIO. Nothing in the README explains how to move from one to the other, whether the code paths differ, or whether the container-oriented variables are aspirational. Anyone planning a shared or server deployment is reading a template, not a guide.
Against BioRender and direct image models: different layer, different trade-off
The obvious comparison is BioRender, which the search questions about this project keep surfacing. The difference is in where the drawing comes from. BioRender gives you a curated library of scientific icons and a drag-and-drop canvas, so the output is assembled from vetted, consistently styled components and you control every position. Academic Figure Generator gives you a generated raster image from a text prompt, so you control the description and the model controls the composition. For a schematic that must match a specific published icon vocabulary, BioRender's approach is the right one. For a conceptual architecture diagram where you can accept the model's layout choices, the prompt route is faster and needs no manual assembly.
The second comparison is using an image model directly, without this project. That is essentially what the skill mode does, and it is a legitimate choice. What the full platform adds over a raw chat with an image model is the paper-parsing step and the project structure: documents, prompts and images grouped per project, with the prompt persisted so you can regenerate or edit it later. If you already write good figure prompts without help, the platform's parsing layer is the only thing you would be paying for in setup time.
A third option worth naming is drawing the figure in code, with matplotlib or TikZ. That gives exact control, reproducible output and vector files, at the cost of writing the layout yourself. Academic Figure Generator is not competing with that; it is aimed at the case where you would rather describe the figure than construct it.
Maintenance, licence and what an upgrade actually costs
The repository is not archived, and the last push was on 2026-03-27. The most recent releases are desktop-v0.1.0-build.34, build.33 and build.32, all dated 2026-03-14. Those build numbers, three releases within about half an hour on the same day, look like iteration on a desktop packaging pipeline rather than a series of feature milestones. The README documents no upgrade procedure, no migration tool and no versioning policy for the SQLite schema, so an upgrade in practice means pulling the repository, reinstalling dependencies in both backend and frontend, and hoping the automatic database creation handles an existing file. That is the kind of thing to test on a copy of backend/data/app.db before doing it on the only copy.
The licence is MIT, stated in the README badge and the LICENSE file at the repository root. MIT is permissive: it allows commercial use and modification with attribution and no warranty. Two things sit outside that grant and are worth separating. The generated images come from a third-party image API under that provider's terms, and the prompt generation goes through Anthropic's SDK under Anthropic's terms. The MIT licence on this repository says nothing about either. If you plan to publish figures in a journal, check the image provider's terms for your own situation rather than assuming the repository licence covers the output. This is not legal advice, and the README does not discuss output ownership at all.
Editorial conclusion
Academic Figure Generator fits a single researcher or a small group who already has an Anthropic key and a NanoBanana or Gemini image key, wants paper-to-figure prompts generated locally, and is willing to run a Python backend and a Vite frontend. It does not fit anyone who needs a hosted service with no keys, or a team expecting a shared multi-user deployment: the README describes a personal local version, the default database is a single SQLite file under backend/data/app.db, and the .env.example hints at Postgres, Redis, Celery and MinIO without any documented migration path. Before committing, verify three things yourself: that Claude Code CLI is installed and logged in on the machine, that your image endpoint returns a result for the model name in NANOBANANA_MODEL, and that a test paper parses cleanly, since the README does not document what happens when PDF extraction mangles a two-column layout. If the figure prompts are all you want, install the academic-figure-prompt skill instead and skip the platform.
Frequently asked questions
Which AI tool is best for creating scientific figures?
The README positions Academic Figure Generator as a local pipeline that reads a paper with Claude and renders the figure through a NanoBanana or Gemini endpoint, which suits single-image diagrams such as architecture, module detail and flow charts. It does not assemble figures from a vetted icon library, so a schematic that must match a specific published icon set is a different kind of tool.
Is there a free app for scientific illustration?
The software is MIT-licensed and free to run, but it requires an ANTHROPIC_API_KEY and a NANOBANANA_API_KEY, both marked as required in the README's environment variable table. There is no offline mode or bundled model, so any cost comes from those two services rather than from the project.
Is there a free version of BioRender available?
The README does not discuss BioRender pricing or a free tier. It does describe Academic Figure Generator as an MIT-licensed alternative path: a generated raster image from a text prompt rather than a drag-and-drop canvas built from a curated icon library.
What is the best AI image generator for this workflow?
The project does not pick one for you. The README defaults NANOBANANA_MODEL to gemini-3-pro-image-preview and NANOBANANA_API_BASE to https://api.keepgo.icu, and both can be overridden in the .env file, so the image model is whatever your endpoint serves.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ligphidonk-academic-figure-generator)