Model or dataset
llmsresearch/paperbanana avatar
llmsresearch/paperbanana

PaperBanana: generating academic diagrams through a two-phase agent pipeline

Open source implementation and extension of Google Research’s PaperBanana for automated academic figures, diagrams, and research visuals, expanded to new domains like slide generation.

2,384 stars338 forksPythonMIT

At a glance

What is it?
llmsresearch/paperbanana is a community implementation of Google Research's PaperBanana paper, packaged as a Python CLI, a Python API and an MCP server. It turns text descriptions of a method into diagrams and statistical plots, and it is honest about being unofficial.
Who is it for?
Adopt paperbanana if you already write method sections in plain text and want a first draft diagram you will edit afterwards, and if you are comfortable with a project the pyproject.toml still classifies as Development Status :: 3 - Alpha.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 14 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What PaperBanana solves for people writing method sections

The problem is narrow and real. A methods section describes a pipeline in prose: an encoder, a retrieval step, a fusion module, a decoder. Turning that prose into a figure that a reviewer will accept means either learning a diagram tool or handing a rough sketch to a designer. PaperBanana takes the prose and produces an image. The README describes it as an agentic framework for generating publication-quality academic diagrams and statistical plots from text descriptions. The intended audience is stated in the project's own subtitle: AI scientists. That means people who can describe their architecture in a text file, not people who want a general-purpose illustration tool. The repository also lists slide generation among its topics, which suggests the maintainers see the same pipeline serving talks as well as papers. Two things shape how you should read the rest of this page. First, the README carries an explicit disclaimer that this is an unofficial, community-driven implementation of the paper by Dawei Zhu, Rui Meng, Yale Song, Xiyu Wei, Sujian Li, Tomas Pfister and Jinsung Yoon, and that it is not affiliated with or endorsed by the original authors or Google Research. Second, the pyproject.toml classifies the package as Development Status :: 3 - Alpha. Neither is a defect. Both are constraints on when you should reach for it.

The two-phase agent pipeline and where your API key fits

The README describes a two-phase multi-agent pipeline with iterative refinement, plus an input optimization layer that rewrites your description before generation. What the documentation does not do is publish the internal prompt chain or the exact model calls per phase, so anyone expecting to audit the agent graph will be disappointed. What is visible is the provider surface. A VLM handles the language side and an image model handles the pixels, and they are configured separately: VLM_PROVIDER and IMAGE_PROVIDER select the pair, while OPENAI_VLM_MODEL, OPENAI_IMAGE_MODEL, GOOGLE_VLM_MODEL and GOOGLE_IMAGE_MODEL name the specific models. The .env.example defaults to gemini-2.5-flash for the VLM and gemini-3-pro-image-preview for images, with gpt-5.2 and gpt-image-1.5 as the OpenAI pair. That split matters because the two halves have different failure modes. A VLM that misreads your description produces a confidently wrong layout; an image model that ignores layout instructions produces something pretty and useless. Auto-refine mode and run continuation with user feedback exist precisely because the first pass is usually not the last. The continuation flow is the more interesting design choice: instead of rerunning from scratch, you feed feedback into an existing run. The README does not document rollback, so if a refinement makes things worse you are comparing output files by hand.

Installing paperbanana and generating your first diagram

The package requires Python 3.10 or newer. The README gives a single-command install from PyPI, which pulls the core dependencies listed in pyproject.toml (pydantic, typer, matplotlib, pandas, pillow and others). Provider SDKs are optional extras, so a bare install will not give you the Gemini or OpenAI client.

bash
pip install paperbanana

For development, or when you want the provider extras in one step, the README clones the repository and installs it editable with the dev, openai and google groups. Each extra resolves to a pinned range in pyproject.toml, for example openai>=1.0 and google-genai>=1.65.

bash
git clone https://github.com/llmsresearch/paperbanana.git
cd paperbanana
pip install -e ".[dev,openai,google]"

Credentials go in a .env file copied from the example. The example shows the Gemini key as the free path, with an OpenAI key and an Azure OpenAI or Foundry base URL as alternatives. Only fill in the provider you intend to use; the other keys can stay empty.

bash
cp .env.example .env
# Edit .env and add your API key:
#   OPENAI_API_KEY=your-key-here
#   GOOGLE_API_KEY=your-key-here

There is also a setup wizard for Gemini, which is the shortest path if you have not configured a key before. After that, generation takes an input text file and a caption. The README points at examples/sample_inputs/transformer_method.txt as the sample input, so you can confirm the whole chain works before writing your own description.

bash
paperbanana setup

paperbanana generate \
  --input examples/sample_inputs/transformer_method.txt \
  --caption "Overview of our encoder-decoder architecture"

If you would rather not install anything locally, the README links a Colab quickstart notebook that walks through install, API key and diagram generation end to end. For container users, the Dockerfile builds from a clone and runs as a non-root user with /work as the writable directory. Note that the image installs the google, openai and pdf extras, and that the wheel embeds prompts/, data/ and configs/ so the build context is deleted after install.

bash
docker build -t paperbanana .
docker run --rm -e GOOGLE_API_KEY \
  -v "$(pwd)/method.txt:/work/method.txt:ro" \
  -v "$(pwd)/outputs:/work/outputs" \
  paperbanana generate --input method.txt --caption "Overview of our framework"

Batch runs, plots and the Studio interface

Single diagrams are the demo. The batch features are what make the tool usable across a paper. A manifest file in YAML or JSON drives multiple diagrams in one run, and examples/batch_manifest.yaml plus examples/composite_batch_manifest.yaml show the expected shape. There is a separate path for statistical plots: paperbanana plot-batch reads a manifest where each item points at CSV or JSON data, which is a different job from generating a schematic. The README also documents PDF inputs for methodology context through the optional paperbanana[pdf] extra backed by PyMuPDF, with per-page selection, so you can point the tool at a draft rather than retyping your method. PaperBanana Studio is a local Gradio web UI started with paperbanana studio, covering diagrams, plots, evaluation, batch and a run browser. The evaluation surface is worth noting: the repository ships a bench-data-v1 release described as a PaperBananaBench dataset mirror, which means the project has a benchmark to measure against rather than only a demo gallery.

Where paperbanana is the wrong tool

Image generation is not deterministic. The same input and caption can produce different figures across runs, and the README offers no seed control or reproducibility guarantee. For a camera-ready figure that must match a specific layout, that is disqualifying, and you should expect to redraw the output in a vector tool anyway. The second constraint is data egress. Every provider in the .env.example that is configured by default (OpenAI, Azure OpenAI, Gemini, Atlas Cloud) is a hosted service, so your method description leaves your machine. The file does list local options, ollama with OLLAMA_BASE_URL and openai_local for vLLM or llama.cpp, but the README does not document quality expectations for those paths, and a local vision model is unlikely to match a hosted one on layout reasoning. If your work is under embargo or your institution forbids sending unpublished methods to third parties, check that section carefully before you start. Third, the project is alpha and unofficial. The README states the implementation is based on the publicly available paper and may differ from the original system. If you need the exact system the paper describes, this is not it. Fourth, the tool generates figures, not the underlying data. A plot batch needs real CSV or JSON behind it; paperbanana will not invent your results, and you should not want it to.

paperbanana against nano banana and TikZ export

The comparison people search for is paper banana vs nano banana, and the difference is scope rather than model. Nano Banana is an image generation model; paperbanana is a pipeline that wraps image models, including Gemini image models, with a VLM stage that reads your method description and an input optimization layer that rewrites it. If you already have a prompt you like and you just want an image, calling the image model directly is fewer moving parts and one less API key to manage. PaperBanana earns its place when the input is a methods section rather than a prompt, and when you want the same description to drive several figures through a manifest. The other alternative is programmatic diagramming: TikZ, Mermaid or Graphviz, where the figure is source code you can diff and review. The repository acknowledges this path by shipping examples/tikz_export.py, so the two approaches are not mutually exclusive. TikZ gives you byte-identical output on every build and a figure that survives a journal's production process. PaperBanana gives you a starting layout in seconds from prose you already wrote. For a paper with a hard submission deadline and a strict style guide, generate the draft here and commit the TikZ.

Maintenance, licence and upgrade cost

The last push was on 2026-09-09, which is recent, and the repository is not archived. Releases are less frequent: v0.3.0 landed on 2026-06-12, one day after v0.2.0, and the bench-data-v1 dataset mirror arrived on 2026-06-11. That pattern suggests bursty development, with feature work landing in a short window and the repository seeing smaller commits since. The licence is MIT, declared in both the LICENSE file and the pyproject.toml, which is permissive and places few restrictions on how you use or redistribute the code. Two caveats belong here rather than in a legal opinion. MIT covers the code, not the images you generate; those are governed by your provider's terms, and the .env.example shows several different providers with presumably different terms. The README's own disclaimer about non-affiliation with Google Research is also worth keeping in mind if you cite the tool in a paper. Upgrade cost is low in the normal case, since the package is a CLI and a Python API rather than a framework you embed, but the provider model names are pinned in .env.example and will need revisiting as vendors retire preview models.

Editorial conclusion

Adopt paperbanana if you already write method sections in plain text and want a first draft diagram you will edit afterwards, and if you are comfortable with a project the pyproject.toml still classifies as Development Status :: 3 - Alpha. Do not adopt it if you need deterministic, reproducible figures for a camera-ready submission, if you cannot send your method description to a hosted model provider, or if you expected an official Google Research release, because the README states plainly that this is an unofficial community implementation. Before committing, verify three things yourself: that your chosen provider returns a usable image for one of the files in examples/sample_inputs, that the run continuation flow with user feedback produces a second iteration you actually prefer, and that the MIT licence plus your provider's terms cover how you intend to publish the output. The last push was on 2026-09-09 and the most recent release is v0.3.0 from 2026-06-12.

Frequently asked questions

What is PaperBanana?

It is an open source, community-driven implementation of the PaperBanana paper from Google Research, described in the README as an agentic framework for generating publication-quality academic diagrams and statistical plots from text descriptions. It ships as a Python CLI, a Python API and an MCP server.

How do I use paperbanana?

Install it with pip install paperbanana, put a provider key in a .env file copied from .env.example, then run paperbanana generate with an input text file and a caption. The README points at examples/sample_inputs/transformer_method.txt as a working sample input.

Is paperbanana free?

The code is MIT licensed and the package itself is free to install. Generation is not free in general, because it calls a hosted model provider, though the README notes the Google Gemini API key is free and the .env.example points at Google AI Studio to obtain one.

How do I access paperbanana?

There are four documented routes: pip install paperbanana for local use, the Colab quickstart notebook linked from the README, the Docker image built from the repository's Dockerfile, and the HuggingFace Space demo. The local Gradio interface is started with paperbanana studio.

What is paper banana vs nano banana?

Nano Banana is an image generation model, while paperbanana is a pipeline that puts a VLM stage and an input optimization layer in front of image models. The README lists Google Gemini among the supported providers, so the two can sit in the same stack rather than competing.

Official sources

  1. Issues
  2. License: MIT
  3. llmsresearch/paperbanana on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/llmsresearch-paperbanana.svg)](https://hysenlabs.com/projects/llmsresearch-paperbanana)