# Paper2Poster: A Multi-Agent Pipeline That Turns a Paper PDF into an Editable Conference Poster

> Paper2Poster generates academic conference posters from paper PDFs using PosterAgent, a three-stage multi-agent system that parses the paper into a structured asset library, plans a binary-tree layout preserving reading order and spatial balance, and refines each panel through a Painter-Commentor loop that uses VLM feedback to eliminate overflow and alignment errors. The output is an editable PPTX file.

**Paper2Poster/Paper2Poster** — [NeurIPS 2025] Open-source Multi-agent Poster Generation from Papers

- Repository: https://github.com/Paper2Poster/Paper2Poster
- Website: https://paper2poster.github.io/
- Stars: 3,971 · Forks: 287
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/paper2poster-paper2poster

## The Problem: Poster Creation from Papers Is Time-Consuming and Layout-Intensive

Academic conference posters require taking a paper's key figures, results, and text, fitting them into a constrained visual layout, balancing whitespace, and producing a file that prints at large format (commonly 48 by 36 inches). Most researchers produce posters manually in PowerPoint or a design tool, a process that takes hours and requires separate decisions about which content to include and how to arrange it.

Paper2Poster addresses the entire pipeline: from paper.pdf to poster.pptx without manual layout work. The README describes this as the question the project addresses: how to create a poster from a paper, and how to evaluate whether the poster is good.

The system is called PosterAgent and is described in a NeurIPS 2025 Dataset and Benchmark Track paper. The same repository also contains the Paper2Poster benchmark, which pairs generated posters with human-authored ones for evaluation across metrics including Visual Quality, Textual Coherence, VLM-as-Judge, and PaperQuiz. PaperQuiz is described as a novel evaluation that assumes a good poster should convey core paper content visually.

## How PosterAgent Works: Parser, Planner, and Painter-Commentor

The README describes PosterAgent as a top-down, visual-in-the-loop multi-agent system. It operates in three stages.

The Parser reads the paper PDF and distills it into a structured asset library: figures, tables, equations, key text passages, and metadata like the paper title, authors, and abstract. This stage determines what material is available for the poster.

The Planner takes the asset library and produces a binary-tree layout. The layout preserves reading order and spatial balance, assigning text and visual pairs to panels in a hierarchical structure. The README describes this as aligning text-visual pairs into a binary-tree layout.

The Painter-Commentor loop then refines each panel. The Painter generates rendering code for a panel; the Commentor uses a vision language model to inspect the rendered result and identify overflow, misalignment, or other issues. The loop iterates until the panel meets quality criteria. This visual feedback step is what the README calls visual-in-the-loop: the system checks its own output visually rather than relying on layout constraints alone.

The result is a poster.pptx file in which each panel is an editable PowerPoint element. Researchers can open the file and adjust the layout, font sizes, or content placement before submitting to a conference.

## Installation and Running the Pipeline

The full pipeline requires Python, LibreOffice, Poppler, and an API key (for cloud-based models) or a local vLLM deployment:

```bash
pip install -r requirements.txt
```

```bash
sudo apt install libreoffice
```

```bash
conda install -c conda-forge poppler
```

Create a .env file in the project root:

```bash
OPENAI_API_KEY=<your_openai_api_key>
```

The directory structure for a paper is simple: create a folder named after the paper under the dataset directory and place the PDF there as paper.pdf.

High-performance generation using GPT-4o for both the language model and vision model:

```bash
python -m PosterAgent.new_pipeline \
    --poster_path="${dataset_dir}/${paper_name}/paper.pdf" \
    --model_name_t="4o" \
    --model_name_v="4o" \
    --poster_width_inches=48 \
    --poster_height_inches=36
```

Economic mode using a local Qwen model for text and GPT-4o for vision:

```bash
python -m PosterAgent.new_pipeline \
    --poster_path="${dataset_dir}/${paper_name}/paper.pdf" \
    --model_name_t="vllm_qwen" \
    --model_name_v="4o" \
    --poster_width_inches=48 \
    --poster_height_inches=36 \
    --no_blank_detection
```

Fully local generation with both models running via vLLM uses vllm_qwen for model_name_t and vllm_qwen_vl for model_name_v. Parallel section generation is enabled by adding --max_workers to the command.

Docker is also supported. Build the image:

```bash
docker build -t paper2poster .
```

The Dockerfile sets up the CUDA-enabled environment, installs LibreOffice and Poppler, and uses an entrypoint script that writes the OPENAI_API_KEY from the environment into the .env file at runtime.

## The Lightweight Agent SKILL for Codex and Claude Code

In addition to the full pipeline, the repository introduced a lightweight Paper2Poster SKILL in June 2026 (documented in the README update log). Located in skills/, it works without installing the pipeline's dependencies and is designed for AI agents like Codex or Claude Code.

The skill helps AI agents prepare poster-ready content from a paper: identifying key contributions, structuring sections, and producing content in a format suitable for use with PosterAgent or manual poster assembly. The README gives this example invocation in Codex:

```
Use $paper2poster-poster to turn this paper into a poster package.
```

More usage examples are in skills/README.md. The SKILL is intended for contexts where the full Python pipeline is not available or not necessary, such as writing assistance or content organisation without automated layout generation.

## Limitations and Cases Where Paper2Poster Is the Wrong Tool

The pipeline's dependency list in requirements.txt is long and includes deep learning libraries. The Dockerfile base image is nvidia/cuda:12.6.0-devel-ubuntu24.04, making GPU availability the expected environment for production use. Running the full pipeline on a machine without a GPU or without CUDA will require configuration changes not documented in the README.

The default configuration uses the OpenAI API. Each poster generation run makes multiple API calls for parsing, planning, and the Painter-Commentor loop, so the cost per poster depends on the number of panels and iterations. The README does not document estimated cost per run.

The output is a PPTX file. Researchers who work in LaTeX-based poster tools (Beamer poster, tikzposter) will need to treat the PPTX as a reference rather than a deliverable, since there is no documented export path to LaTeX.

The poster layout is determined algorithmically. For papers with unusual visual structures (many small figures, complex tables, multi-column layouts), the binary-tree layout planner may not produce a result that matches the visual hierarchy a human designer would choose. The Painter-Commentor loop refines alignment and overflow but does not redesign the layout.

The last push to the repository was on 2026-06-08. The project does not have GitHub releases; the main branch is the current version.

## Maintenance, Benchmark Dataset, and Licence

The last push to the repository was on 2026-06-08. The repository is not archived. The Paper2Poster benchmark dataset is hosted on Hugging Face at Paper2Poster/Paper2Poster and contains human-authored poster and paper pairs for evaluation. A Gradio demo is available, and Docker support was added in October 2025 per the README update log.

The project was accepted to NeurIPS 2025 Dataset and Benchmark Track and has a corresponding arXiv paper at arxiv.org/abs/2505.21497. The Paper2Video follow-up project was published in October 2025 and extends the approach to video generation from papers.

Paper2Poster is MIT-licensed. The repository includes a demo/ directory with a Gradio application, a PosterAgent/ directory for the main pipeline code, and Paper2Poster-eval/ for evaluation tooling. The camel/ directory appears to contain CAMEL framework components used in the multi-agent implementation.

## Conclusion

Paper2Poster is a practical tool for researchers who regularly produce conference posters and want to automate the first draft. The pipeline produces an editable PPTX rather than a locked image, so the output is a starting point rather than a finished product. It requires an OpenAI API key by default (or a locally deployed vLLM instance for open-source models), LibreOffice, and Poppler, making the dependency footprint non-trivial. The last push was on 2026-06-08. Teams running on CPU-only infrastructure or wanting to avoid external API dependencies should use the local Qwen model configuration rather than GPT-4o. The benchmark dataset on Hugging Face pairs generated posters with human-authored references for evaluation, which is useful for anyone who wants to measure output quality rather than assess it visually.

## FAQ

### How do I use Paper2Poster?

Install the dependencies with pip, apt, and conda as documented in the README, create a .env file with your OpenAI API key, place your paper.pdf in a named folder under the dataset directory, and run python -m PosterAgent.new_pipeline with the poster path and size arguments. Docker is also supported.

### What models does Paper2Poster support besides GPT-4o?

The pipeline supports flexible LLM and VLM combinations. The README documents Qwen-2.5-7B-Instruct for the language model and a Qwen VL model for the vision model, both deployed via vLLM. Any model accessible through vLLM can be configured in the get_agent_config() function in utils/wei_utils.py.

### Does Paper2Poster produce an editable output or a fixed image?

The output is an editable poster.pptx file. Each panel is a PowerPoint element that can be resized, repositioned, or edited after generation. The README describes this as a key design goal: the poster is a starting point, not a locked deliverable.

## Sources

- [Issues](https://github.com/Paper2Poster/Paper2Poster/issues)
- [License: MIT](https://github.com/Paper2Poster/Paper2Poster/blob/main/LICENSE)
- [Paper2Poster/Paper2Poster on GitHub](https://github.com/Paper2Poster/Paper2Poster)
- [Project website](https://paper2poster.github.io/)
- [README](https://github.com/Paper2Poster/Paper2Poster/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/paper2poster-paper2poster
