PPTAgent: An AI Framework for Generating and Revising PowerPoint Presentations
An Agentic Framework for Reflective PowerPoint Generation
At a glance
- What is it?
- PPTAgent is a MIT-licensed Python framework, backed by two peer-reviewed papers, that generates editable PowerPoint files through a coding agent interface such as Claude Code or Codex. It adds a visual review step using a multimodal model and exports actual PPTX files rather than rendered images, making it suited to engineers who want presentation generation inside their existing AI-native workflow.
- Who is it for?
- PPTAgent is a reasonable choice for any developer who already uses Claude Code, Codex, or OpenCode and wants to generate or revise PowerPoint files from within that environment. The system dependencies, particularly LibreOffice and Playwright, add setup overhead that rules it out for anyone wanting a lightweight package.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 9 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What PPTAgent Does and Who It Is For
PPTAgent is a framework for generating PowerPoint presentations through an agentic pipeline. It is not a command-line converter that transforms a document directly into slides. Instead, it operates as a skill registered with a coding agent: the user writes a request in natural language inside Claude Code, Codex, or OpenCode, and PPTAgent handles the presentation structure, content, and visual review.
The project has produced two peer-reviewed publications. PPTAgent, the original system, was accepted to EMNLP 2025 and is available at code version v0.2.0. DeepPresenter, an expanded version with free-form visual design and text-to-image generation, was accepted to ACL 2026 and is available at code version v1.1.38. The current development branch is v2.x and introduces integration with the Atria Dawn Preview model.
The primary audience is developers and researchers who already work inside an AI coding assistant and want presentation generation to be part of that workflow, not a separate application. The requirement for LibreOffice, a Python virtual environment, and Playwright means that non-technical users are not the intended audience.
Installing the PPTAgent Skill on Claude Code or Codex
Installation is split into two phases: setting up system dependencies and registering the skill with the coding agent.
On Debian or Ubuntu, install the required system packages:
sudo apt-get install npm libreoffice-impressOn macOS, install the dependencies and create the expected path for LibreOffice:
brew install node
brew install --cask libreoffice google-chrome
ln -sf "$(command -v soffice)" "$(brew --prefix)/bin/libreoffice"Next, clone the repository, navigate to the skill directory, and create the Python environment:
git clone https://github.com/icip-cas/PPTAgent.git
cd PPTAgent/skills/pptagent
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python -r requirements.txt
.venv/bin/python -m playwright install --with-deps chromiumThe skill uses uv for the Python virtual environment. Python 3.12 is specified explicitly. Playwright installs a Chromium browser used for rendering slide previews.
To register the skill with Claude Code:
.venv/bin/python scripts/install.py --client claudeFor Codex:
.venv/bin/python scripts/install.py --client codexFor OpenCode:
.venv/bin/python scripts/install.py --client opencodeAfter registration, verify the installation:
.venv/bin/python scripts/pptagent.py doctorThe doctor command checks that all dependencies are available and that the skill is correctly registered with the selected coding agent.
The Text and Multimodal Review Pipelines
PPTAgent supports two generation modes. In text mode, a language model writes the slides and a separate multimodal vision model reviews rendered images of each slide to identify layout or readability issues. In multimodal mode, a single model capable of processing images handles both writing and visual review.
Configuring text mode requires a config.yaml file in the skills/pptagent/ directory. The README provides an example configuration that connects to a vision model through an OpenAI-compatible API:
mode: text
visual:
base_url: "<OpenAI-compatible API base URL from the Duanyan console>"
model: "deepseek-v4-flash-vision"
api_key_env: VISUAL_API_KEY
timeout_seconds: 300
delivery:
mode: strictThe README recommends `deepseek-v4-flash-vision` as the visual review model. The API key is stored in a .env file:
VISUAL_API_KEY=<your-duanyan-api-key>The README notes that the skill appends `/chat/completions` to the `base_url` value, and instructs users to provide the API prefix from the console rather than the Token Plan webpage URL. This distinction matters because the two URLs have different path structures.
The delivery.mode setting controls how strictly the visual review results gate the output. The `strict` setting requires all review checks to pass before the presentation is delivered.
Launching Claude Code with the Atria Dawn Preview Model
The README documents a quick-start workflow using the Atria Dawn Preview model, which is a jointly released model from Shanghai AI Laboratory and partner institutions. To use it with Claude Code, set the Atria API base URL and key as environment variables:
export ATRIA_API_KEY="<your-atria-api-key>"
mkdir -p ~/pptagent-demo
cd ~/pptagent-demoThen launch Claude Code pointing to the Atria endpoint:
ANTHROPIC_BASE_URL=https://api.atria-asi.ai \
ANTHROPIC_AUTH_TOKEN="$ATRIA_API_KEY" \
claude --model Atria-Dawn-PreviewThis approach routes Claude Code's API calls through Atria rather than Anthropic's servers. The Atria API is OpenAI-compatible, so it can also be used with Codex by adding a `[model_providers.atria]` section to `~/.codex/config.toml`.
The README offers free API token access through two programs: Discovery and Atria API, both linked from the repository. The free tier is intended for exploration rather than production workloads.
Research Basis: PPTAgent and DeepPresenter Publications
The project has two distinct research identities. PPTAgent, published at EMNLP 2025, introduced the reflective generation framework: the agent generates slides, reviews them against the source content, and revises. The paper is at arxiv.org/abs/2501.03936.
DeepPresenter, published at ACL 2026, extended the framework with several additions. The README lists deep research integration, free-form visual design, autonomous asset creation, text-to-image generation, and an agent environment with a sandbox and more than 20 tools. Fine-tuned models and the taskset from DeepPresenter are published on Hugging Face under the ICIP collection.
For a developer who wants to reproduce results from either paper, the README provides pinned version references: v0.2.0 for PPTAgent and v1.1.38 for DeepPresenter. The current development branch is not intended for paper reproducibility and includes incompatible changes.
Where PPTAgent Has Limitations
The system dependencies are substantial. LibreOffice must be installed and accessible as `libreoffice` on the PATH. On macOS this requires a manual symlink from the LibreOffice binary to that name. A Chromium browser is installed by Playwright during setup. Node.js and npm must also be present. Any environment where these packages are restricted or unavailable will not run PPTAgent.
The visual review step requires access to a multimodal vision model API. The README suggests applying for a free token through the Duanyan token plan, but this is a third-party service outside the repository. A team that cannot use external APIs cannot use the visual review feature, which reduces the quality gate to the text model's self-assessment only.
The repository is structured around the PPTAgent skill in skills/pptagent/. The earlier runtime code from the papers is at pinned tags. Running the development branch for production use means accepting that it diverges from the published research without a new peer review.
Offline mode is listed in the changelog for January 2026 as supported for freeform and template generation with PPTX export. However, the visual review step requires a live API call and does not work offline.
Comparing with python-pptx Scripting
python-pptx is a widely used Python library for creating and modifying PowerPoint files programmatically. It provides direct access to the OOXML structure of .pptx files, which gives complete control over every element: text boxes, shapes, charts, images, and transitions. The difference in approach is the level of abstraction. python-pptx requires the developer to write explicit slide layout code. PPTAgent generates the content and structure from a natural language prompt and the agent's reasoning, without the developer specifying each element.
For a developer who needs precise control over slide geometry, branding rules, or a fixed template structure, python-pptx is the more direct tool because it maps one-to-one to the OOXML format. PPTAgent is better suited to generating first-draft content where the structure should emerge from the content rather than a predefined layout.
Maintenance and License
The most recent release in the repository is v1.1.38, tagged on 2026-09-20. The last push to the repository was on 2026-09-14. The project is MIT-licensed, which permits unrestricted use, modification, and redistribution.
The repository includes a CLAUDE.md and an AGENTS.md, which are used by coding agents that clone the repository to understand its conventions. The pre-commit configuration in .pre-commit-config.yaml runs formatting and lint checks on commits.
The README points contributors to the skills/pptagent/ directory for the current implementation and to the pinned tags for the research code. Issues about the development branch and issues about reproducing paper results should be filed separately, as the two codebases have diverged.
Editorial conclusion
PPTAgent is a reasonable choice for any developer who already uses Claude Code, Codex, or OpenCode and wants to generate or revise PowerPoint files from within that environment. The system dependencies, particularly LibreOffice and Playwright, add setup overhead that rules it out for anyone wanting a lightweight package. The two research publications at EMNLP 2025 and ACL 2026 establish that the generation quality has been evaluated against academic benchmarks, which is more than most scripting tools offer. Before installing, verify that LibreOffice is available as `libreoffice` on the PATH, that uv is installed, and that Node 22 or later is present.
Frequently asked questions
What is PPT Agent?
PPTAgent is a Python framework that generates and revises editable PowerPoint files through an agentic loop. It registers as a skill in coding agents such as Claude Code, Codex, and OpenCode, allowing users to request presentations in natural language and receive .pptx output with an optional visual review step.
Which coding agents does PPTAgent support?
The README lists Claude Code, Codex CLI, and OpenCode as supported clients. The install script at scripts/install.py accepts a --client flag for each. An MCP server configuration is also documented in the v1.1.37 release notes.
Does PPTAgent work without an internet connection?
The README states that freeform and template generation with PPTX export and offline mode were added in January 2026. The visual review step requires a live API call to a multimodal vision model and does not function offline.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/icip-cas-pptagent)