# Banana Slides: AI-Native Presentation Generator with Conversational Editing and PPTX Export

> Banana Slides is a self-hosted, AGPL-3.0 application that generates editable PowerPoint presentations from a single prompt, an outline, or uploaded documents. It connects to a configurable LLM provider and uses the nano banana model to generate each slide's visual layout rather than filling a fixed template.

**Anionex/banana-slides** — 一站式原生AI PPT生成应用，几分钟内生成一套幻灯片; 支持上传任意模板图片，上传任意素材&智能解析，一句话/大纲/页面描述自动生成PPT，口头修改指定区域、一键导出可编辑ppt、视频等 - An AI-native slides generator based on nano banana pro🍌

- Repository: https://github.com/Anionex/banana-slides
- Website: http://bananaslides.online
- Stars: 15,674 · Forks: 1,789
- Language: TypeScript
- License: AGPL-3.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/anionex-banana-slides

## The Problem Banana Slides Addresses and Who It Is For

Traditional AI presentation generators follow one pattern: they fill predefined visual templates with LLM-generated text. The README identifies five specific failures in that approach: users can only pick from preset templates, multi-round edits are hard to execute, the output looks the same across users, materials used are generic, and image-text layout feels disconnected.

Banana Slides targets several groups. Students and professionals who need to produce a presentation on short notice and want to focus on content, not layout, are the primary audience. The README also identifies PPT specialists looking for design inspiration, educators converting course content into slide decks, and small-business owners needing quick visual proposals.

The technical answer to the template limitation is the nano banana model, described in the repository as the foundation that makes layout-native generation possible. Rather than inserting LLM text into a fixed theme, Banana Slides instructs the model to generate the visual composition of each slide according to a style prompt and optional reference images you supply. This approach can produce slides whose layout matches a reference image rather than matching a template author's predefined choices.

## How Banana Slides Generates Slides: Three Input Modes

The application supports three starting points: a one-sentence idea, an outline, or individual page descriptions. Each maps to a different level of control over the final structure.

The one-sentence path asks the AI to generate the entire outline and per-page descriptions automatically. The outline path lets users define the structure first, then have AI generate page-level content for each section. The page-description path provides maximum control, letting users write the descriptive intent for each slide individually. All three paths funnel into the same generation step where the nano banana model renders each slide as a composed image.

After generation, the conversational editing feature allows spoken-language modifications: a request like "change the third page to a case study" is interpreted by the AI and applied to the outline or description before a regeneration. The README describes this as a Vibe-PPT approach, meaning the user directs intent rather than manipulating layout controls directly.

The backend uses a Python/Flask architecture. Dependencies in pyproject.toml include python-pptx for assembly of the final PPTX file, Pillow and OpenCV for image processing, and edge-tts and ElevenLabs for narration in exported videos. The frontend is TypeScript. The project requires Python 3.10 or newer.

## Deploying Banana Slides with Docker Compose

The primary deployment method is Docker Compose. The docker-compose.yml defines two services: a Python backend and a TypeScript frontend. By default the backend binds to port 5011 on the host and the frontend to port 3011:

```yaml
services:
  backend:
    ports:
      - "${BACKEND_PORT:-5011}:5000"
  frontend:
    ports:
      - "${FRONTEND_PORT:-3011}:80"
```

Before starting, create a `.env` file from the provided `.env.example` and set the AI provider. The minimum configuration for Gemini looks like this:

```
AI_PROVIDER_FORMAT=gemini
GOOGLE_API_KEY=your-api-key-here
GOOGLE_API_BASE=https://generativelanguage.googleapis.com
```

Once `.env` is configured, start the service:

```bash
npm run start
```

The package.json expands this to `docker compose up -d`. The backend health check polls `http://localhost:5000/health` every 30 seconds with a 10-second timeout and three retries. The frontend container depends on the backend being healthy before starting.

For development without Docker, run the backend directly:

```bash
npm run dev:backend
```

This executes `cd backend && uv run python app.py`. A separate terminal runs the frontend with `npm run dev:frontend`, which starts the Vite/TypeScript dev server.

## Conversational Editing, Template Control, and PPTX Export

Four editing mechanisms are documented in the README. Conversational editing applies spoken-language changes to the outline or individual page descriptions in real time. The material toolbox, released in April 2026, adds three image editing modes: whole-image editing, region selection (overlay or replace), and smart erasure. Template control, added in June 2026, supports both a unified template across all slides and per-page template assignment. Users can upload image or PDF files to build a project template library, and the AI parses the uploaded file's visual style.

Export produces a PPTX file using the python-pptx library. The README notes an option in settings under Export Options for how backgrounds are fetched, with a generative fetch mode available. Alongside PPTX, a video narration export path exists, using edge-tts or ElevenLabs for voice generation; the top-level repository contains a TODO file (TODO_video_narration_fixes.md) indicating this feature was under active refinement as of the most recent commits.

The CLI path, documented at docs.bananaslides.online/cli, exposes slide generation commands for integration with agent workflows. The project includes a `skills/` directory in the repository root, suggesting it supports agent-skill invocation from external systems.

## LLM Provider Configuration: Thirteen Providers via LazyLLM

The .env.example file shows five top-level provider format modes: gemini, openai, volcengine, anthropic, and lazyllm. The lazyllm mode activates a provider-abstraction layer that the README describes as supporting eleven providers in the desktop-packaged version: Qwen, Doubao, DeepSeek, GLM, Kimi, MiniMax, SenseNova, SiliconFlow, PPIO, Aiping, and OpenAI.

The pyproject.toml lists the Python SDK packages for each: dashscope (Qwen), zhipuai (GLM), volcengine-python-sdk (Doubao via Volcengine), anthropic (Claude), google-genai (Gemini), openai, and lazyllm. This breadth is unusual for a presentation tool; it reflects a design choice to be LLM-agnostic rather than locked to one provider's pricing.

A practical limitation of this breadth is configuration complexity. Each provider uses different env variable names, different base URLs, and different auth patterns. The .env.example documents all of them, but a new deployment requires identifying which provider will handle text generation and which will handle image generation separately, since those are different steps in the pipeline. Selecting a provider that lacks image generation support breaks the visual output.

Version 0.9.0 RC7, released on 2026-09-05, fixed a Codex (OpenAI OAuth) image generation issue by updating the internal text model from gpt-5.4 to gpt-5.6-terra while keeping the image model as gpt-image-2.

## Template-Based Alternatives, Limitations, and AGPL-3.0 Licensing

The established alternative approach is a tool that fills pre-configured slide templates (PowerPoint masters or HTML themes) with LLM-generated text. Gamma is an example of this category: it maps content to predefined card-style layouts. Banana Slides differs by generating the visual composition itself using the nano banana model, without requiring you to select a template theme first. The trade-off is generation time and dependency on a capable image model. Template-fill tools are faster because they skip image generation; Banana Slides requires a model that can render a stylistically consistent image per slide.

A concrete limitation is the AGPL-3.0 license. If you modify Banana Slides and run the modified version as a network service that outside users access, the license requires you to make the modified source code available. Internal deployment within a company for its own employees is a different case, but any externally facing service built on modifications triggers the copyleft requirement.

A second limitation is the video narration path: the TODO_video_narration_fixes.md file in the repository root indicates this feature was not complete as of the most recent commits. Audio-narrated video export may not work reliably in the current release.

The last push to the repository was on 2026-09-27. Version 0.9.0 RC7 was released on 2026-09-05, with RC6 on 2026-08-30 and RC5 on 2026-08-29. The project has not yet reached a stable 1.0 release, and the RC designation means API surfaces and configuration formats may change between releases.

## Conclusion

Teams that need a self-hosted pipeline for generating editable PPTX files from text prompts will find Banana Slides deployable after a Docker Compose setup and an API key configuration. The AGPL-3.0 license requires publishing source modifications if you run the app as a network service for outside users, so operators considering a commercial deployment should read the license terms before committing. Verify that your chosen LLM provider supports image generation at the quality level you need, since Banana Slides delegates visual composition to the nano banana model and the results depend directly on the configured provider's image generation capability.

## FAQ

### Does Banana Slides export real, editable PowerPoint files?

Yes. The backend uses the python-pptx library to assemble the final output, which produces a .pptx file that opens in Microsoft PowerPoint or LibreOffice Impress. The README notes that a generative background option is configurable in the export settings for richer background images.

### Which LLM providers does Banana Slides support?

The .env.example documents support for Gemini, OpenAI (including Codex via OAuth), Anthropic (Claude), Volcengine (Doubao), and a LazyLLM abstraction layer that covers Qwen, DeepSeek, GLM, Kimi, MiniMax, SenseNova, SiliconFlow, PPIO, and Aiping. Text generation and image generation are configured separately since different providers may handle each step.

### Can Banana Slides be run entirely without the hosted demo?

Yes. The repository includes a Docker Compose file that runs the backend and frontend as separate containers on configurable ports (defaults: 5011 for the backend, 3011 for the frontend). There is also an all-in-one Dockerfile (Dockerfile.allinone) in the repository root for single-container deployments.

## Sources

- [Anionex/banana-slides on GitHub](https://github.com/Anionex/banana-slides)
- [License: AGPL-3.0](https://github.com/Anionex/banana-slides/blob/main/LICENSE)
- [Project website](http://bananaslides.online)
- [README](https://github.com/Anionex/banana-slides/blob/main/README.md)
- [Releases](https://github.com/Anionex/banana-slides/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/anionex-banana-slides
