CLI tool
sunbigfly/ppt-agent-skills avatar
sunbigfly/ppt-agent-skills

PPT Agent recovers by reading its own output files instead of a checkpoint

A code-driven presentation generation framework. 像构建软件工程一样生成演示文稿。

902 stars90 forksPythonNOASSERTION

At a glance

What is it?
A presentation framework that runs as an agent skill, with four isolated subagent stages, a per-page visual audit loop and a state machine whose resume point is inferred from artefacts on disk. The exports claim perfect fidelity in absolute terms and no measurement is attached.
Who is it for?
Use this if you hand decks to an assistant regularly and keep getting the same two failures: overlapping elements and text that ignores the slide it is on. The mechanism aimed at those failures is concrete, since every page gets a data contract validated before rendering, then a screenshot that a model reads, then a rewrite of the structure rather than of the spacing.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 119 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The resume point is inferred from files on disk

There is no checkpoint file. The whole flow is described as stateless with respect to progress, and the recovery rule is to scan the artefacts that already exist and work out where to continue:

code
interview-qa.txt → requirements-interview.txt
  → search.txt + search-brief.txt | source-brief.txt
  → outline.txt → style.json
  → planningN.json → slide-N.html → slide-N.png
  → preview.html → presentation-{png,svg}.pptx → delivery-manifest.json

The presence of the outline file means the outline stage finished. The presence of a style file means the style lock is done. A numbered image means that page was rendered. There is no ambiguity about which page you are on, and no ambiguity about what to redo.

The cost of that design is trust. If you edit an outline by hand after a run is interrupted, recovery will treat your edit as the stage output and continue from it, which is convenient until it is not. The final artefact in the chain is a delivery manifest, so the run also ends with a machine-readable record of what was produced rather than only with two files.

Everything is written under a per-run directory keyed by a run identifier, inside an output directory, so concurrent runs do not collide.

A failed gate rolls back one step, not the whole deck

The pipeline is written as stages, and each stage's artefact has to pass a gate before the next one starts:

code
P0 采访   →  P1 分支确认
P2A 联网检索 / P2B 本地资料压缩
P3 叙事大纲  →  P3.5 全局风格锁定
P4 逐页并行生产(Planning → HTML → Visual QA)
P5 Preview + 双 PPTX 导出

The important structural detail is at stage four, which is per page and parallel. Each page goes through planning, then HTML, then the visual audit. A failure rewinds only the current step, and the stated guarantee is that other pages' progress is unaffected.

That is the difference between a gate and a checkpoint. A run that fails on page nine of fifteen keeps pages one to eight, and the design intends you to fix the problem and continue rather than restart the deck. Combined with the disk-based recovery, the two mechanisms cover different failures: the gate handles a bad page inside a live run, and the file scan handles a process that stopped altogether.

The earlier stages are sequential by nature, with a branching point where web research and local material compression are alternatives, and a separate style-lock step between the outline and page production so that every page inherits one visual system.

Four stages, four contexts, and no default model fallback

The multi-agent story is about isolation rather than parallelism. Four stages, named research, outline, style and planning, each run as a separate subagent, and the claim is that their contexts do not contaminate each other.

That is the right shape for this problem. An outline writer that has read eleven pages of prose will start making layout decisions it has no business making, and a style agent that has read the outline will start rewriting the argument. Keeping them apart is what makes a later stage's output trustworthy.

The second half of that bullet is the detail worth knowing. Every subagent must be created with an explicit model parameter, and falling back to a default is prohibited. That is a small implementation rule with a large effect on cost and on reproducibility: you cannot accidentally have one stage running on an expensive model because that was the ambient default, and you cannot have the same deck produced by two different models because a fallback kicked in.

It also means model choice is a per-stage decision rather than a per-project one, which is why the repository later adds a logger that records the stage commands and run logs for the two page-level agents, making a rewind auditable and replayable.

The visual audit rewrites structure, not spacing

Every page's HTML is screenshotted as soon as it is built, and a model reads the screenshot to audit it. When the audit finds a layout overflow, the stated remedy is a rewrite of the DOM and CSS structure rather than a nudge to margins and padding.

That distinction is the whole point of the loop, and it is a correction to how this is usually done. Adjusting spacing moves the symptom; if two elements collide because the layout assumed one line count and got three, the fix belongs in the layout. A pipeline that only adjusts spacing converges on a slightly larger number of nudges, and the next long string breaks it again.

The current version tightened this. The visual checker gained a second layer of assertions alongside the image check, so it now verifies structure and the decoration budget as well as what the screenshot looks like. That means the audit is partly a rules check again, which is the right shape for anything measurable.

The data layer is separated the same way. A page's content is generated as a JSON contract that a planning validator checks before any rendering happens, and a physical structure detector at the HTML level intercepts non-standard skeletons in generated static files.

Two export paths and three absolute claims

The deck is exported twice, and the two formats serve different purposes. A raster path through images is described as guaranteeing visual fidelity across platforms, and a vector path that preserves fonts and stays independently editable in a graphics editor.

Both are the right call for different people, and offering both is unusually thoughtful for a generated deck. The image path is what you send to someone who will not have your fonts; the vector path is what a designer will ask for and what the image path cannot satisfy.

The claims attached to them are where the caution goes. The image export is described as guaranteeing one hundred percent visual restoration across platforms. The screenshot engine is described as abandoning hard delays entirely, hooking font readiness and per-node listeners, and achieving zero lost fonts and zero seconds of waiting. There is no benchmark, no test matrix and no platform list attached to any of it in the visible documentation.

A design philosophy is also stated in absolute terms, promising the model full privilege over typography, overlap depth and whitespace while sealing off absolute physical safety walls underneath. The interesting half is the sealing off, which is the same idea as the validator contracts: constrain what must be true, free everything else.

The layout directory is singular and the repository is plural

The repository structure block is named for a directory that does not match the repository name. It opens with a singular skill folder name, while the repository is the plural form. Everything inside the block is consistent with what is actually at the root, so the mismatch is only in the name used to introduce the tree.

The layout itself is worth reading because it tells you where the token weight of this project sits. A single control file holds the state machine, the gates and the recovery rules. A scripts directory holds the validators, the harness and the exporters. A references directory holds the knowledge that gets mounted on demand rather than loaded up front, split into five kinds: per-stage subagent playbooks, theme style specifications, layout resources, chart templates and interface blocks.

That split is a design decision rather than an accident. The stages are long and procedural, and the reference material is long and not always needed, so keeping them out of the always-loaded context is what makes the per-stage isolation affordable.

Two smaller things sit alongside it. The repository has a Chinese readme and an English one linked from its header, so this material exists in both languages. And the examples section is an empty disclosure widget: the expandable block in the readme contains no rendered output.

The only version marker lives inside the readme

There are no published releases, and there is no version field anywhere in the visible material. The single place a version appears is a dated entry in the readme itself, version 4.1, dated in early April 2026.

The last push came about two months after that date, so the readme's own change log is already behind the working tree. That is normal for a project without release discipline and it matters when you are deciding whether a reported problem is fixed, since there is no tag to compare against.

That entry is also the best description of what changed, and it reads like a response to a specific failure. A density budget became an explicit, checkable constraint rather than a guideline, by completing a bias parameter in the planning stage and a label and contract pair in the page stage. So page density went from something the model was asked to respect to something a validator can fail on.

The last two items in the same entry are a logging addition for the two page-level agents and the dual assertion in the visual checker. Three of the four changes in one entry are about verification rather than about generating slides, which is the clearest signal of where the project thinks its problems were.

Editorial conclusion

Use this if you hand decks to an assistant regularly and keep getting the same two failures: overlapping elements and text that ignores the slide it is on. The mechanism aimed at those failures is concrete, since every page gets a data contract validated before rendering, then a screenshot that a model reads, then a rewrite of the structure rather than of the spacing. Two cautions. The quality claims are absolute, with perfect cross-platform fidelity and no font loss stated without a measurement, so check a rendered deck against your own template before trusting the output on a branded one. And the recovery design has a cost, since inferring the resume point from files on disk means a hand-edited artefact will be trusted, so do not fix an outline by hand mid-run.

Frequently asked questions

What is the PPT agent skills repository?

A code-driven presentation framework that runs as an agent skill with no separate deployment. A strict state machine drives multiple agents, each stage running its own subagent with an isolated context, and a one-line brief becomes a PPTX file. Outputs include a web preview and two PPTX formats.

How do I install the ppt agent skills?

Through the skills CLI, with npx skills add sunbigfly/ppt-agent-skills. There is nothing to deploy: in any agent environment that supports skills you type the requirement and the full pipeline runs. All artefacts are written under a per-run directory inside the output directory, including the preview and both PPTX formats.

What stages does the PPT agent workflow run?

An interview stage, then a branch confirmation, then either web research or local material compression, then a narrative outline, a global style lock, per-page parallel production through planning, HTML and visual QA, and finally a preview with dual export. Each stage's artefact is checked by a gate, and a failure rewinds only the current step without affecting other pages.

Does the PPT agent save progress between runs?

No progress state file is used. After an interruption the flow scans the artefacts already on disk, such as the interview transcript, the outline, the style file and the per-page images, and infers the resume point from which ones exist.

What license does ppt-agent-skills use?

The readme states MIT and links to a licence file at the repository root, while no licence name is recorded for the project. The repository also ships a Chinese and an English readme, with the English one linked from the header of the Chinese one.

Official sources

  1. Issues
  2. README
  3. sunbigfly/ppt-agent-skills on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/sunbigfly-ppt-agent-skills.svg)](https://hysenlabs.com/projects/sunbigfly-ppt-agent-skills)