# PaperOrchestra skill pack: run a five-agent paper pipeline inside your coding agent

> Ar9av/PaperOrchestra turns the prompts from the PaperOrchestra arXiv paper into host-agent-executable skills. It ships no API keys and no LLM SDKs; your coding agent does the reasoning and the searching.

**Ar9av/PaperOrchestra** — An automated AI research-paper writer based off Google's PaperOrchestra paper's implementation through a skills -  benchmark + autoraters using any coding agent (Claude Code, Cursor, Antigravity, Cline, Aider). No API keys, no LLM SDKs.

- Repository: https://github.com/Ar9av/PaperOrchestra
- Website: https://arxiv.org/pdf/2604.05018
- Stars: 666 · Forks: 92
- Language: Python
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ar9av-paperorchestra

## What PaperOrchestra solves, and for whom

The project targets a specific gap: you have research materials, and you want a LaTeX paper, but you do not want to build an LLM application to get one. The README describes the repository as a pluggable skill pack that lets any coding agent which can run the PaperOrchestra multi-agent pipeline turn unstructured research materials into a submission-ready LaTeX paper. The named hosts are Claude Code, Cursor, Antigravity, Cline, Aider and OpenCode.

The intended user is someone already sitting in one of those agents. If you have been running experiments through Claude Code or Cursor and never wrote a clean experiment log, the optional agent-research-aggregator skill exists for exactly that case. If you already have workspace/inputs/idea.md and workspace/inputs/experimental_log.md, the aggregator skips itself and the pipeline proceeds directly.

The paper behind the repository, cited as Song, Y., Song, Y., Pfister, T., Yoon, J., PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing, arXiv:2604.05018, 2026, defines a five-agent pipeline: Outline, Plotting, Literature Review, Section Writing and Content Refinement. The repository's claim is that this pipeline substantially outperforms single-agent and tree-search baselines on the PaperWritingBench benchmark, with a 50 to 68 percent absolute win margin on literature review quality and 14 to 38 percent on overall quality. Those are the paper's numbers, not measurements made here.

## Skills as instruction documents, not as an LLM application

The architecture is unusual and worth stating plainly. There are no API keys, no SDK dependencies and no embedded LLM calls. Each skill is a SKILL.md instruction document the host agent reads and follows, a references/ directory holding verbatim paper prompts from Appendix F plus JSON schemas, rubrics, halt rules and example outputs, and a scripts/ directory of purely deterministic local helpers. Those helpers do JSON schema validation, Levenshtein fuzzy matching, BibTeX formatting, dedup, LaTeX sanity checks and coverage gates. No network, no LLM, no API keys.

Everything that requires judgement is delegated to the host agent by instruction: LLM reasoning, web search, Semantic Scholar lookups and LaTeX compilation. That delegation is the whole design. It means the repository stays small and auditable, and it means the quality ceiling is set by whatever model and tools your coding agent has. A host with weak web search will produce weak literature review candidates, and the deterministic scripts will faithfully validate them.

The seven skills map onto the paper's steps. paper-orchestra is the top-level driver coordinating the other six. outline-agent corresponds to Step 1 and makes one LLM call, turning idea, log, template and guidelines into structured outline JSON with plotting, lit review and section plans. plotting-agent is Step 2 at roughly 20 to 30 calls. literature-review-agent is Step 3 at roughly the same count, and its documented job includes web-search candidates, Semantic Scholar verification with Levenshtein above 70, cutoff and dedup, then drafting Intro and Related Work with at least 90 percent citation integration. section-writing-agent is a single multimodal call for the remaining sections, tables and figure splicing. content-refinement-agent runs a simulated peer review across roughly five to seven calls. The README states that Steps 2 and 3 run in parallel.

## Installing the skill pack and running a first pipeline

The repository ships a setup.sh at the top level and a requirements.txt that contains only deterministic helpers: jsonschema, python-Levenshtein, reportlab, matplotlib and pypdf. The comment at the top of that file is explicit that all web search, LLM calls and Semantic Scholar fetches are delegated to the host coding agent and that the repo ships zero API integrations.

Install the Python dependencies first.

```bash
pip install -r requirements.txt
```

Then make the skills visible to your host. The README gives the Claude Code symlink as the worked example, pointing at the skills directory in the cloned repository.

```bash
ln -sf ~/paper-orchestra/skills/agent-research-aggregator \
       ~/.claude/skills/agent-research-aggregator
```

For Cursor, Antigravity, Cline and Aider the README says to follow skills/paper-orchestra/references/host-integration.md, which documents per-host invocation. Read that file before assuming the symlink pattern transfers, because the README does not reproduce the per-host steps inline.

The first real use depends on whether you have inputs. If workspace/inputs/idea.md and workspace/inputs/experimental_log.md exist, the aggregator is skipped and the pipeline runs. If they do not, point your agent at a directory and let the aggregator work. Its discovery phase is deterministic: discover_logs.py walks the search roots you pass and catalogs relevant log files across agent caches, printing a summary for review before anything is read.

```bash
python discover_logs.py --search-roots
```

Extraction then runs per roughly 50 KB batch, strips PII and flags unverified numbers as [UNVERIFIED]. Synthesis merges redundant experiment records into a single narrative and pauses to ask the user if it detects multiple disconnected projects. Formatting converts the synthesis into idea.md in Sparse Idea format and experimental_log.md in the paper's Appendix D.3 format, at which point paper-orchestra can run.

## Where the delegation model breaks down

The absence of API keys is presented as a feature, and for adoption friction it is. It is also the main failure mode. Because the skills are instructions rather than code, a host agent that ignores or half-follows a SKILL.md will produce a paper that looks structurally correct while violating the halt rules the paper specifies. The deterministic scripts catch schema violations and LaTeX sanity problems. They cannot catch a literature review whose citations were verified against the wrong records, because the verification path depends on the agent performing the Semantic Scholar lookup correctly.

The literature review step is the most exposed. Its documented acceptance bar is a Levenshtein similarity above 70 for title matching, plus cutoff and dedup, plus at least 90 percent citation integration. A fuzzy title match above 70 will accept some near-miss titles, and the README does not describe a human review gate at that point. Treat the citation list as something to inspect rather than something to trust.

The refinement step is the other weak spot. It runs a simulated peer review and accepts or reverts changes per strict halt rules, with safety constraints intended to prevent gaming the evaluator. That is a sensible design, but the README does not document rollback, so it is unclear from the repository description how to recover a draft that refinement made worse. Keep your own copy of the pre-refinement LaTeX.

Finally, this is the wrong tool if you want an unattended pipeline. There is no CLI that runs the five steps end to end. The orchestrator is a skill your agent follows, and the aggregator explicitly pauses to ask the user when it detects multiple disconnected projects.

## PaperOrchestra compared with The AI Scientist

The AI Scientist is the obvious reference point, and the two projects sit at different layers. The AI Scientist is a research system: it generates ideas, runs experiments and writes them up, which means it needs its own model access and its own execution environment. PaperOrchestra assumes the experiments already happened. Its inputs are idea.md and experimental_log.md, and its output is a LaTeX paper.

That difference shows up in the dependency list. PaperOrchestra's requirements.txt has five entries and the header comment says no LLM SDKs and no network clients. The whole reasoning budget comes from the coding agent you already pay for. If you have already done the work and the bottleneck is the write-up, this layering is the right one. If the bottleneck is having no results at all, PaperOrchestra has nothing to offer, because the aggregator can only extract experiments that exist in your agent caches.

The second difference is verifiability. PaperOrchestra ships paper-autoraters, which runs the paper's own autoraters: Citation F1 at P0 and P1, LitReview quality on six axes, SxS paper quality and SxS litreview quality. It also ships paper-writing-bench, which reverse-engineers raw materials from an existing paper to build benchmark cases. That is a self-evaluation loop the README describes in enough detail to be useful, though the autoraters are themselves LLM-driven through the host, so their verdicts inherit the same host dependency.

## Maintenance, licensing and the cost of upgrading

The repository is not archived, and the last push was on 2026-08-09. Releases are v0.1.0 on 2026-04-09 and v0.2.0 on 2026-04-25, the latter described as multi-agent skills, PaperBanana and new search backends. A CHANGELOG.md sits at the top level, which is where upgrade notes should appear.

Upgrade cost is low in the dependency sense, since requirements.txt pins nothing beyond lower bounds and pulls in no model client. It is higher in the prompt sense. The skills embed verbatim paper prompts from Appendix F, so a release that revises those prompts changes agent behaviour without changing any Python. Diffing SKILL.md and references/ between tags is the practical way to see what moved, and the CHANGELOG is the place to confirm it.

The licence is the item to check yourself. The repository reports a LICENSE file but the licence identifier is NOASSERTION, meaning GitHub could not classify it. The README does not discuss licence terms or commercial use. For a paper-generation tool whose output you may submit somewhere, read the LICENSE file directly and decide with whoever owns that decision. Nothing here is legal advice, and the repository description does not resolve the question.

## Conclusion

Adopt it if you already work inside a coding agent and have raw materials (an idea file and an experimental log) that you want turned into LaTeX without wiring up API keys. Do not adopt it if you expect a self-contained command that runs unattended, or if you need a licence you can read at a glance: the repository reports NOASSERTION and the README does not document rollback for refinement edits. Before committing, verify that your host agent is covered by host-integration.md, and check the LICENSE file directly rather than relying on the GitHub label.

## FAQ

### How do I use PaperOrchestra?

Install the deterministic helpers with pip install -r requirements.txt, symlink or otherwise expose the skills directory to your coding agent following skills/paper-orchestra/references/host-integration.md, then have the agent run the paper-orchestra skill. If workspace/inputs/idea.md and experimental_log.md are missing, the optional agent-research-aggregator skill extracts them from your agent caches first.

### Does PaperOrchestra need an API key or an LLM SDK?

No. The README states there are no API keys, no SDK dependencies and no embedded LLM calls. All LLM reasoning, web search and Semantic Scholar lookups are delegated to the host coding agent, and requirements.txt contains only deterministic helpers.

### Which coding agents does PaperOrchestra support?

The README lists Claude Code, Cursor, Antigravity, Cline, Aider and OpenCode. Per-host invocation details are in skills/paper-orchestra/references/host-integration.md, and the README only shows the Claude Code symlink inline.

### What inputs does PaperOrchestra expect before it will write a paper?

It expects workspace/inputs/idea.md and workspace/inputs/experimental_log.md. When those are absent, the optional agent-research-aggregator skill discovers logs, extracts experiments in roughly 50 KB batches, synthesises them and formats the result into those two files.

## Sources

- [Ar9av/PaperOrchestra on GitHub](https://github.com/Ar9av/PaperOrchestra)
- [Issues](https://github.com/Ar9av/PaperOrchestra/issues)
- [Project website](https://arxiv.org/pdf/2604.05018)
- [README](https://github.com/Ar9av/PaperOrchestra/blob/main/README.md)
- [Releases](https://github.com/Ar9av/PaperOrchestra/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ar9av-paperorchestra
