Model or dataset
Spark-To-Paper-Skills/spark-to-paper-skills avatar
Spark-To-Paper-Skills/spark-to-paper-skills

spark-to-paper-skills installs with a git clone and loops its reviewers until they find nothing

One sentence in, one draft paper out: spark-to-paper-skills automatically reviews papers, plans and runs experiments, and writes a research draft—typically using about $10 in API costs per paper.

1,217 stars20 forksPythonMIT

At a glance

What is it?
Thirteen agent skills and one orchestrator that take a pasted sentence to a compiled paper, with deterministic gates on citations, source-traced numbers, vector structure and LaTeX, and a proposal mode that leaves result cells blank on purpose. No application, no server, no tests, and one secrets file that can borrow another tool's credentials.
Who is it for?
The parts worth keeping are the gates and the two modes. Requiring deterministic checks on citation records, claim to citation links, numbers in prose, figure structure and the LaTeX before anything is called done is the right shape for a system that writes its own references, and separating a proposal that is allowed to leave result cells blank from a paper about data you actually hold is an unusually honest distinction.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 42 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The install is a repository cloned into the directory your agent reads

There is no package to install and no binary to run. One command clones the repository into a skills directory inside your home configuration, and a note in the changelog says it restructured itself as a proper plugin with a plugin manifest, so it loads on the next agent session. Everything after that is a sentence typed into a session, with the readme telling you to paste your idea, a proposal or data and the orchestrator routing it, choosing a mode and running the chain. Thirteen skills are composable underneath, and the install is a single line:

bash
# Install — auto-loads on next Claude Code session
git clone https://github.com/Spark-To-Paper-Skills/spark-to-paper-skills.git ~/.claude/skills/spark-to-paper-skills

The command clones a branch of a live repository straight into a configuration directory, with no checksum, no tag and no pin. The consequence is worth stating plainly: what you have installed is a directory of instructions and scripts that an agent reads on every session, and the update path is a git pull rather than a versioned release.

The adversarial review loops until it finds nothing, with no stated cap

The review stage is described as isolated reviewers who read the whole paper with a verbatim-quote rule designed to stop them skimming, followed by perspective-diverse skeptics who try to refute each issue the reviewers raised, and then a loop that runs until it is dry. That last clause is the one to notice: dry means no new issues, so termination depends on a model eventually running out of objections. The documentation states no iteration cap, no token budget and no time limit for that stage. The project describes its typical cost as about ten dollars of API spend per paper, and this loop is where that number is decided. The other half of the integrity story is cheaper and more concrete: deterministic gates check citation records, claim to citation links, numbers traced to sources in the prose, vector figure structure, and the LaTeX, and a violation fails the build.

Two integrity modes, and the difference is grammatical tense

The most interesting design decision on the page is a distinction most paper generators do not make. In proposal mode the writing is forward-looking and the result cells in tables are deliberately left blank, which is what a proposal is allowed to contain. In data-aware mode every number is traced to your real data and the prose is written in the past tense, and the whole thing is machine-audited. The same gates run over both, so a blank cell in proposal mode is a legitimate state rather than a missing one. The audit trail is part of the output rather than an afterthought: every stage writes an input, decisions and output log, and a single gate entry point has to pass across four families, citations, the draft, the figures and the LaTeX, before the build is called good.

Citation checking is hedged, and the figure engine redraws rather than traces

The reference file is described as real BibTeX entries whose citation records are checked via web search and Crossref when available. That hedge is the honest weak point in an otherwise strong claim, because a gate that checks citation records is only as good as the check behind it. The figure pipeline is more interesting. An external figure tool renders candidate images, and then a skill learns the design language of that render and redraws the figure natively from the paper's own facts, iterating against a geometry audit written using the standard library alone until it passes. Output is vector, in three formats, with the original render always kept alongside. The stated benefit over the usual approach is that the result contains live text elements rather than a traced bitmap.

One secrets file, and an empty key silently borrows another tool's credentials

Configuration is a single file, copied into the project, ignored by git, auto-loaded by the scripts, and overridden by any real exported environment variable. Three things in it deserve a second look. The vision key is documented as optional, and when it is unset the scripts fall back to a credentials file belonging to a different tool on your machine; a legacy code path is documented as authenticating the same way. The embedding model field ships empty, so you must choose one. And the image model must be picked explicitly from a five-entry list, alongside a switch that chooses between two different HTTP endpoints, one for image generation and one for chat completions, because the correct setting depends on which model you chose. The default output size is a fixed landscape rectangle.

The figure runtime downloads a gated model, and Stage eight can reach Overleaf

Provisioning the vector redraw runtime fetches three models from Hugging Face: a segmentation model that is gated, meaning you need a token and you have to accept its terms on the model's page yourself, plus an optical character recognition model and a background removal model. The comment is clear that this is a one-time step and is not used at run time once the models are deployed, which is the right shape for a multi-gigabyte dependency. The other outbound path is Overleaf, documented as being only for the experiment stage and only if you enable it, off by default behind a configuration flag in a config file that does not appear among the repository's top-level entries. Stage eight also writes and runs experiment code against the data you hand it, which is the stage where this project touches your machine most directly. Nothing in the page describes a sandbox, a container or a permission boundary around that execution, so the isolation is whatever your own agent session already has. The outputs list one further artefact worth naming: an experiments directory holding the run code and the filled result tables, so the numbers in the paper are backed by something re-runnable rather than by prose.

Three version strings, two template lists, and a comparison table that stops mid-sentence

Small inconsistencies that a careful reader will hit. The release tags are 1.0.1, 1.1.0 and 1.2, while the changelog entry for the newest calls it 1.2.0 and a plain version file sits at the repository root, so the tag and the documented version for one release differ. The sample set leads with the official ICML 2025 style among its entries, while the feature table says the bundled templates are NeurIPS and one engineering conference. The comparison table scores four named alternatives on a three-level scale and is cut off partway through its source list. And version 1.0.1 added a soft update check that queries the GitHub releases API on each run, with a day-long cache, silent when current, and never blocking, which is a network request per invocation for a project whose whole selling point is that there is no server.

Editorial conclusion

The parts worth keeping are the gates and the two modes. Requiring deterministic checks on citation records, claim to citation links, numbers in prose, figure structure and the LaTeX before anything is called done is the right shape for a system that writes its own references, and separating a proposal that is allowed to leave result cells blank from a paper about data you actually hold is an unusually honest distinction. The risks are concrete. The install is a repository cloned straight into the directory your agent reads, so you are auditing a whole tree before you trust it. The adversarial review loop has no iteration cap in the documentation, and the project states about ten dollars of API spend per paper, so that loop is your cost. Stage eight writes and runs code. And the secrets file will silently borrow another tool's credentials when you leave a key empty. Read it, cap it, and run it on data you can lose.

Frequently asked questions

What is spark-to-paper-skills?

One orchestrator and 13 composable Claude Code skills that turn a one-line idea into a compiled paper PDF, with no separate application or orchestration server. You install it by cloning the repository into your skills directory and it loads on the next agent session.

What does spark-to-paper-skills cost to run?

The project description states it typically uses about ten dollars in API costs per paper. Several model settings are yours to choose, including an embedding model that ships empty in the example configuration and an image model you must select explicitly from a list of five.

What is the difference between proposal mode and data-aware mode?

Proposal mode is forward-looking and leaves result cells in tables blank. Data-aware mode traces every number to your real data and writes in the past tense. Both are machine-audited, and deterministic gates fail the build when they find violations.

How does spark-to-paper-skills make figures?

An external figure tool renders candidate images, then a skill learns the design language of that render and redraws the figure natively from the paper's own facts, iterating against a geometry audit written with the standard library alone until it passes. Output is vector in SVG, PDF and PPTX, with the original render kept.

Does spark-to-paper-skills check its citations?

It emits real BibTeX entries and describes citation records as checked via web search and Crossref when available. A deterministic gate then checks citation records and the links between claims and citations, failing the build on violations.

What does run_gates.py do?

It is the single entry point that must pass before the paper builds, covering four gate families: citations, the draft, vector figure structure and the LaTeX output.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. Releases
  5. Spark-To-Paper-Skills/spark-to-paper-skills on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/spark-to-paper-skills-spark-to-paper-skills.svg)](https://hysenlabs.com/projects/spark-to-paper-skills-spark-to-paper-skills)