Model or dataset
Spark-To-Paper-Skills/paperjury avatar
Spark-To-Paper-Skills/paperjury

PaperJury is a prompt package with a deterministic clerk and two npm scripts

Pre-submission AI review stress-test for research papers. A Claude Code skill: review, verdict, revise, verify.

1,214 stars44 forksJavaScriptMIT

At a glance

What is it?
PaperJury is a pre-submission review skill that runs a loop of review, verdict, revision and re-check on a LaTeX paper. The model reads, judges and drafts; state, voting, patch application and the stopping condition belong to code. It publishes a careful evaluation table whose own footnotes contain the most important caveats on the page.
Who is it for?
The design here is the interesting part, and two decisions are worth copying whatever else you think of it. Breaking the memory chain between rounds so a jury cannot anchor on its own earlier conclusions, while a clerk outside the model's view keeps the whole record, is a real answer to a real failure mode.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 50 days ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A prompt package with two npm scripts

The clearest thing about this repository is what it is not. The package manifest has no dependencies at all, an engine floor of Node 18, and exactly two scripts: a doctor script and a test command that invokes Node's own built-in test runner across a glob of test files.

That is the whole JavaScript surface. There is no framework, no model client, no build step. The behaviour lives in a skill definition file at the repository root, a workflows directory, a references directory and a plugin manifest directory. It is a prompt package with a small deterministic harness around it, not an application.

Which also explains the two-command install. It is added as a plugin from a plugin marketplace, not installed from a package registry:

text
/plugin marketplace add Spark-To-Paper-Skills/paperjury
/plugin install paperjury@Spark-To-Paper-Skills

The version in the manifest matches the newest release, and the test script is the whole quality gate:

bash
node --test tests/*.test.js

The manifest is marked as publishable, which is a leftover of it once being a package rather than a plugin.

The interesting structural decision is that there are two of these. A separate repository holds the version for a different coding assistant, and the first stable release of this one is described on the page as aligned with that other version's first stable release. Maintaining the same skill for two hosts in two repositories with version parity as an explicit goal is a real ongoing cost, and it is the kind of thing a single-skill project usually avoids by dropping the second host.

The model judges and the code keeps score

The subtitle is the architecture: the model is responsible for reading, judging and drafting, while state management, voting rules, patch application and the stopping condition are controlled by deterministic code.

Four responsibilities, and all four are the ones a language model is bad at. Keeping state across rounds is bookkeeping. Counting votes is arithmetic. Applying a patch is a file operation that must happen exactly once. Deciding when to stop is a termination condition that has to be evaluated identically every time.

The stopping rule is worth spelling out because it is the kind of thing that ends up either missing or hand-waved. The loop continues until a round finds nothing new. Each round is an independent re-check of the current manuscript; a deterministic clerk writes every round's result into one ledger; and when a round surfaces no new problems, that round is the signal to stop.

There is a second control in the same family. Before edits are applied, a fresh challenger re-examines findings that were already rejected, to reduce misjudgement. So a rejection is not final: the loop has an appeal path, and the appeal is heard by someone who has not seen the original objection.

The jury is blind to its own earlier verdicts

Multi-round review has a well known failure mode. A reviewer that remembers what it concluded last round stops looking, and the rounds converge on the first opinion. The usual fix is more rounds, which makes it worse.

This tool takes the opposite approach: each round reviews only the current manuscript, and the jury is not shown the previous round's ledger, specifically so earlier conclusions cannot anchor it. The ledger still exists, and the deterministic clerk keeps appending to it. The record is complete and the model's view of it is not.

That is a neat separation. The information a reviewer needs to notice something new is the current text; the information an auditor needs is the whole history; and the two wants are incompatible, so the history lives somewhere the reviewer cannot read.

The escalation ladder is on the same theme. Mechanical or minor findings go one way, substantive findings go to a side hearing between two parties, and a five-member jury deliberates in isolation. If there is no clear majority the panel is enlarged, to twelve. If every juror says it cannot judge from the context it was given, the item goes back to the author.

That last branch is the one to notice. A system that returns the question is a different thing from a system that always produces a verdict, and it is the behaviour most review pipelines cannot express because their output type has nowhere to put abstention.

The evaluation table's own footnotes are the story

The comparison is twelve held-out papers, four papers each from three areas, four baselines, and an audit by blind expert reviewers. Three numbers are lifted to the top of the page as headline tiles with the strongest baseline struck through underneath them.

Read the footnotes and the tiles stop being headlines.

The first footnote is the important one. The unsafe-edit rate is the proportion of applied edits that violated a safety rule, and the page says it must be read together with how many edits were applied: thirteen point four per paper for this tool against fourteen point three for the baseline it is compared to. So a fourfold reduction in an unsafe rate is meaningful here precisely because the two methods apply a comparable number of edits. That caveat is the difference between a rate that means something and a rate that was won by editing less.

The second footnote deflates a second number. The round counts are three point zero eight on average with a standard deviation of zero point six seven, and no paper hit the ceiling, while the strongest baseline averaged three point three three with a standard deviation above one and hit the ceiling on two of the twelve papers. The baseline was truncated; this one was not.

The third is a presentational wrinkle rather than an error. The tiles show one precision figure for the audit and the table shows another under a different column heading, and the footnote defines the table's version as blind-expert agreement on the final verdict. So the two numbers may be different measures, or one may be stale, and the page does not say which.

The method that wins is the second most expensive row

The cost column is the one the page does not lift to a tile, and it is the most decision-relevant column there is.

A forward-only rewriter takes just over a quarter of an hour per paper. An LLM critic takes about half an hour. The LLM-as-judge loop takes just over two hours. This tool takes two and a half. A naive unbounded generator takes eight and a third.

So the two cheap methods are cheap because they do not do the task: both produce no findings list at all, and the table marks their quality score as not applicable. Among the methods that actually produce a review, this tool costs about twenty percent more time per paper than the baseline it improves on, and its saving is against the naive generator, at roughly a quarter of the hours and a fifth of the tokens.

That is a defensible trade and it is worth stating plainly rather than discovering later. If you want the verdict quality, this is one of the more expensive ways to get it, and the reason is structural: a jury that deliberates in isolation, escalates to twelve, and re-checks its own rejections is doing many more model calls than a critic that answers once.

The per-paper figures are also modest in absolute terms, which makes the comparison easy to run yourself. Two and a half hours and under seven million tokens is a figure an author can budget for one manuscript without a procurement conversation.

Auto mode is opt-in and the compile step degrades honestly

There are three modes and the tool picks between two of them from your description. Ask it to tighten one paragraph and it drafts a patch and waits. Ask it to review and it runs the loop. The unattended mode, where safe edits are applied without asking, must be switched on explicitly, which is the correct default for a tool that edits your manuscript.

The guardrails are graded by risk rather than applied uniformly. A safe fix is accepted against a frozen anchor, a per-passage edit cap, and an audit of both the anchor and whether the meaning of other sections still holds. In the two interactive modes every change waits for you. In the unattended mode the safe ones apply under the policy you authorised in advance and the high-risk ones go into a to-do queue and come back to you.

The verification step is the part most tools skip and this one does not. It compiles the paper on your machine and reports errors, undefined references, overflowing boxes and page count. If you have no toolchain it says which checks it could not run and degrades to structural lint, rather than reporting success it did not achieve.

Alongside that sits a set of deterministic compliance checks, and the first one is the one that matters most in practice: anonymisation leaks. A conference submission whose source reveals who wrote it is rejected before review. Also checked are margin changes, drift in the document class declaration, missing required sections and page limit overruns. Those are desk-reject risks, and having them checked by code rather than by a model is exactly the right division.

The terms are in a callout and the news leads with a share count

The most important paragraph on the page is a callout near the top, and it is worth reading before the feature tables.

It says this is a self-check tool for before submission, that it cannot replace the author's scientific judgement and cannot replace peer review, that it must not be used to invent experiments, fabricate results, add claims without evidence, or conceal limitations, and that anything requiring new experiments, missing evidence, the author's private knowledge, or research judgement goes back to the author.

That is an unusually explicit statement for a tool in this category, and it lines up with the design rather than sitting beside it. The verdict that routes to the author is the mechanism, and the callout is the promise that the mechanism will be used.

The other thing the page asks you to do first is look at a worked example. There is a sample directory containing a real draft taken through one full review in the unattended mode, with the manuscript before and after as PDFs and a run report that a person checked by hand. The page recommends reading it before deciding whether to use the tool on your own paper, which is a better onboarding order than most projects manage.

The news log, though, opens with an audience milestone on a social platform and lists it above the paper publication and above the first stable release. The newest release is from mid-June and the last commit to the branch is from mid-August, so the log is two months behind the code while leading with the number that reads best.

Editorial conclusion

The design here is the interesting part, and two decisions are worth copying whatever else you think of it. Breaking the memory chain between rounds so a jury cannot anchor on its own earlier conclusions, while a clerk outside the model's view keeps the whole record, is a real answer to a real failure mode. And having every juror decline to answer when it lacks context, instead of forcing a verdict, is the behaviour most review systems cannot express. Three things to check before you rely on it. Read the cost column rather than the accuracy column, because the method that wins the comparison is also the second most expensive row in the table. Read the safety rate next to the edit count, which the page itself insists on. And read the terms, because the tool states plainly that it cannot make the judgement calls that matter and hands those back, which means a paper that comes back clean has been checked, not approved.

Frequently asked questions

What is PaperJury?

It is a pre-submission review tool for research papers, delivered as a coding assistant skill, that organises a self-check into a loop of review, verdict, revision and re-check. The model reads, judges and drafts, while state management, voting rules, patch application and the stopping condition are handled by deterministic code.

What verdicts does PaperJury return?

Three. Uphold and fixable for text-level problems such as unclear writing or an over-strong claim, which get a minimal patch through guardrail checks. Return to the author when experiments, data or evidence are missing. And reject when the reviewer misread the paper, in which case the finding is recorded and later re-examined by a fresh challenger before edits are applied.

How does PaperJury avoid repeating itself across rounds?

Each round re-checks only the current manuscript, and the jury is not shown the previous round's ledger so earlier conclusions cannot anchor it. A deterministic clerk writes every round's results into one ledger, and the loop ends when a round finds nothing new.

How much does PaperJury cost per paper?

The published table gives 2.47 hours and 6.76 million tokens per paper, against 2.06 hours for the strongest baseline it improves on and 8.37 hours and 31.4 million tokens for a naive unbounded generator. The two cheapest methods take under an hour each but produce no findings list at all.

What does PaperJury check before submission?

It runs a real LaTeX compilation on your machine when a toolchain is available and reports errors, undefined references, overflowing boxes and page count, stating explicitly which checks it could not complete otherwise and degrading to structural lint. Compliance checks cover anonymisation leaks, margin changes, documentclass drift, missing required sections and page limit overruns.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. Spark-To-Paper-Skills/paperjury on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/spark-to-paper-skills-paperjury.svg)](https://hysenlabs.com/projects/spark-to-paper-skills-paperjury)