PaperForge Research OS v3: an evidence-first paper pipeline that gates its own claims
End-to-end AI-powered academic paper writing system — from idea generation and literature search to experiment execution, result backfill, and LaTeX paper compilation. Supports multi-LLM routing, SSH remote training, incremental sync, and anti-AI-detection writing style.
At a glance
- What is it?
- PaperForge v3 chains literature search, controlled experiments and LaTeX publication into one SQLite-backed runtime, with three execution policies and a release verifier. It is opinionated about what counts as proof, and it needs a specific provider key to write at all.
- Who is it for?
- Adopt PaperForge if you already run experiments on SSH or Slurm machines and want every claim in the paper tied to a stored run, artifact hash and citation, because that linkage is the product. Do not adopt it if you only need prose: the writing path is pinned to the bailu provider and the bailu-turing model, and the README states that config.json cannot switch the production writing model, so a different gateway means editing the runtime rather than a config file.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 7 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem PaperForge targets: claims that cannot be traced back to a run
Most automated writing tools stop at text. PaperForge v3 starts from the opposite end. The README describes it as an evidence-first research system, and the mechanism that carries that idea is the SQLite Scientific Memory: claims, evidence, runs and artifacts live in one store, and a public assertion in the paper has to be linked through a claim_id to a source, a citation, a run result, an experiment result or a licence record. That is a different contract from a chat assistant that writes a paragraph about a table.
The audience follows from that design. This is for a researcher or a small group who already runs training or inference jobs, often over SSH, Docker, Slurm or Kubernetes, and who wants the paper, the code and the numbers to stay in one lineage. The repository ships launch_scientist.py, launch_mvp_workflow.py and a remote.example.yaml, which suggests the authors expect remote compute rather than a laptop-only workflow. If your work is a literature review with no experiments, the controlled experiment machinery is dead weight you will still have to step around.
One runtime behind the CLI, the browser front end and the v2 entry points
The architecture diagram in the README is the clearest statement of intent. The unified CLI, the browser front end and the v2 compatibility entries all converge on PaperForgeService. From there, work passes through an ExecutionPolicy, a Workflow Engine and a ResearchOSRuntime, which in turn talks to typed agents, a Provider Registry, compute backends and domain plugins. The Workflow Engine writes into SQLite Scientific Memory and the Artifact Registry, and the Publication Engine reads from both before the Release Verifier decides whether the workflow may move to COMPLETED.
The v3 migration table makes the change concrete. In v2, writeup, research_partner, mvp and scientist each maintained their own flow; in v3 they all enter the same service. Execution permission moved from calling convention and prompting to three enforced policies: writing-only, research and full. Experiments moved from ad hoc organisation inside a workflow to Proposal, Static Check, Mini Experiment, Full Experiment, with approval required between stages. State is per workspace under .paperforge/paperforge.db, and writes use version numbers, idempotency keys and transactions so that paperforge resume can pick up an interrupted run.
The honest reading is that v3 is a consolidation release. The individual capabilities existed in some form before; what is new is that they share a database, a policy layer and a release gate. That is good for auditability and bad for anyone who liked swapping one stage out.
Installing PaperForge and running a first writing-only job
The README requires Python 3.10, 3.11 or 3.12 on macOS, Linux or Windows, with Git, pip and venv. Clone the repository and check out the v3.0.0 tag if you want a fixed point; the README notes you can stay on main to follow later updates.
git clone https://github.com/QJHWC/PaperForge.git
cd PaperForge
git checkout v3.0.0Dependencies are split into extras, and the README is explicit that paper writing alone does not need PyTorch, Transformers, Datasets or the training stack. A writing install looks like this.
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e ".[writing]"
paperforge preflight --workspace .Before any model call you need credentials. PaperForge refuses plaintext keys as command line arguments and does not write keys into the workspace. The file lives in the user config directory, and on macOS and Linux the README requires restrictive permissions.
mkdir -p ~/.config/paperforge
chmod 600 ~/.config/paperforge/credentials.jsonThe JSON holds a single named key, and the README warns against setting conflicting base URL variables at the same time.
{
"bailu_primary": "<your key here>"
}With that in place, a first real job uses the writing-only profile, which the README says refuses training, inference and experiment commands, SSH and container backends, weight files such as .pt and .ckpt, experimental data formats, and run_* result directories at four separate layers.
mkdir -p workspace/my-paper
paperforge run \
--profile writing-only \
--workspace workspace/my-paper \
--title "Paper title" \
--topic "Research topic" \
--instructions "Write the paper only; run no training, inference or experiments"After it starts, paperforge status --workspace workspace/my-paper reports the latest state, and paperforge resume --workspace workspace/my-paper continues an interrupted run. Expect the preflight to report CODE_VERIFIED when Python and the workspace are usable; that status says nothing about whether your key works. Adding --live-provider to the preflight is what makes one minimal real request and can return EXTERNAL_SERVICE_VERIFIED or AUTH_BLOCKED.
Where PaperForge pushes back: provider lock-in, LaTeX gaps and the publish gate
The sharpest limitation is stated in the README itself. The current v3 production writing path is fixed to the bailu provider and the bailu-turing model, and the README tells you not to rely on config.json to switch the production writing model, because the CLI runtime does not inject provider, model or generation_profile from that file into the writing agent. The Bailu endpoint is OpenAI-compatible and defaults to https://bailucode.com/openapi/v1, and the request builder strips fields Bailu does not accept: reasoning_effort, seed, n and stop. You can override the gateway with OPENAI_BASE_URL, but the model choice is not a configuration knob in this release. Anyone who assumed multi-LLM routing meant free model selection should read that paragraph twice.
The second gap is tooling. The README states that without LaTeX or Poppler, research and state management still work but paperforge publish cannot complete a formal PDF release. The static preflight reports latex_compiler, bibtex and pdftoppm separately and does not downgrade the top-level status to FAILED when they are missing. A missing pdftoppm therefore looks like a healthy environment until the release verifier refuses the PDF, and the README notes that the release gate requires claims, PDF, protected blocks, page checks, secret scanning and a release manifest to pass together before the workflow reaches COMPLETED.
Third, the licence field in the repository metadata is NOASSERTION while pyproject.toml declares MIT and lists LICENSE. The README carries a disclaimer that the project is for study and research only and not for commercial use. Those two statements do not obviously agree, and the repository does not resolve the conflict. Treat the licence question as open until you read LICENSE and THIRD_PARTY_NOTICES.md yourself.
How PaperForge differs from a LaTeX-first assistant such as Aider
The natural comparison is a coding assistant pointed at a paper repository, and Aider is the concrete one here because aider-chat==0.86.2 is a dependency of the writing extra and the README mentions an Aider request path. Aider edits files in a repository on request. It has no notion of a claim, an experiment run or a release manifest. If you ask it to change a number in a table, it changes the number.
PaperForge inverts the control. The Publication Engine reads from Scientific Memory and the Artifact Registry, and the Release Verifier checks the claim set before the workflow can complete. The README also describes up to three rounds of compilation, rendering, diagnosis and bounded layout repair across Generic, CVPR, IEEE and Elsevier templates, plus protected blocks that the repair pass is not allowed to rewrite. That is a different failure mode: Aider fails by producing plausible text, PaperForge fails by refusing to publish.
If your paper has no experiments, the comparison stops mattering and a plain LaTeX editor plus a coding assistant is the lighter choice. PaperForge earns its complexity only when there is a run to point back to.
Maintenance status, upgrade cost and what the release actually covers
The repository is not archived, and the last push was on 2026-07-26. There are no retrieved releases, so the v3.0.0 version string in the README and pyproject.toml is the reference point rather than a published artifact. The presence of CHANGELOG.md, a CI workflow badge and a dev extra with pytest, ruff and mypy indicates the project is maintained, but nothing in the repository describes a release cadence or a support window.
Upgrade cost is concentrated in the v2 to v3 move. The migration table shows entry points collapsing into PaperForgeService, execution permissions becoming enforced policies, and evidence moving into SQLite. A v2 workspace whose results were stored as loose documents and result files has to be re-expressed as claims, evidence and runs. The README documents a resume path for interrupted workflows but does not document rollback from v3 to v2, and the migration section is not reproduced in the repository files examined here. Plan for that gap.
On licensing, pyproject.toml declares MIT and ships LICENSE and THIRD_PARTY_NOTICES.md, while the metadata says NOASSERTION and the README carries a non-commercial disclaimer. The third_party directory and the notices file suggest bundled components with their own terms. This is not legal advice; read both files before you ship anything built on it.
Editorial conclusion
Adopt PaperForge if you already run experiments on SSH or Slurm machines and want every claim in the paper tied to a stored run, artifact hash and citation, because that linkage is the product. Do not adopt it if you only need prose: the writing path is pinned to the bailu provider and the bailu-turing model, and the README states that config.json cannot switch the production writing model, so a different gateway means editing the runtime rather than a config file. Before committing, run paperforge preflight --workspace . with --live-provider to confirm the key works, and check that latexmk, bibtex and pdftoppm are on PATH, since the static preflight still reports CODE_VERIFIED when they are missing and the failure only surfaces at paperforge publish.
Frequently asked questions
What is PaperForge Research OS v3?
It is a Python research system that covers planning, controlled experiments, paper writing, review, typesetting and release verification, with all entry points converging on one workflow database and runtime. The README describes it as evidence-first, meaning public claims in the paper are linked to sources, runs or artifacts through a claim_id.
How do I install PaperForge?
Clone the repository, create a Python 3.10 to 3.12 virtual environment, and install the writing extra with python -m pip install -e ".[writing]". Then run paperforge preflight --workspace . to check the environment. LaTeX and Poppler are only needed for paperforge publish.
Which model does PaperForge use for writing?
The README states that the current v3 production writing path is fixed to the bailu provider and the bailu-turing model, and that config.json cannot switch it because the CLI runtime does not inject provider, model or generation_profile from that file. The endpoint is OpenAI-compatible and defaults to https://bailucode.com/openapi/v1.
Why does paperforge publish fail when preflight says the environment is fine?
The static preflight only reflects whether Python and the workspace are usable, and it reports latex_compiler, bibtex and pdftoppm as separate fields without failing the top-level status. If those tools are missing, the release gate cannot pass and paperforge publish cannot complete a formal PDF release.
Can PaperForge run training or experiments over SSH?
Yes, the experiment extra includes paramiko and the repository ships remote.example.yaml, and the README lists SSH among the unified compute backends alongside Local, Docker, Slurm, Kubernetes and Cloud SSH. The writing-only profile refuses SSH, containers and remote compute, so you need the research or full policy instead.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/qjhwc-paperforge)