Model or dataset
going-doer/Paper2Code avatar
going-doer/Paper2Code

Paper2Code: Automating ML Research Code Generation with a Multi-Agent LLM Pipeline

Paper2Code: Automating Code Generation from Scientific Papers in Machine Learning

4,975 stars686 forksPythonApache-2.0

At a glance

What is it?
Paper2Code introduces PaperCoder, a multi-agent LLM system that transforms a machine learning research paper into a code repository through a structured three-stage pipeline of planning, analysis, and code generation. It is for researchers and engineers who want to reproduce ML papers without writing the implementation from scratch, and it requires either an OpenAI API key or a vLLM-compatible model server.
Who is it for?
Paper2Code is a practical starting point for ML researchers who need to reproduce a paper's code and want a structured approach rather than prompting a single LLM directly. The three-stage pipeline is more systematic than ad-hoc generation, and the evaluation tools let you measure result quality against a reference repository.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Paper2Code Solves for ML Researchers

Reproducing a machine learning paper's results requires implementing the architecture, training loop, evaluation harness, and data pipeline described in the paper text. For complex papers, this can take days or weeks even when the paper is written clearly. Paper2Code frames this as an automation problem: given the text of a paper as input, produce a working code repository as output.

The system introduced in this repository is called PaperCoder. According to the repository's README, it was accepted at the International Conference on Learning Representations (ICLR) 2026. The paper describing PaperCoder is available at arxiv.org/abs/2504.17192. PaperCoder is designed specifically for machine learning papers, not for papers in other scientific fields. The approach uses a multi-agent LLM pipeline rather than a single large prompt, on the premise that breaking the task into structured stages produces more faithful and higher-quality implementations.

The repository also ships the Paper2Code benchmark dataset, hosted on Hugging Face, which allows the evaluation of any code-generation system against a standardized set of ML papers with known reference implementations.

The Three-Stage Pipeline: Planning, Analysis, and Code Generation

PaperCoder's pipeline divides the work into three sequential stages, each handled by specialized agents.

The planning stage reads the paper and produces a high-level plan: which components are needed, how they relate to each other, and what the repository structure should look like. The analysis stage takes that plan and goes deeper, examining each component's requirements and identifying implementation decisions that are underspecified in the paper. The code generation stage implements each component using the outputs of the previous stages as context.

The output directory structure shows how the stages are recorded. After running PaperCoder on a paper, the `outputs/` directory contains three subdirectories:

bash
outputs
├── Transformer
│   ├── analyzing_artifacts
│   ├── coding_artifacts
│   └── planning_artifacts
└── Transformer_repo

The `planning_artifacts`, `analyzing_artifacts`, and `coding_artifacts` directories hold the intermediate outputs from each agent stage. The final `Transformer_repo` directory is the generated repository. Keeping the intermediate artifacts means you can inspect what each agent decided, which is useful for understanding why the generated code took a particular approach.

Running PaperCoder with the OpenAI API

The quick start path uses OpenAI's API. The README notes that the example paper is the Attention Is All You Need Transformer paper. The estimated cost for running the example with o3-mini is $0.50 to $0.70.

Install the OpenAI Python package:

bash
pip install openai

Set your API key as an environment variable:

bash
export OPENAI_API_KEY="<OPENAI_API_KEY>"

Then change into the scripts directory and run the included shell script:

bash
cd scripts
bash run.sh

For papers where you have the LaTeX source rather than the PDF, a separate script handles that path:

bash
export OPENAI_API_KEY="<OPENAI_API_KEY>"

cd scripts
bash run_latex.sh

The scripts handle the full pipeline, including calling the planning, analysis, and code generation agents in sequence. If you want to run PaperCoder on your own paper rather than the bundled example, modify the environment variables in the script to point to your paper's JSON or LaTeX input.

Using Open-Source Models via vLLM

For users who cannot or do not want to use the OpenAI API, PaperCoder supports open-source models through vLLM. The README notes that if you encounter issues installing vLLM, the official vLLM repository should be consulted for troubleshooting. The default model for this path is `deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct`.

Install vLLM:

bash
pip install vllm

Run the vLLM-based pipeline on the example paper:

bash
cd scripts
bash run_llm.sh

For the LaTeX source variant:

bash
cd scripts
bash run_latex_llm.sh

All four dependencies are also listed in `requirements.txt` if you want to install them all at once:

bash
pip install -r requirements.txt

The full list requires `openai>=1.65.4`, `vllm>=0.6.4.post1`, `transformers>=4.46.3`, and `tiktoken>=0.9.0`. The README recommends using a Python virtual environment before installing these, given the size of the dependencies.

Converting a PDF Paper to the Required JSON Format

PaperCoder accepts papers in a structured JSON format. If you have the LaTeX source, you can use it directly with the `run_latex.sh` scripts. If you only have a PDF, you need to convert it first using the `s2orc-doc2json` tool.

Clone the converter:

bash
git clone https://github.com/allenai/s2orc-doc2json.git

Start the Grobid service:

bash
cd ./s2orc-doc2json/grobid-0.7.3
./gradlew run

Then convert the PDF:

bash
mkdir -p ./s2orc-doc2json/output_dir/paper_coder
python ./s2orc-doc2json/doc2json/grobid2json/process_pdf.py \
    -i ${PDF_PATH} \
    -t ./s2orc-doc2json/temp_dir/ \
    -o ./s2orc-doc2json/output_dir/paper_coder

The README notes that in the original Paper2Code experiments, all papers were converted from PDF to JSON using this process. The examples directory includes a pre-converted `Transformer_cleaned.json` so you can run the pipeline without going through the conversion step first.

Evaluating the Quality of Generated Repositories

The repository includes evaluation tooling to score the quality of a generated repository. The model-based evaluation uses o3-mini-high to critique implementation components and assign a correctness score from 1 to 5, averaged over 8 samples.

For reference-free evaluation, which scores the generated repository against the paper alone:

bash
cd codes/
python eval.py \
    --paper_name Transformer \
    --pdf_json_path ../examples/Transformer_cleaned.json \
    --data_dir ../data \
    --output_dir ../outputs/Transformer \
    --target_repo_dir ../outputs/Transformer_repo \
    --eval_result_dir ../results \
    --eval_type ref_free \
    --generated_n 8 \
    --papercoder

For reference-based evaluation, add the `--gold_repo_dir` argument pointing to the official author-released implementation. The evaluation also requires installing `tiktoken` and setting the OpenAI API key. The README shows that the example Transformer paper scores 4.5 on an 8-sample reference-based evaluation with o3-mini.

Limitations, Cost, and When Paper2Code Is the Wrong Tool

PaperCoder does not claim to produce runnable, bug-free code in all cases. The evaluation framework measures correctness on a 1-to-5 scale, which implies that scores below 5 are expected, particularly for papers that describe architectures with incomplete implementation details. The README's experiments use the Transformer paper as an example precisely because it is one of the most thoroughly documented papers in the field.

The cost to run the full pipeline with o3-mini is estimated at $0.50 to $0.70 per paper for the example case. More complex papers with longer descriptions will cost more. The cost is tied to the OpenAI API pricing for o3-mini, which is an external dependency not controlled by this project.

The repository has no GitHub releases, and the last push was on 2026-03-25. For researchers who need the latest model compatibility or support for newer paper formats, the maintenance status is a consideration.

PaperCoder is designed for machine learning papers. It is not tested on papers from other scientific domains, and there is no documentation suggesting it would perform well on, for example, systems engineering or biology papers. The benchmark dataset covers ML papers specifically, and the agent prompts are designed with ML implementation patterns in mind.

Editorial conclusion

Paper2Code is a practical starting point for ML researchers who need to reproduce a paper's code and want a structured approach rather than prompting a single LLM directly. The three-stage pipeline is more systematic than ad-hoc generation, and the evaluation tools let you measure result quality against a reference repository. It is not appropriate for papers outside machine learning, and the quality of the output depends heavily on how clearly the paper documents its methods. The last push to the repository was on 2026-03-25. Before using it, confirm that your OpenAI API access covers the o3-mini model, or that your vLLM deployment meets the version requirements listed in requirements.txt.

Frequently asked questions

What models does Paper2Code support for generating code?

Paper2Code supports OpenAI API models (with o3-mini as the documented option) and open-source models via vLLM, with deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct as the default vLLM model.

Does Paper2Code work with papers that only exist as PDF files?

Yes, but you must first convert the PDF to JSON using the s2orc-doc2json tool with a running Grobid service. If you have the LaTeX source, you can skip this step and use the run_latex.sh or run_latex_llm.sh scripts directly.

How does PaperCoder evaluate the quality of a generated repository?

The evaluation script uses o3-mini-high to critique key implementation components and assign a correctness score from 1 to 5, averaged over 8 model samples. Both reference-free and reference-based evaluation modes are supported.

Official sources

  1. going-doer/Paper2Code on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/going-doer-paper2code.svg)](https://hysenlabs.com/projects/going-doer-paper2code)