# paper2code: A Claude Code Skill for Citation-Anchored ML Paper Implementations

> paper2code is a Claude Code skill that takes an arXiv URL and produces a Python implementation where every code decision is anchored to a specific paper section. It audits every implementation choice before writing a line, marking unspecified hyperparameters explicitly rather than filling them in silently.

**PrathamLearnsToCode/paper2code** — Agent skill to turn any arxiv paper into a working implementation

- Repository: https://github.com/PrathamLearnsToCode/paper2code
- Stars: 1,522 · Forks: 179
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/prathamlearnstocode-paper2code

## The Problem: Papers That Lie by Omission

Machine learning papers routinely omit critical details. Hyperparameters appear in appendices, prose contradicts equations, and phrases like "standard settings" refer to nothing documented. Anyone who has tried to reproduce a paper from scratch knows that the actual implementation work is less about coding and more about cross-referencing footnotes, checking the authors' GitHub if it exists, and making judgment calls about things the paper never addresses.

General-purpose code generation tools make this worse in a specific way. When you ask a coding assistant to implement a paper directly, it fills every gap silently and confidently. The resulting code runs, but you cannot tell which parts came from the paper and which were invented by the model. You end up with a plausible-looking implementation that may diverge from the paper in a dozen places, with no map back to the original text.

paper2code addresses this by refusing to generate code until it has classified every implementation choice. Choices get one of three labels: SPECIFIED (the paper states it clearly), PARTIALLY_SPECIFIED (the paper mentions it but is ambiguous, with a quoted passage), or UNSPECIFIED (the paper is silent, so a common default is used and alternatives are listed). The README describes this as "appendix mining": appendices, footnotes, and figure captions are treated as first-class sources rather than supplementary material.

## Citation Anchoring: How Code Stays Traceable

The core mechanism is a set of inline tags that link generated code to the paper it came from. Every non-trivial code decision carries a tag in a comment at the exact line where the choice is made. The README defines six tags:

- `§X.Y` marks a decision directly specified in paper section X.Y.
- `§X.Y, Eq. N` marks code that implements equation N from section X.Y.
- `[UNSPECIFIED]` marks a choice the paper does not address, always accompanied by the default used and at least one alternative.
- `[PARTIALLY_SPECIFIED]` marks an ambiguous mention in the paper, with the relevant quote included.
- `[ASSUMPTION]` marks a reasonable inference from context, with the reasoning explained.
- `[FROM_OFFICIAL_CODE]` marks a detail taken from the authors' own released implementation.

The README provides a concrete example showing a TransformerBlock where the residual connection is tagged `§3.2` and the attention computation is tagged `§3.2, Eq. 2`. A LayerNorm epsilon that the paper never specifies gets an `[UNSPECIFIED]` comment listing `1e-6` as the choice, `1e-5` as the PyTorch default, and `1e-8` as another common alternative.

```python
# §3.2, Eq. 2 — attention_weights = softmax(QK^T / sqrt(d_k))
attn_out = self.attention(self.norm1(x))  # (batch, seq_len, d_model)
x = x + attn_out  # §3.2 — residual connection
```

This lets a reader verify any line against the paper without tracing through the whole file. It also makes disagreements explicit: if you believe the paper actually specifies a different epsilon, the comment is the place to start the argument.

## Installing the Skill and Running Your First Paper

paper2code is distributed as a Claude Code skill, not a standalone Python package. Installation uses the `skills` command:

```bash
npx skills add PrathamLearnsToCode/paper2code/skills/paper2code
```

The installer prompts you to select which coding agents should have the skill, choose between global and project-level scope, and pick between a symlink installation and a copy. The README recommends both global scope and symlink. Once installed, open Claude Code:

```bash
claude
```

Then run the skill with an arXiv URL:

```bash
/paper2code https://arxiv.org/abs/1706.03762
```

The skill also accepts bare arXiv IDs (`/paper2code 1706.03762`) and supports several flags. `--framework jax` specifies the target framework when the paper does not prescribe one. `--mode full` adds a training loop and data pipeline to the output. `--mode educational` adds extra comments and a pedagogical notebook aimed at readers studying the paper rather than reproducing it.

The README does not document a default framework, so if the paper is framework-agnostic, the skill will make a choice and likely mark it `[ASSUMPTION]`.

## What the Skill Produces

Each run creates a directory named after the paper, containing a consistent set of files. The README describes the layout for the "Attention Is All You Need" paper as an example:

- `README.md` summarizes the paper, states the contribution, and gives a quick start.
- `REPRODUCTION_NOTES.md` is the full ambiguity audit: every implementation choice, whether the paper specified it, and what alternatives exist for unspecified ones.
- `requirements.txt` lists pinned dependencies.
- `src/model.py` contains only the architecture, with each class mapped to a paper section and variable names matching paper notation.
- `src/loss.py` holds loss functions with equation references.
- `src/data.py` provides a Dataset class skeleton with preprocessing instructions but no actual data download logic.
- `src/train.py` appears when `--mode full` is used.
- `base.yaml` is the single source of truth for all hyperparameters.
- `walkthrough.ipynb` runs on CPU with reduced dimensions, quotes paper passages next to the corresponding code, and runs shape checks.

The key design choice is that `model.py` is architecture only. It does not include training infrastructure, logging, or experiment management. The README states this explicitly: no distributed training, no experiment tracking, no checkpointing beyond what the paper's contribution requires.

## Six Things paper2code Will Not Do

The README lists the skill's limits in direct terms. Understanding them before running the skill avoids disappointment.

First, the skill does not guarantee correctness. If the paper describes something incorrectly, the code will faithfully implement the incorrect description. The citation anchoring helps you catch this by making every decision traceable, but the skill does not cross-check the paper against external sources.

Second, it does not invent details. An unspecified hyperparameter gets a common default and an `[UNSPECIFIED]` flag, not a guess that looks authoritative. This is a deliberate contrast with general-purpose LLM code generation, which fills gaps without marking them.

Third, `data.py` provides a Dataset class skeleton with instructions on where to get the data. It does not download datasets or set up preprocessing pipelines beyond what the paper describes.

Fourth, the output contains no distributed training, no experiment tracking tool integration, and no checkpointing system unless the paper specifically describes one of these as a contribution.

Fifth, the skill implements only the core contribution of the paper. Baselines are not included.

Sixth, standard components referenced by name in the paper are imported or noted as dependencies. The skill does not reimplement a standard transformer encoder from scratch if the paper says to use one.

The alternative to paper2code is asking a general-purpose coding assistant to implement a paper directly. That approach generates code faster, but the result is harder to verify: the assistant fills all gaps confidently, variable names may diverge from paper notation, and there is no REPRODUCTION_NOTES.md to anchor the review process.

## Repository Layout and Maintenance

The repository contains three top-level entries: LICENSE, README.md, and a `skills/` directory. The skill itself lives under `skills/paper2code/`. The README describes a contribution workflow for worked examples: run the skill on a well-known paper, save the full output to `skills/paper2code/worked/{paper_slug}/`, and write a `review.md` evaluating what the skill got right, what it correctly flagged as unspecified, and any mistakes.

The project is licensed under the MIT licence, which permits use in commercial and research contexts without restriction beyond attribution.

The last push to the repository was on 2026-04-03. The repository has no published GitHub releases. The README includes a placeholder note for an animated GIF showing the pipeline, indicating the project is still developing its documentation alongside its code.

## Conclusion

paper2code is worth adopting if you need verifiable ML paper implementations and can accept that the output is a starting point rather than a finished, runnable experiment. Skip it if you want code that executes end-to-end without first resolving flagged unknowns. Before running it, check that your target paper contains numbered sections and explicit equations: papers with heavy informal descriptions and no equation numbering give the citation anchoring system less to work with, and the REPRODUCTION_NOTES.md will contain more UNSPECIFIED entries than resolved ones.

## FAQ

### How can I convert my research paper into working code with paper2code?

Install the skill with `npx skills add PrathamLearnsToCode/paper2code/skills/paper2code`, open Claude Code, and run `/paper2code https://arxiv.org/abs/YOUR_PAPER_ID`. The skill audits the paper's implementation choices before generating code, producing a directory with citation-anchored source files and a REPRODUCTION_NOTES.md that lists every unspecified hyperparameter.

### Which Python frameworks does paper2code support for generating implementations?

The README shows PyTorch examples in its citation anchoring documentation and lists PyTorch in the example requirements, but the `--framework jax` flag demonstrates that other frameworks are supported. The README does not enumerate every supported framework explicitly.

### What is the REPRODUCTION_NOTES.md file that paper2code generates?

REPRODUCTION_NOTES.md is the ambiguity audit produced alongside the implementation. It lists every implementation choice the skill encountered, classifies each as SPECIFIED, PARTIALLY_SPECIFIED, or UNSPECIFIED, and for unspecified choices documents which default was used and what alternatives exist.

## Sources

- [Issues](https://github.com/PrathamLearnsToCode/paper2code/issues)
- [License: MIT](https://github.com/PrathamLearnsToCode/paper2code/blob/main/LICENSE)
- [PrathamLearnsToCode/paper2code on GitHub](https://github.com/PrathamLearnsToCode/paper2code)
- [README](https://github.com/PrathamLearnsToCode/paper2code/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/prathamlearnstocode-paper2code
