# arxiv-latex-cleaner: preparing a LaTeX paper for arXiv submission

> A Google Research command line tool that copies a LaTeX project into a cleaned folder, stripping comments, unused files and oversized images before you upload the ZIP to arXiv. It is narrow by design, and the narrowness is the point.

**google-research/arxiv-latex-cleaner** — arXiv LaTeX Cleaner: Easily clean the LaTeX code of your paper to submit to arXiv

- Repository: https://github.com/google-research/arxiv-latex-cleaner
- Stars: 7,061 · Forks: 421
- Language: Python
- License: Apache-2.0
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/google-research-arxiv-latex-cleaner

## The submission checklist problem arxiv-latex-cleaner targets

arXiv accepts a source upload, not just a PDF, and the source you keep on your laptop is rarely the source you want to publish. It carries commented-out paragraphs, notes to coauthors, stale figure files, and a bibliography you may or may not want to ship. There is also a 50MB limit on submissions, which the README names directly as the reason for the size-oriented features. So the last hour before a deadline is usually spent deleting comments by hand, chasing which .png files are actually referenced, and guessing at image resolution until the ZIP fits.

This tool automates that pass. You point it at a folder of LaTeX code, for example /path/to/latex/, and it writes a new folder /path/to/latex_arXiv/ that the README describes as ready to ZIP and upload. The audience is narrow and clear: authors of LaTeX papers who are submitting to arXiv and who want the source cleaned without editing the original project. It is not a typesetting system, not a PDF builder, and not useful to anyone working in Word or Markdown.

## What the cleaner actually rewrites in your source tree

The mechanism is a source-to-source copy with rules applied along the way. Nothing in the README suggests the tool compiles your document, so the cleaning is textual and file-level.

On the privacy side, it removes auxiliary files such as .aux, .log and .out, and it strips comments. The comment removal is broader than a line-by-line pass: the README states it also handles \begin{comment}\end{comment}, \iffalse\fi and \if0\fi blocks, which are the usual ways authors hide text. It can additionally delete user-defined commands you name with commands_to_delete, the README's example being a \todo{} macro you redefine as empty.

On the size side, the tool deletes .tex files that are neither in the root nor included by another .tex file, and images that no used .tex file references. It can resize images to a longest-side pixel count, compress PDFs with ghostscript on Linux and macOS, and convert PNG to JPG. The --images_allowlist flag takes a dictionary mapping a path to a per-file size, pixels for images and dpi for PDFs, so a logo or a plot with fine text can keep its resolution while everything else shrinks.

There is one ordering detail worth knowing: custom regex replacements are processed before \includegraphics commands are handled, so figures introduced by a replacement rule still get copied. That ordering is documented in the README and it is the kind of thing that would otherwise produce a cleaned folder with missing images.

## Installing arxiv-latex-cleaner with pip and running a first clean

The README states the package requires Python 3.9 or newer. Install it from PyPI:

```bash
pip install arxiv-latex-cleaner
```

On macOS there is also a Homebrew formula:

```bash
brew install arxiv_latex_cleaner
```

If you prefer to run from source, the README gives this sequence, which also prints the help text and confirms the module is importable:

```bash
git clone https://github.com/google-research/arxiv-latex-cleaner
cd arxiv-latex-cleaner/
python -m arxiv_latex_cleaner --help
```

The README also mentions installing as a command line program from the source tree with python setup.py install.

A first real run needs only the input folder. This is the README's example call, which resizes images to a longest side of 500 pixels and lets one image keep 2000 pixels:

```bash
arxiv_latex_cleaner /path/to/latex --resize_images --im_size 500 --images_allowlist='{"images/im.png":2000}'
```

After it finishes you should have a sibling folder named latex_arXiv next to your input folder. Check that it contains your .tex files, your referenced figures and nothing else, then compile it once before zipping. The README does not describe a dry-run mode, so the first run is also the moment you find out whether a file you needed was classified as unused.

## Driving the cleaner from cleaner_config.yaml

The command line gets long once you add resizing, PDF compression and deletion lists, so the README offers a config file path instead:

```bash
arxiv_latex_cleaner /path/to/latex --config cleaner_config.yaml
```

The repository ships a cleaner_config.yaml at the top level, and the README points to it for the pattern syntax. The documented shape of a pattern is a dictionary with pattern, insertion and description keys, where the insertion template is filled from named regex groups captured in the pattern. The README's example turns a \figcomp{path}{w1}{w2} command into a parbox wrapping an includegraphics call, using the named groups first, second and third.

This is the part of the tool that rewards reading before running. A pattern is a full regular expression, and a replacement that matches more text than you intended will rewrite your source in the cleaned copy. Keep the original folder intact, which the copy-to-a-new-folder design already does, and diff the cleaned .tex files against the originals after any pattern change.

## TikZ externalization and the filename contract

Papers with heavy TikZ pictures often carry source code or raw simulation data inside the tikzpicture environment, and the README treats hiding that as a feature. The cleaner replaces a tikzpicture environment with an \includegraphics pointing at a compiled PDF in an external folder, which you supply by compiling the pictures yourself. The README points to section 52 of the PGF/TikZ manual for the externalization library.

The constraint is strict: only environments preceded by a \tikzsetnextfilename{picture_name} command are replaced, and the externalized PDF must be named picture_name.pdf to match. If your document does not already use \tikzsetnextfilename, this feature does nothing for you, and adding it across a large paper is manual work. The README does not describe a fallback for pictures that were never externalized, so treat this as a feature for projects already set up that way rather than a retrofit.

## Where arxiv-latex-cleaner is the wrong tool

The deletion rules are heuristic and the README does not document an undo. Files are removed because they are not in the root and not included by another .tex file, or because no used .tex file references the image. A build that pulls in inputs through a mechanism the cleaner cannot see, such as a generated include list, a Makefile that assembles the document, or a figure path built from a macro, risks losing a file that the paper needs. The cleaned folder is the artifact to inspect, not the original, and there is no rollback command to run if the result is wrong: you re-run against the untouched source.

Ghostscript PDF compression is listed as Linux and Mac only, so Windows users get the other features but not that one. The tool also has nothing to say about the PDF you upload alongside the source, about arXiv's metadata fields, or about whether your document compiles. It cleans a source tree and stops there. If your problem is a broken bibliography or a font that will not embed, this is not the program that fixes it.

## How this differs from using latexmk or a packaging script

The nearest real alternative is a latexmk-driven build plus a hand-written shell script that deletes .aux and .log files and zips the result. That approach is more general: latexmk knows how to run the compilation chain, resolves dependencies through the LaTeX toolchain rather than by scanning for \include and \includegraphics, and will keep any file the build actually needs. It is also entirely manual on the parts this tool automates, since you write the comment stripping, the unused-file detection and the image resizing yourself, and comment stripping in particular is awkward to do correctly with sed given the \iffalse and comment-environment forms.

A converter such as pandoc sits in a different place: it translates between document formats and is not aimed at producing an arXiv-ready source folder. The cleaner's distinguishing choice is that it never compiles. It reasons about your project as text and file references, which is why it can run in seconds on a paper it cannot build, and also why a dynamically assembled document can defeat it. If your project already builds cleanly and you only need to delete auxiliary files, a three-line script is enough. If you need comments gone, unused figures gone, and images shrunk to fit a 50MB cap in one command, the cleaner is doing work you would otherwise redo before every submission.

## Maintenance, licence and what a version bump costs you

The repository is not archived. Its most recent push was on 2026-03-27, and the release list shows v1.0.11 on 2026-03-27, v1.0.10 on 2026-03-16, and v1.0.8 back on 2024-07-21. The gap between v1.0.8 and v1.0.10 is roughly twenty months, so the project moves in bursts rather than continuously, and you should not plan around frequent releases. The README's usage banner names the version, arxiv_latex_cleaner@v1.0.11, which is a small but useful signal that the help text tracks releases.

The runtime dependencies are listed in requirements.txt as absl_py, pillow, pyyaml and regex, all with loose or no upper bounds, which means a fresh install pulls current versions of each. For a tool you run once per submission, that is low risk. Upgrading means re-running the cleaner and diffing the output, since a change in image handling or deletion rules changes the cleaned folder rather than something you can unit-test in your own paper.

The licence is Apache-2.0, stated in the LICENSE file and in setup.py, which permits commercial and academic use with the usual attribution and notice conditions. This is a summary, not legal advice; if you are redistributing the tool or a modified version, read the LICENSE file itself.

## Conclusion

Adopt arxiv-latex-cleaner if you are preparing a LaTeX source upload for arXiv and want comments, unused files and oversized images handled in one pass instead of by hand. Do not adopt it as a general LaTeX build tool or a document converter; it copies and rewrites a source tree, it does not compile a PDF. Before your first real submission, verify two things: that your Python is 3.9 or newer, since the README states that is the floor, and that the cleaned folder still compiles, because the tool removes files it judges unused and a build system that resolves inputs dynamically can lose one.

## FAQ

### How do I use arxiv-latex-cleaner on my paper?

Run it against the folder holding your LaTeX code, for example arxiv_latex_cleaner /path/to/latex, and it writes a cleaned copy to a sibling folder named latex_arXiv. Add flags such as --resize_images and --im_size, or pass --config cleaner_config.yaml, to control resizing and pattern replacement.

### Does arxiv-latex-cleaner remove comments from my LaTeX source?

Yes. The README lists comment removal as a privacy feature and states that it also covers \begin{comment}\end{comment}, \iffalse\fi and \if0\fi blocks, not just ordinary commented lines.

### Can arxiv-latex-cleaner shrink my submission below the arXiv size limit?

It has size-oriented features aimed at the 50MB limit the README names: removing unused .tex files and unreferenced images, resizing images to a longest-side pixel count, compressing PDFs with ghostscript on Linux and macOS, and converting PNG to JPG.

### What happens to my TikZ pictures when I run arxiv-latex-cleaner?

Environments preceded by \tikzsetnextfilename{picture_name} are replaced with an includegraphics call pointing at an externally compiled picture_name.pdf in the folder you pass with --use_external_tikz. Pictures without that preceding command are not replaced.

## Sources

- [google-research/arxiv-latex-cleaner on GitHub](https://github.com/google-research/arxiv-latex-cleaner)
- [Issues](https://github.com/google-research/arxiv-latex-cleaner/issues)
- [License: Apache-2.0](https://github.com/google-research/arxiv-latex-cleaner/blob/main/LICENSE)
- [README](https://github.com/google-research/arxiv-latex-cleaner/blob/main/README.md)
- [Releases](https://github.com/google-research/arxiv-latex-cleaner/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/google-research-arxiv-latex-cleaner
