arxiv-latex-cleaner: preparing a LaTeX paper for arXiv submission
arXiv LaTeX Cleaner: Easily clean the LaTeX code of your paper to submit to arXiv
At a glance
- What is it?
- A Google Research command line tool that copies a LaTeX project into a cleaned folder, stripping comments, unused files and oversized images before you upload the ZIP to arXiv. It is narrow by design, and the narrowness is the point.
- Who is it for?
- Adopt arxiv-latex-cleaner if you are preparing a LaTeX source upload for arXiv and want comments, unused files and oversized images handled in one pass instead of by hand. Do not adopt it as a general LaTeX build tool or a document converter; it copies and rewrites a source tree, it does not compile a PDF.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The submission checklist problem arxiv-latex-cleaner targets
arXiv accepts a source upload, not just a PDF, and the source you keep on your laptop is rarely the source you want to publish. It carries commented-out paragraphs, notes to coauthors, stale figure files, and a bibliography you may or may not want to ship. There is also a 50MB limit on submissions, which the README names directly as the reason for the size-oriented features. So the last hour before a deadline is usually spent deleting comments by hand, chasing which .png files are actually referenced, and guessing at image resolution until the ZIP fits.
This tool automates that pass. You point it at a folder of LaTeX code, for example /path/to/latex/, and it writes a new folder /path/to/latex_arXiv/ that the README describes as ready to ZIP and upload. The audience is narrow and clear: authors of LaTeX papers who are submitting to arXiv and who want the source cleaned without editing the original project. It is not a typesetting system, not a PDF builder, and not useful to anyone working in Word or Markdown.
What the cleaner actually rewrites in your source tree
The mechanism is a source-to-source copy with rules applied along the way. Nothing in the README suggests the tool compiles your document, so the cleaning is textual and file-level.
On the privacy side, it removes auxiliary files such as .aux, .log and .out, and it strips comments. The comment removal is broader than a line-by-line pass: the README states it also handles \begin{comment}\end{comment}, \iffalse\fi and \if0\fi blocks, which are the usual ways authors hide text. It can additionally delete user-defined commands you name with commands_to_delete, the README's example being a \todo{} macro you redefine as empty.
On the size side, the tool deletes .tex files that are neither in the root nor included by another .tex file, and images that no used .tex file references. It can resize images to a longest-side pixel count, compress PDFs with ghostscript on Linux and macOS, and convert PNG to JPG. The --images_allowlist flag takes a dictionary mapping a path to a per-file size, pixels for images and dpi for PDFs, so a logo or a plot with fine text can keep its resolution while everything else shrinks.
There is one ordering detail worth knowing: custom regex replacements are processed before \includegraphics commands are handled, so figures introduced by a replacement rule still get copied. That ordering is documented in the README and it is the kind of thing that would otherwise produce a cleaned folder with missing images.
Installing arxiv-latex-cleaner with pip and running a first clean
The README states the package requires Python 3.9 or newer. Install it from PyPI:
pip install arxiv-latex-cleanerOn macOS there is also a Homebrew formula:
brew install arxiv_latex_cleanerIf you prefer to run from source, the README gives this sequence, which also prints the help text and confirms the module is importable:
git clone https://github.com/google-research/arxiv-latex-cleaner
cd arxiv-latex-cleaner/
python -m arxiv_latex_cleaner --helpThe README also mentions installing as a command line program from the source tree with python setup.py install.
A first real run needs only the input folder. This is the README's example call, which resizes images to a longest side of 500 pixels and lets one image keep 2000 pixels:
arxiv_latex_cleaner /path/to/latex --resize_images --im_size 500 --images_allowlist='{"images/im.png":2000}'After it finishes you should have a sibling folder named latex_arXiv next to your input folder. Check that it contains your .tex files, your referenced figures and nothing else, then compile it once before zipping. The README does not describe a dry-run mode, so the first run is also the moment you find out whether a file you needed was classified as unused.
Driving the cleaner from cleaner_config.yaml
The command line gets long once you add resizing, PDF compression and deletion lists, so the README offers a config file path instead:
arxiv_latex_cleaner /path/to/latex --config cleaner_config.yamlThe repository ships a cleaner_config.yaml at the top level, and the README points to it for the pattern syntax. The documented shape of a pattern is a dictionary with pattern, insertion and description keys, where the insertion template is filled from named regex groups captured in the pattern. The README's example turns a \figcomp{path}{w1}{w2} command into a parbox wrapping an includegraphics call, using the named groups first, second and third.
This is the part of the tool that rewards reading before running. A pattern is a full regular expression, and a replacement that matches more text than you intended will rewrite your source in the cleaned copy. Keep the original folder intact, which the copy-to-a-new-folder design already does, and diff the cleaned .tex files against the originals after any pattern change.
TikZ externalization and the filename contract
Papers with heavy TikZ pictures often carry source code or raw simulation data inside the tikzpicture environment, and the README treats hiding that as a feature. The cleaner replaces a tikzpicture environment with an \includegraphics pointing at a compiled PDF in an external folder, which you supply by compiling the pictures yourself. The README points to section 52 of the PGF/TikZ manual for the externalization library.
The constraint is strict: only environments preceded by a \tikzsetnextfilename{picture_name} command are replaced, and the externalized PDF must be named picture_name.pdf to match. If your document does not already use \tikzsetnextfilename, this feature does nothing for you, and adding it across a large paper is manual work. The README does not describe a fallback for pictures that were never externalized, so treat this as a feature for projects already set up that way rather than a retrofit.
Where arxiv-latex-cleaner is the wrong tool
The deletion rules are heuristic and the README does not document an undo. Files are removed because they are not in the root and not included by another .tex file, or because no used .tex file references the image. A build that pulls in inputs through a mechanism the cleaner cannot see, such as a generated include list, a Makefile that assembles the document, or a figure path built from a macro, risks losing a file that the paper needs. The cleaned folder is the artifact to inspect, not the original, and there is no rollback command to run if the result is wrong: you re-run against the untouched source.
Ghostscript PDF compression is listed as Linux and Mac only, so Windows users get the other features but not that one. The tool also has nothing to say about the PDF you upload alongside the source, about arXiv's metadata fields, or about whether your document compiles. It cleans a source tree and stops there. If your problem is a broken bibliography or a font that will not embed, this is not the program that fixes it.
How this differs from using latexmk or a packaging script
The nearest real alternative is a latexmk-driven build plus a hand-written shell script that deletes .aux and .log files and zips the result. That approach is more general: latexmk knows how to run the compilation chain, resolves dependencies through the LaTeX toolchain rather than by scanning for \include and \includegraphics, and will keep any file the build actually needs. It is also entirely manual on the parts this tool automates, since you write the comment stripping, the unused-file detection and the image resizing yourself, and comment stripping in particular is awkward to do correctly with sed given the \iffalse and comment-environment forms.
A converter such as pandoc sits in a different place: it translates between document formats and is not aimed at producing an arXiv-ready source folder. The cleaner's distinguishing choice is that it never compiles. It reasons about your project as text and file references, which is why it can run in seconds on a paper it cannot build, and also why a dynamically assembled document can defeat it. If your project already builds cleanly and you only need to delete auxiliary files, a three-line script is enough. If you need comments gone, unused figures gone, and images shrunk to fit a 50MB cap in one command, the cleaner is doing work you would otherwise redo before every submission.
Maintenance, licence and what a version bump costs you
The repository is not archived. Its most recent push was on 2026-03-27, and the release list shows v1.0.11 on 2026-03-27, v1.0.10 on 2026-03-16, and v1.0.8 back on 2024-07-21. The gap between v1.0.8 and v1.0.10 is roughly twenty months, so the project moves in bursts rather than continuously, and you should not plan around frequent releases. The README's usage banner names the version, [email protected], which is a small but useful signal that the help text tracks releases.
The runtime dependencies are listed in requirements.txt as absl_py, pillow, pyyaml and regex, all with loose or no upper bounds, which means a fresh install pulls current versions of each. For a tool you run once per submission, that is low risk. Upgrading means re-running the cleaner and diffing the output, since a change in image handling or deletion rules changes the cleaned folder rather than something you can unit-test in your own paper.
The licence is Apache-2.0, stated in the LICENSE file and in setup.py, which permits commercial and academic use with the usual attribution and notice conditions. This is a summary, not legal advice; if you are redistributing the tool or a modified version, read the LICENSE file itself.
Editorial conclusion
Adopt arxiv-latex-cleaner if you are preparing a LaTeX source upload for arXiv and want comments, unused files and oversized images handled in one pass instead of by hand. Do not adopt it as a general LaTeX build tool or a document converter; it copies and rewrites a source tree, it does not compile a PDF. Before your first real submission, verify two things: that your Python is 3.9 or newer, since the README states that is the floor, and that the cleaned folder still compiles, because the tool removes files it judges unused and a build system that resolves inputs dynamically can lose one.
Frequently asked questions
How do I use arxiv-latex-cleaner on my paper?
Run it against the folder holding your LaTeX code, for example arxiv_latex_cleaner /path/to/latex, and it writes a cleaned copy to a sibling folder named latex_arXiv. Add flags such as --resize_images and --im_size, or pass --config cleaner_config.yaml, to control resizing and pattern replacement.
Does arxiv-latex-cleaner remove comments from my LaTeX source?
Yes. The README lists comment removal as a privacy feature and states that it also covers \begin{comment}\end{comment}, \iffalse\fi and \if0\fi blocks, not just ordinary commented lines.
Can arxiv-latex-cleaner shrink my submission below the arXiv size limit?
It has size-oriented features aimed at the 50MB limit the README names: removing unused .tex files and unreferenced images, resizing images to a longest-side pixel count, compressing PDFs with ghostscript on Linux and macOS, and converting PNG to JPG.
What happens to my TikZ pictures when I run arxiv-latex-cleaner?
Environments preceded by \tikzsetnextfilename{picture_name} are replaced with an includegraphics call pointing at an externally compiled picture_name.pdf in the folder you pass with --use_external_tikz. Pictures without that preceding command are not replaced.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/google-research-arxiv-latex-cleaner)