rmlint: finding duplicate files and other filesystem lint from the command line
Extremely fast tool to remove duplicates and other lint from your filesystem
At a glance
- What is it?
- rmlint is a C tool that scans a filesystem for duplicate files, duplicate directories, broken symlinks, empty files and more, then hands you a script or JSON instead of deleting anything itself. This review covers what it does, how to install and run it, and where it stops being the right tool.
- Who is it for?
- rmlint suits engineers and administrators who want a scriptable, non-interactive duplicate finder that runs on Linux, FreeBSD, Darwin and Solaris and can emit JSON, a shell script or a Python script for review before anything is touched. It is the wrong tool if you want a graphical file manager to delete files as it finds them, or if you need a documented rollback path, since the README does not document one.
- Can I use it commercially?
- Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository last received commits 11 days ago.
- What is it written in?
- Mainly C, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What rmlint actually looks for, and who needs it
rmlint is not only a duplicate finder. The feature list in the README names five categories: duplicate files and duplicate directories, nonstripped binaries (binaries that still carry debug symbols), broken symbolic links, empty files and directories, and files whose user or group ID no longer resolves. That mix is the point. A disk that has filled up usually has more than one cause, and a single traversal that reports all of them is cheaper than running four separate tools over the same tree.
The audience is anyone doing filesystem hygiene from a shell: backup operators checking whether a snapshot really added data, developers cleaning build trees and vendored copies, administrators hunting zero-byte files left behind by failed jobs. It is explicitly not interactive. The README lists "No interactivity" as a difference from other duplicate finders, which means it never asks you a question mid-scan and never deletes on its own. You get output, and you decide what to do with it. That is a deliberate design stance, and it is the main reason to pick rmlint over a file manager's built-in duplicate search.
How the scan pipeline works: traversal, hashing, output formats
The README describes the project as written in C and says it runs and compiles under most Unices, including Linux, FreeBSD, Darwin and Solaris, with Linux as the main target. It also warns that some optimisations may not be available on other platforms, which is worth taking at face value: the portability claim and the performance claim are not the same claim.
The comparison section lists the mechanisms that matter. Paranoia mode exists "for those who do not trust hashsums", meaning the default path compares files by hash and paranoia mode adds byte-level comparison as a second check. There is caching and replaying, so a scan can be stored and reused rather than recomputed. There is BTRFS support, which lets the tool take advantage of that filesystem's own facilities rather than treating it as a generic block device. There is an mtime filter, so you can restrict the search to files newer than a given modification time, and many output formats, which is what makes the tool scriptable at all.
What the README does not give is a walkthrough of the internal architecture: no description of the traversal order, the hashing algorithm, the thread model or the cache file format. If you need that level of detail before adopting it, the repository's src/ and lib/ directories are where you would have to read, because the top-level README does not cover it.
Installing rmlint and running a first scan
The README's installation section is short and defers to your distribution first: "Chances are that you might have rmlint already as readily made package in your favourite distribution. If not, you might consider compiling it from source." So the first step is checking your package manager rather than cloning anything.
The README does not name a package manager command, so the install path it points at is the source build instructions at rmlint.readthedocs.org under install.html, plus whatever your distribution provides. The repository uses SConstruct at the top level, which tells you the build system is SCons rather than make or CMake, but the README does not spell out the build command sequence, so follow the install page rather than guessing flags.
Once installed, the README does not print a canonical example command either. What it does document is the feature set: an mtime filter, many output formats, caching and replaying, and no interactivity. The documented entry points for learning the actual flags are the tutorial at rmlint.rtfd.org/en/latest/tutorial.html and the online manpage at rmlint.rtfd.org/en/latest/rmlint.1.html. Read those before running anything against a real tree, because the README itself stops at describing capabilities.
The output is a script, not a deletion
The most important behavioural fact in the README is the phrasing that rmlint "finds space waste and other broken things on your filesystem and offers to remove it." Offers. Combined with the explicit "No interactivity" bullet, this means the tool's job ends at producing a description of what could be removed, in one of its many output formats. The cleanup is a separate step that you perform.
That separation is what makes the tool safe to run against a real filesystem and also what makes it inconvenient. You cannot point it at a directory, walk away, and come back to a smaller directory. You have to read the output, or run the generated script, or feed the JSON into something else. For a one-off cleanup of a personal photo library, that is friction. For a scheduled job on a server, it is exactly the right shape, because the decision about what counts as a duplicate stays in your hands and in your version control.
The README does not document an undo command, a trash integration, or a rollback mechanism. If the generated script deletes the wrong file, nothing in the documented feature set brings it back. That is a property of handing you a script, not a bug, but it should shape how you test.
Where rmlint is the wrong choice
Three cases stand out.
First, small trees. The README's headline claim is speed, and speed matters when you are traversing millions of files. On a directory of a few hundred documents, the time you spend reading rmlint's output and deciding what to do with it will exceed the time the scan saved you over simply sorting by size in a file manager.
Second, non-Linux platforms where you need the fast path. The README is candid that the main target is Linux and that some optimisations might not be available elsewhere, while still claiming FreeBSD, Darwin and Solaris support. Those two statements together mean you should expect correct results but not necessarily the same throughput, and the README does not quantify the gap.
Third, anything where you want the tool to decide and act. If your workflow is "find duplicates and delete the older copy automatically", rmlint gives you the raw material for that but not the policy. You would be writing the policy yourself on top of its output. A tool with a built-in interactive deletion prompt fits that workflow better, at the cost of not being scriptable.
There is also a version caveat. The most recent release listed is v2.11.0-rc.0, a release candidate dated 2026-09-18, following v2.10.3 from 2025-03-22. If you install from a distribution package you are likely getting a stable release, not the candidate. The README does not describe what changed between them; CHANGELOG.md in the repository is where that would be.
Alternatives: fdupes, jdupes and rdfind
The related searches show people comparing rmlint against fdupes, jdupes and rdfind, so the comparison is worth making concretely.
fdupes and jdupes are duplicate finders in the narrow sense: they locate identical files and report them, and jdupes is a fork of fdupes that continued development after the original slowed. Their scope is duplicates. rmlint's scope is duplicates plus nonstripped binaries, broken symlinks, empty files and directories, and unresolvable user or group IDs. If all you need is a list of identical files, the narrower tools do that without the extra categories, and their output conventions are what most existing scripts already expect.
rdfind is closer in spirit: it also targets duplicate files and is designed for scripted use. The difference the rmlint README emphasizes is the combination of paranoia mode, caching and replaying, BTRFS support, mtime filtering and multiple output formats in one binary. Whether that combination is worth switching for depends on whether you actually use more than one of those features. If you only ever want "list duplicates as text", the extra machinery is weight you carry without using.
None of these comparisons is settled by the README, which does not benchmark against any of them. The honest position is that rmlint's distinguishing feature set is documented and its relative speed is asserted rather than shown in the README.
Maintenance, licensing and what to verify before adopting
The repository is not archived, and the last push was on 2026-09-19, which is recent. The release cadence visible in the release list is uneven: v2.10.2 in August 2023, v2.10.3 in March 2025, then v2.11.0-rc.0 in September 2026. The AUTHORS section shows a maintainer timeline with gaps, from Christopher Pahl (2010-2017) through Daniel Thomas (2014-2021) and Cebtenzzre (2021-2023) to Vassili Tchersky from 2025 onward. That history is relevant if you are planning to depend on the project long term: the current maintainer's tenure is the shortest in the list, and the release candidate status of the newest version means the stable line is still v2.10.3.
Licensing is GPL-3.0, stated in the README and distributed as the COPYING file in the repository root. The practical implication is the usual one for GPL: if you link rmlint's code into your own program and distribute it, the GPL's terms attach to that distribution. Running the rmlint binary as a separate process from your own tooling is a different situation from incorporating its source. This is not legal advice, and the COPYING file is the authoritative text.
Upgrade cost is low if you install from a distribution, since that is your package manager's problem. If you build from source, the SCons-based build means your toolchain needs SCons, and the README's install page is the reference rather than the README itself. The caching and replaying feature is the one upgrade consideration worth naming: if you rely on a stored scan, check whether a version change invalidates it before you upgrade in place.
Editorial conclusion
rmlint suits engineers and administrators who want a scriptable, non-interactive duplicate finder that runs on Linux, FreeBSD, Darwin and Solaris and can emit JSON, a shell script or a Python script for review before anything is touched. It is the wrong tool if you want a graphical file manager to delete files as it finds them, or if you need a documented rollback path, since the README does not document one. Before trusting it on a real tree, inspect the generated script and check whether your distribution already ships rmlint so you do not build from source unnecessarily.
Frequently asked questions
How do I use rmlint?
Install it from your distribution if a package exists, otherwise build from source as the README's install page describes. Then point it at a directory and read the output, which lists duplicate files and other issues rather than deleting anything. The tutorial at rmlint.rtfd.org covers the features in detail.
Does rmlint delete files by itself?
No. The README says rmlint finds space waste and offers to remove it, and lists no interactivity as a deliberate difference from other duplicate finders. It produces output in one of several formats, and acting on that output is a separate step you take.
What is the rmlint GUI?
The repository contains a gui/ directory at the top level, and the related searches include rmlint --gui, so a graphical front end is part of the project. The README itself does not document the GUI, so the docs at rmlint.rtfd.org are the place to check what it offers.
Which platforms does rmlint run on?
The README states it runs and compiles under most Unices, including Linux, FreeBSD, Darwin and Solaris, with Linux as the main target. It adds that some optimisations might not be available on other platforms.
What licence is rmlint under?
GPL-3.0, according to the README, with the full text in the COPYING file distributed with the source.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/sahib-rmlint)