Open-source project
adrianlopezroche/fdupes avatar
adrianlopezroche/fdupes

fdupes: Finding and Deleting Duplicate Files from the Command Line

FDUPES is a program for identifying or deleting duplicate files residing within specified directories.

3,016 stars221 forksCNOASSERTION

At a glance

What is it?
fdupes is a C program that scans directories for duplicate files and can delete them interactively. It is aimed at people already comfortable in a shell, and its deletion mode carries documented data-loss risks worth reading before you run it.
Who is it for?
Adopt fdupes if you work in a shell, want grouped duplicate output you can pipe into other tools, and are willing to read the manpage before using --delete. Do not adopt it if you need a graphical interface or a Windows-native binary: the README documents no GUI, and Windows users are expected to run it under a POSIX layer.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 169 days ago.
What is it written in?
Mainly C, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What fdupes Solves, and Who It Is Aimed At

Disk space fills up quietly. Photo libraries accumulate second copies after imports, source trees collect build artifacts that match files elsewhere, and backup drives end up holding the same archive under three names. fdupes exists to answer one question across a set of directories: which files are byte-identical to each other? The README states the purpose plainly, calling it a program for identifying or deleting duplicate files residing within specified directories. That is the whole scope. It does not deduplicate at the filesystem level, does not compress, and does not index your machine in the background.

The audience is people who live in a terminal. Every option is a short flag, output is line-oriented, and there is no configuration file to learn. If you are comfortable with find, xargs and sort, fdupes fits the same mental model. If you want a window with thumbnails and a recycle bin, this is the wrong tool, and the README offers no GUI to point you toward.

How fdupes Decides Two Files Are Duplicates

The repository layout shows the mechanism split across several C modules: dir.c walks directories, fmatch.c handles file matching, confirmmatch.c performs byte-for-byte confirmation, and hashdb.c backs the optional cache. A separate md5/ directory sits at the top level, which tells you the signature used for the first pass is MD5. So the flow is a cheap filter followed by an expensive check: candidates are grouped by signature, and then, unless you ask otherwise, the program confirms equality by reading the bytes.

Two flags change that flow in ways worth understanding. --quicksummary, documented as skipping the slower byte-for-byte match confirmation, trusts the signature alone and is therefore faster and weaker. --deferconfirmation pushes the byte comparison to just before deletion in interactive mode. The cache is the other moving part. With --cache, signatures are stored in a database so repeat runs can skip re-reading files, and the cache has its own sub-options for pruning orphaned entries, clearing everything, or vacuuming the database file. Those cache options can be used without passing a directory at all, which makes them maintenance commands in their own right.

Default behaviour around links is worth noting. The README says that normally, when two or more files point to the same disk area, they are treated as non-duplicates, and that --hardlinks changes this. Symlinks are not followed unless you pass --symlinks.

Installing fdupes and Running a First Scan

The repository ships an INSTALL file and an INSTALL.enduser file, plus configure.ac and Makefile.am, so the documented path is the usual Autotools sequence. You need a C toolchain and a POSIX environment. On Linux distributions fdupes is commonly available as a package, though this repository does not list package names, so check your own distribution's index rather than trusting a name from a blog post.

The README gives the usage line as fdupes [options] DIRECTORY..., so a first scan is just a directory argument:

bash
fdupes -r -S DIRECTORY

Here -r follows subdirectories and -S shows the size of each duplicate file. The README explains the output format: unless -1 or --sameline is given, duplicate files are listed together in groups, each file on its own line, with blank lines separating the groups. Read that output before doing anything else. It is the inventory you are about to act on.

For scripted use, -1 puts each set of matches on a single line, and the README notes that spaces and backslash characters in filenames are then escaped with a backslash. That escaping matters if you plan to split the line and feed the paths to another command.

The README does not spell out a build command sequence, so consult INSTALL and INSTALL.enduser in the repository before compiling.

The Deletion Mode Is Where fdupes Gets Dangerous

The --delete flag is the reason to read the manpage rather than skim the help text. The README's own wording says that under particular circumstances, data may be lost when using this option together with -s or --symlinks, or when specifying a particular directory more than once. It then explains the failure modes rather than hand-waving them. In the symlink case, a user could accidentally preserve a symlink while deleting the file it points to. In the repeated-directory case, every file inside that directory is listed as its own duplicate, so preserving a file without its duplicate means preserving nothing and deleting the original.

That second case is the one that catches people. Passing the same path twice, or passing a parent and a child directory, produces output that looks like a normal duplicate set. It is not. There is no guard documented in the README that detects this and refuses to proceed.

The safer flags are the ones that remove the human from the loop in a controlled way. --noprompt, used with --delete, preserves the first file in each set and deletes the rest without asking. --immediate deletes duplicates as they are encountered and implies --noprompt, which is faster and gives you no review step at all. --log writes deletion choices to a file, and --order controls which file lands first in each group, by modification time, status change time, or filename, with --reverse flipping the sort. If you are going to automate deletion, --order plus --log plus --noprompt is the combination that leaves a record of what happened.

fdupes Compared with jdupes, rdfind and fclones

The most common comparison is with jdupes, a fork that has diverged over time. The practical difference visible from this repository is that fdupes keeps its matching pipeline in C modules with an MD5 first pass and an optional signature database, and it ships an ncurses-based interactive deletion interface built from the ncurses-*.c files in the tree. That interactive screen-mode prompt is a specific fdupes feature, and --plain exists to fall back to the older line-based prompt instead.

rdfind and fclones take a different angle: they are built around the idea of a ranked or automated deduplication run rather than an interactive review. fdupes defaults to showing you groups and letting you decide. That is a design choice, not a deficiency, but it means fdupes is a worse fit when you want a single unattended pass over a large tree with a policy applied consistently.

For graphical needs, the README points nowhere. Tools like czkawka or dupeguru operate in that space, and fdupes has no equivalent. If your workflow involves looking at files before deciding, and you are not comfortable doing that in a terminal, fdupes is the wrong layer of the stack.

Filtering, Ordering and the Cache

A raw scan of a home directory is noisy, and fdupes gives you filters rather than a query language. --minsize and --maxsize bound the byte range under consideration, which is how you avoid matching thousands of tiny config files. --noempty excludes zero-length files, and --nohidden excludes hidden files. --permissions refuses to treat files with different owner, group or permission bits as duplicates, which is useful when identical content has deliberately different access rights.

The cache deserves separate attention because it changes the cost profile of repeated runs. Keeping signatures in a database means the second scan over the same tree does less I/O. The trade-off is state: the database can accumulate entries for files that no longer exist, which is what the prune option is for, and it can grow, which is what vacuum addresses. A cache that is stale in the wrong direction is worse than no cache, so if you enable it, plan to run the maintenance sub-options occasionally rather than never.

Note also that --summarize and --quicksummary exist for reporting only. --quicksummary explicitly skips the byte-for-byte confirmation, so its numbers are an estimate built on signatures, not a verified count.

Licence and Maintenance Status

The repository's licence field reports NOASSERTION, which means the hosting platform could not classify the licence file automatically. The tree does contain a LICENSE file, so the terms exist, but this article will not guess at them. If you plan to redistribute fdupes or ship it inside a product, read that file directly and get your own advice; nothing here substitutes for that.

On activity, the last push to the master branch was on 2026-04-14, and the most recent tagged release in the repository is v2.4.0 from 2025-03-30. The repository is not archived. Those two facts are the whole of what can be said: the project has seen commits within the last several months and a release within the last year and a half. The README does not document an upgrade procedure, and there is no migration guide in the repository, so upgrading between versions means rebuilding from source with the same Autotools sequence and reading CHANGES for behaviour differences. If you depend on a specific flag's semantics, pin the version you built.

Editorial conclusion

Adopt fdupes if you work in a shell, want grouped duplicate output you can pipe into other tools, and are willing to read the manpage before using --delete. Do not adopt it if you need a graphical interface or a Windows-native binary: the README documents no GUI, and Windows users are expected to run it under a POSIX layer. Verify two things first: that your build produces the version you expect from ./configure and make, and that you understand the symlink and repeated-directory warnings attached to --delete, because those are the paths where the documentation says data may be lost.

Frequently asked questions

What does fdupes do?

It identifies duplicate files within directories you specify, listing matches in groups, and can optionally delete duplicates interactively or without prompting. The README describes it as a program for identifying or deleting duplicate files residing within specified directories.

How do I install fdupes?

From a source checkout the repository provides INSTALL and INSTALL.enduser files with configure.ac and Makefile.am, so the documented route is the Autotools sequence those files describe. Many Linux distributions also package it, though this repository does not name specific packages.

How do I use fdupes?

Run it with one or more directory arguments, for example fdupes -r -S DIRECTORY to recurse into subdirectories and show file sizes. Without -1 or --sameline the output lists each group of duplicates on separate lines, with blank lines between groups.

Is fdupes safe to use?

Scanning is non-destructive, but the README warns that data may be lost when --delete is combined with -s or --symlinks, or when the same directory is specified more than once, because files may be listed as their own duplicates. Using --log and reviewing groups before deleting reduces the risk.

What are the main differences between fdupes and jdupes?

jdupes is a fork of fdupes that has diverged. In this repository, fdupes uses an MD5 first pass with byte-for-byte confirmation and ships an ncurses-based interactive deletion interface, with --plain falling back to the older line-based prompt.

Official sources

  1. adrianlopezroche/fdupes on GitHub
  2. Issues
  3. README
  4. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/adrianlopezroche-fdupes.svg)](https://hysenlabs.com/projects/adrianlopezroche-fdupes)