fastdup: the Dockerfile installs Python 3.8 and never copies the source
fastdup is a powerful, free tool designed to rapidly generate valuable insights from image and video datasets. It helps enhance the quality of both images and labels, while significantly reducing data operation costs, all with unmatched scalability.
At a glance
- What is it?
- visual-layer/fastdup finds duplicates, outliers, mislabels and broken images in image and video datasets, and can delete duplicates from disk. Its container builds Python 3.8 from a deadsnakes repository and pip installs the published package instead of the tree, its licence record disagrees with its own badge, and its newest release is from June 2024.
- Who is it for?
- fastdup fits someone with a large image or video dataset who needs a duplicate and outlier report before training, and who has the disk space to let a nearest neighbour index run over it. Its deleting convenience is a separate decision: a function that removes duplicates from disk at a similarity threshold is not something to point at an archive without dry_run first.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 47 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The Dockerfile installs Python 3.8 and installs fastdup from PyPI
The container recipe is short enough to read in full, and it does not do what you would expect. It starts from ubuntu:20.04, adds the deadsnakes repository for extra Python versions, installs python3.8, then libopencv-dev and libgl1, then pip, upgrades pip, and finishes with a single install line pulling in fastdup together with matplotlib, matplotlib-inline, torchvision, pillow and pyyaml. Two things are missing that a build of this repository would need. There is no COPY of the project, so the fastdup inside the image comes from PyPI rather than from the source sitting next to the Dockerfile. And there is no CMD or ENTRYPOINT, so the image ends as an environment with no command to run. The version story also disagrees with the documentation: the badge advertises Python 3.9, 3.10, 3.11 and 3.12, while the container builds 3.8 on a distribution that has been past end of life for years.
One call deletes from disk, and the knob is called distance
Duplicate removal is a single call rather than a report you export:
import fastdup
fastdup.remove_duplicates("IMAGE_FOLDER/")The documentation is clear about the default, which is similarity above 0.96, and clear that files are found and deleted directly from disk. There is a dry_run flag to preview which files would be removed, and a parameter to move the threshold. The naming is where the friction is: the prose calls the value a similarity threshold while the parameter is called `distance`, so a reader who assumes the direction of the inequality will get the opposite of what they intended, and the default of 0.96 is more natural as a similarity than as a distance. The reporting path is better behaved, since it only writes static HTML galleries and leaves the files where they are.
Releases are per platform builds and the newest is from June 2024
The release list is not a version stream, it is a set of platform builds. The two newest tags are v2.2_3.8 and v2.2_3.7, both titled Centos 7.0.9 Stable release, published on 2024-06-14 two seconds apart. Below them sits v1.119, titled Ubuntu 18, from 2024-04-04. So the tag carries a version, a separator and something else, the release title carries an operating system and its version, and there is no single artefact that represents the current code for every platform. CentOS 7 and Ubuntu 18 are both long past their support windows, which makes the naming a historical record of what was built rather than a current support matrix. Meanwhile the branch's last recorded push is 2026-08-23, more than two years after the newest release, so the tree and the downloadable releases have been diverging for a long time.
The badge says CC BY-NC-ND and the repository record says nothing
The licence is stated in two places that do not agree. The README badge reads License CC BY-NC-ND 4.0 and links to the LICENSE file in the repository, while the repository record for this project carries no recognised licence identifier at all. Nothing in the tree resolves it, and this article will not choose. What can be said is that the terms in the badge are restrictive in two ways that matter for a tool like this. Non commercial rules out using it inside a product, and no derivatives rules out modifying the code, which is an unusual pairing for a package published on an index that people pip install as a dependency. The badge also carries an operating system line for macOS, Linux and Windows through WSL2, which is worth noting next to the licence question, since WSL2 is a qualified rather than native Windows path.
400 million images on one CPU is a claim with no method attached
The differentiators are listed as five qualities, and three of them are measurable while two are not. Quality covers duplicates and near duplicates, outliers, mislabels, broken images and low quality images. Privacy means the tool runs locally or on your own cloud infrastructure. Ease of use covers labeled and unlabeled data in image or video form across the three operating systems. Scale is stated as capable of processing 400M images on a single CPU machine, scaling up to billions. Speed is credited to an optimised C++ engine performing on low resource CPU machines. Neither of the last two is accompanied by a hardware description, a dataset description, an elapsed time or a link to a benchmark in the text of this README, so the numbers are there to be believed rather than checked. The nearest thing to evidence is a separate gallery and a video tutorial linked after the quickstart.
Five galleries, twenty five notebooks, and two with the same name
The example notebooks do most of the explaining. There are twenty five of them, and the naming tells you the scope: analysing datasets from Hugging Face, Kaggle, Roboflow, Labelbox, TensorFlow Datasets and torchvision, cleaning an image dataset, finding and removing duplicates, finding and removing mislabels, generating captions with a BLIP model on LAION captions, enrichment with zero shot classification, detection and segmentation, embeddings from ONNX DINOv2 and from timm, feature vectors, heatmaps, image search and optical character recognition. Two of them are the same notebook twice: finding-removing-duplicates.ipynb and finding_removing_duplicates.ipynb differ only in whether the separators are hyphens or underscores. That is a small blemish, but it is the kind that makes a newcomer wonder which one is maintained, and it is the only duplication visible in a directory otherwise named carefully.
Three links point at the wrong repository, and the homepage says old
The header has link problems worth listing, because two of them send a reader to software that is not this project. The sentence crediting the founders links XGBoost and Apache TVM to the same URL, the Apache TVM repository, so the XGBoost link lands somewhere other than XGBoost. The contributors badge points at a third party README template repository rather than at this repository's own contributors graph. The site link in the repository record goes to a documentation path whose own name contains fastdup_docs_old, while the navigation inside the README links a different documentation host. Five documentation files sit at the root beside the README, covering installation, examples, running, cloud and release notes, so the material is there; it is the entry points that disagree about where to start. There is also a commented out block of social and discussion links left in the markup, which is how a reader ends up with several dead anchors and no working contact.
Editorial conclusion
fastdup fits someone with a large image or video dataset who needs a duplicate and outlier report before training, and who has the disk space to let a nearest neighbour index run over it. Its deleting convenience is a separate decision: a function that removes duplicates from disk at a similarity threshold is not something to point at an archive without dry_run first. Before using it, check four things. Which Python you are on, since the badge advertises 3.9 through 3.12 while the container pins 3.8. Which licence applies, since the badge says CC BY-NC-ND 4.0 while the repository record says nothing at all. Which version you get, since the newest release predates the last push by more than two years. And what your container actually contains, because the Dockerfile builds an environment from PyPI rather than from the checkout next to it, which is a container for the published package wearing this repository's name.
Frequently asked questions
What does fastdup detect in a dataset?
It handles labeled or unlabeled image and video datasets and reports duplicates and near duplicates, outliers, mislabels, broken images and low quality images such as dark, bright or blurry ones. Results are shown as static HTML galleries, including duplicates, outliers, connected components, statistics and similar images.
How do I install fastdup?
With `pip install fastdup` from PyPI. The badge advertises Python 3.9, 3.10, 3.11 and 3.12, and the operating systems listed are macOS, Linux and Windows through WSL2. A Dockerfile is included, but note that it builds Python 3.8 and installs the published package from PyPI rather than building the source in the repository.
Is it safe to let fastdup delete duplicate files?
The convenience call does delete from disk at a similarity above 0.96 by default. The documentation points at `dry_run=True` to preview which files would be removed first, and the threshold is adjustable, so run the preview before the deletion on anything you cannot recreate.
Which licence is fastdup under?
The README badge reads CC BY-NC-ND 4.0 and links to the LICENSE file, while the repository record carries no recognised licence identifier. Since non commercial and no derivatives are both restrictive, read the LICENSE file before using it in a product or modifying it.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/visual-layer-fastdup)