CLI tool
Serene-Arc/bulk-downloader-for-reddit avatar
Serene-Arc/bulk-downloader-for-reddit

Bulk Downloader for Reddit: three modes, three config layers, and no release since 2023

Downloads and archives content from reddit

2,615 stars240 forksPythonGPL-3.0

At a glance

What is it?
Serene-Arc/bulk-downloader-for-reddit installs from PyPI as bdfr and does two different jobs under three command names. Its packaging metadata still points at the previous owner's namespace, its newest release is v2.6.2 from January 2023, and the version number lives in code rather than in pyproject.toml.
Who is it for?
This is a capable archiver with an honest warning attached to its own clone command, and it remains the right tool for pulling a subreddit's linked media or a submission's comment tree into files you control. Two things decide whether it fits your setup.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 176 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Two names for one tool: bdfr on PyPI, bulk-downloader-for-reddit on GitHub

The name you type and the name you search for are different strings. The distribution on PyPI is `bdfr`, which is also the console command, so installation is a one word operation:

bash
python3 -m pip install bdfr --upgrade

The repository is Serene-Arc/bulk-downloader-for-reddit, and the package description inside pyproject.toml repeats the repository's own one-liner, Downloads and archives content from reddit. The metadata records Ali Parlakci as author and Serene Arc as maintainer, so the project changed hands without changing its package name, which is the sensible choice for anyone with `pip install bdfr` already in a script. Three other install paths exist and each costs something different: pipx installs the same package into an isolated environment, Arch users have two AUR packages, one tracking the latest release and one tracking development builds from git, and shell completions are a separate `bdfr completions` step after installation rather than something you get during it.

Every link in pyproject.toml still points at aliparlakci

The URLs section of pyproject.toml lists the homepage, the source repository and the bug tracker, and all three are built on `github.com/aliparlakci/bulk-downloader-for-reddit`. The repository those links describe is not where the code lives. Code, pull requests and the release tags sit under Serene-Arc, and the three most recent tags, v2.6.0, v2.6.1 and v2.6.2, are all published from the Serene-Arc namespace. The continuous integration badge in the documentation points at the workflow file through the same old path, so the badge reports on a workflow URL that has moved. Two other homepage values disagree with each other as well: the repository metadata names the PyPI project page as the homepage while pyproject.toml names a pages site under the previous namespace. Nothing in the tree updates these automatically, because they are hand-written strings, which means a reader who follows the documented bug tracker link is trusting a URL that the project stopped maintaining when it moved.

A Production/Stable classifier over a 2023 release

The classifier list in pyproject.toml includes Development Status :: 5 - Production/Stable. That is a statement about maturity, and the release history is a statement about time: v2.6.0 on 2022-09-27, v2.6.1 on 2022-12-03, and v2.6.2 on 2023-01-31, with nothing tagged after that. The repository itself is not frozen, since the last push is dated 2026-04-12 and the tracker holds 112 open issues against 2,613 stars, so the gap is between the released artifact and the work in progress rather than between an active project and an abandoned one. A user installing from PyPI gets v2.6.2 and its dependency set from that date. An Arch user who picks the development AUR package gets something newer and gets whatever the untagged tree happens to do. The two are different products sharing a name, and the classifier does not distinguish them.

The version number lives in code, not in pyproject.toml

The packaging metadata declares the version as dynamic, then points at an attribute inside the package itself to fill it in. pyproject.toml therefore contains no version string to read, and a checkout on disk carries no number either until something imports the package and asks. That has a practical consequence for anyone comparing a checkout to a tag: the only way to learn the version is to install it and run `bdfr --version`, so a source reader cannot confirm from the tree which release they are looking at. The same pattern means the version string is edited in one place and picked up by every build, which is tidy until a contributor forgets, and then the artifact on PyPI silently disagrees with the tag name. Everything else in the metadata is pinned down: `requires-python` is 3.9 or above, and the classifier list names only 3.9, 3.10 and 3.11, so a machine on a newer interpreter is permitted by the constraint but untested by the declaration.

yt-dlp decides whether a link still resolves

The dependency list is where this project's real exposure sits. Alongside appdirs, beautifulsoup4, click, dict2xml, praw, pyyaml and requests, there is yt-dlp with a floor of 2022.11.11, a version number that dates the constraint to the same window as the last release. yt-dlp is the component that actually retrieves media behind a submission, and it is the component whose own upstream moves fast enough that a pinned lower bound is the normal shape of the dependency rather than a choice. The packaging metadata names four packages for installation, the top level package, an archive_entry package, a site_downloaders package and a fallback_downloaders package beneath it, plus one data file, `bdfr/default_config.cfg`, installed under the name `config`. Per-site handlers living in their own packages is the acknowledgement that a link from any given host needs its own handling, and the fallback package exists for the hosts with none.

clone is not a clone, and the documentation says so

Three command names cover two jobs. `download` fetches the resource a submission links to, images and video, while `archive` stores the submission itself: details, upvotes, text, statistics and the comments, serialised as JSON, XML or YAML. `clone` runs both in one pass and is described as more efficient than running the two separately, which is a claim about the network round trips rather than about fidelity. The documentation is unusually direct about the limit: the clone command is not a true, faithful clone of Reddit, it retrieves much of the raw data Reddit provides, and a true clone needs a different tool, HTTrack by name. That sentence is the most useful line in the whole usage guide, because the word clone invites exactly the wrong expectation. Anyone whose actual goal is a browsable local copy of a site should read that paragraph before planning around the command.

Command line beats opts file beats installed config

Options can arrive from three places, and the precedence between them is stated rather than left to discovery: when the same option appears both in the YAML file and as a command line argument, the command line wins, and the opts file sits above the global config file. The YAML spelling also changes, with `file_scheme` in the file against `--file-scheme` on the command line, which is the sort of difference that produces a silently ignored key. This example is equivalent to the long command below it:

yaml
skip: [mp4, avi]
file_scheme: "{UPVOTES}_{REDDITOR}_{POSTID}_{DATE}"
limit: 10
sort: top
subreddit:
  - EarthPorn
  - CityPorn

Any option that can be repeated on the command line is written as a YAML list, exactly as `subreddit` is above. The installed default config ships as a data file, `--config` points at an alternative one, and appdirs is in the dependency list for per-user config directories, so a user-level override sits in the same precedence chain. Two typos sit in the option descriptions themselves, where the opts file is described as having higher prority and the example is called equilavent.

Filename schemes decide what lands on disk

The options that shape output rather than fetch it are where a run quietly does something you did not intend. `--filename-restriction-scheme` takes `windows` or `linux` and exists to turn off the operating system's own detection, so a Linux filesystem can still be given names that a Windows host would reject, or the reverse. Scheme placeholders like `{POSTID}`, `{UPVOTES}` and `{DATE}` decide whether a rerun collides with an earlier one, and the archive example sets an empty folder scheme to flatten the output. Repeatable options accumulate rather than replace: `--skip mp4 --skip avi` in the expanded command becomes the `skip: [mp4, avi]` list, `--disable-module` takes module names, and `--include-id-file` reads files holding one submission ID per line. `--log` sets the logfile path and is called out as required when running several instances at once, which is how a parallel run avoids two processes writing the same log.

Editorial conclusion

This is a capable archiver with an honest warning attached to its own clone command, and it remains the right tool for pulling a subreddit's linked media or a submission's comment tree into files you control. Two things decide whether it fits your setup. If you need a version pinned to a known date, remember that the newest tagged release is v2.6.2 from 2023-01-31 while commits have continued to 2026-04-12, and that yt-dlp is the dependency most likely to decide whether a link still resolves. If you plan to file bugs or follow releases, check the links in the packaging metadata first, because the project URLs still resolve to the aliparlakci namespace rather than the Serene-Arc one where the code and the release tags now live. Read the config precedence before your first run: the command line overrides the opts YAML file, which overrides the installed default config, and that ordering decides where a surprising filename pattern comes from.

Frequently asked questions

How do you use Bulk Downloader for Reddit?

It has three modes: download fetches the resource linked in a submission, archive saves the submission data itself as JSON, XML or YAML, and clone does both in one pass. Options come from flags, from a YAML file passed with --opts, or from a config file, with the command line taking priority.

Which Reddit downloader is the best?

The project does not compare itself against other downloaders. It names one alternative, HTTrack, and only for a case it says it does not handle, which is producing a faithful copy of a site rather than collecting raw submission data.

Is there a way to download Reddit data?

The archive command stores submission details, upvotes, text, statistics and all the comments on a submission, serialised as JSON, XML or YAML. Its own examples combine archive with --all-comments and --comment-context.

Does BDFR depend on r/all still existing?

Its examples pass --subreddit all, and one chains 'Python, all, mindustry' into a single argument. Nothing in the repository addresses Reddit's policy for that source, so those examples are syntax to copy rather than advice about the source itself.

How do you update or check a bdfr installation?

pip users rerun python3 -m pip install bdfr --upgrade, pipx users run pipx upgrade bdfr, and bdfr --version prints the installed version. Arch users have two AUR packages, python-bdfr for the latest release and python-bdfr-git for the development build.

Official sources

  1. License: GPL-3.0
  2. Project website
  3. README
  4. Releases
  5. Serene-Arc/bulk-downloader-for-reddit on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/serene-arc-bulk-downloader-for-reddit.svg)](https://hysenlabs.com/projects/serene-arc-bulk-downloader-for-reddit)