bup: git-packfile backups with rolling checksums and global deduplication
Very efficient backup system based on the git packfile format, providing fast incremental saves and global deduplication (among and within files, including virtual machine images). Please post problems or patches to the mailing list for discussion (see the end of the README below).
At a glance
- What is it?
- bup splits files with an rsync-like rolling checksum and stores them in git packfiles, so huge VM images and databases can be saved incrementally. It is a good fit for people who want deduplicated backups over ssh and can live with a project that is not as well tested as tar.
- Who is it for?
- bup suits engineers who need incremental backups of very large single files, such as VM disk images or database dumps, and who want the stored data readable as git packfiles. It is the wrong choice if you need a tool as thoroughly tested as tar, or if you are not prepared to run bup fsck and keep par2 available for recovery.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 26 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What bup solves, and who it is for
Backing up a virtual machine image, a database dump or a large XML file with a traditional archiver means every run writes the whole file again. bup exists to avoid that. The README states that bup uses a rolling checksum algorithm similar to rsync to split large files into chunks, so you can back up huge VM disk images, databases and XML files incrementally without spending disk space on multiple full versions. The second half of the problem is deduplication across machines: the README says data is shared between incremental backups automatically, without having to know which backup is based on which other one, even when the backups come from two computers that do not know about each other. The intended audience is therefore someone with large, mostly unchanging binary files and more than one host, who wants a repository they can inspect with ordinary git tooling. If your data is small text files that change completely between runs, the chunking machinery buys you little.
How the rolling checksum and git packfiles fit together
Two mechanisms do the work. The first is content-defined chunking: a rolling checksum scans the file and cuts it into chunks at boundaries determined by the data rather than by fixed offsets, which is what makes an edited VM image produce mostly identical chunks on the next run. The second is storage format. bup writes git packfiles directly, and the README contrasts this with git itself, which has a separate garbage collection and repacking stage. Writing packfiles directly is what the README credits for speed on very large data sets. The same design allows bup to track far more filenames than git (millions) and far more objects (hundreds or thousands of gigabytes), because bup's index formats were built for that scale. An incremental backup is presented as if it were a full backup, so restoring does not mean replaying a chain of increments. The repository can also be mounted as a FUSE filesystem and exported over Samba, which gives you a second way to reach the content when the bup command line is not convenient.
Installing bup and making a first backup
The README does not contain a package installation command; it points readers at the project homepage at bup.github.io and at the mailing list for problems and patches, and it lists the runtime requirements: python 3.8 or newer (though it is automatically tested only against some python versions 3.9 and newer), a C compiler, and an installed git version >= 1.7.2 (automatically tested against git 2.3 and newer). par2 is required if you want fsck to generate the information needed to recover from some types of corruption. The repository ships a configure script and a GNUmakefile at the top level, which is the build path a source checkout follows. The README does not spell out the first bup commands, so check the documentation in the Documentation/ directory of the repository before running anything against real data. What the README does describe is the shape of a backup: a rolling checksum splits the input into chunks, the chunks are written into git packfiles in a bup repository, and each run stores only the data that is not already present. A remote target is set up by installing bup on any machine you can reach over ssh; the README states you can back up directly to a remote bup server without needing large temporary space on the client, and that an interrupted backup resumes on the next run.
The data-loss warning in bup's own README
The README's own section on reasons to avoid bup opens with the admission that it is not remotely as well tested as something like tar, and is therefore more likely to eat your data. That sentence should be read literally rather than as modesty. A second constraint is the test matrix: the project intends to work with python 3.8 or newer and git >= 1.7.2, but the README says it is currently only automatically tested against some python versions 3.9 and newer and git versions 2.3 and newer, and asks for discrepancies to be reported. Platform support is similarly reported rather than guaranteed: Linux, FreeBSD, NetBSD, OS X >= 10.4, Solaris, or Windows with Cygwin or WSL, with patches for other platforms welcome. Finally, corruption recovery depends on par2. If you want fsck to be able to generate the information needed to recover from some types of corruption, par2 has to be installed; without it that recovery path is unavailable. None of this makes bup unusable, but it does mean the tool expects you to verify your restores rather than assume them.
bup compared with rsync-based snapshot tools
The obvious alternative for the same job is a snapshot tool built on rsync and hard links, where each backup directory looks like a full tree and unchanged files are hard-linked to the previous run. The difference is where the deduplication happens. Hard-link snapshots deduplicate whole files, so a VM image that changes by a few blocks is copied in full every time. bup's rolling checksum cuts inside the file, so the unchanged chunks are reused and only the changed ones are stored, which is exactly the case the README calls out for VM disk images, databases and XML files. The trade-off runs the other way for ordinary file trees: a hard-link snapshot is a plain directory you can browse with ls and copy with cp, while a bup repository is a packfile store that you read through bup or by mounting it as a FUSE filesystem. rsync-style tools also have no equivalent of bup's automatic sharing between hosts that do not know about each other. Choose bup when the large-file case dominates; choose the hard-link approach when you want the backup to be a directory you can inspect without any tooling.
Maintenance, releases and licensing
The repository is not archived, and the last push was on 2026-09-04. The README lists release notes from 0.27.1 up to 0.34, each as a separate note file, so upgrades are documented per version rather than in one running changelog. The repository's LICENSE file is present at the top level, but the project metadata reports the licence as NOASSERTION, which means the licence identifier was not recognised automatically; read LICENSE yourself before redistributing bup or bundling it into a product. There is nothing in the README about a paid tier, a hosted service or a support contract, so the maintenance model is the mailing list named at the end of the README. Upgrading means rebuilding from source, since the build path is configure plus GNUmakefile, and checking that your python and git versions still fall inside the tested range. The realistic ongoing cost is not the upgrade itself but the verification habit: you need par2 installed if you want fsck to be able to generate recovery information, and you need to run restores to know the repository is intact.
Editorial conclusion
bup suits engineers who need incremental backups of very large single files, such as VM disk images or database dumps, and who want the stored data readable as git packfiles. It is the wrong choice if you need a tool as thoroughly tested as tar, or if you are not prepared to run bup fsck and keep par2 available for recovery. Before adopting it, verify the versions of python, git and the C compiler on your target machines, confirm that your platform is one of the reported ones (Linux, FreeBSD, NetBSD, OS X >= 10.4, Solaris, or Windows with Cygwin or WSL), and check whether your repository layout allows a remote bup server over ssh.
Frequently asked questions
What is bup?
bup is a backup program whose name is short for backup. It splits large files with a rolling checksum similar to rsync and stores the result in git's packfile format, which gives incremental saves and deduplication across and within files.
What platforms does bup support?
The README says bup has only been reported to work on Linux, FreeBSD, NetBSD, OS X >= 10.4, Solaris, or Windows with Cygwin or WSL, and that patches for other platforms are welcome.
What do I need to build and run bup?
The README lists python 3.8 or newer, a C compiler, and an installed git version >= 1.7.2, while noting automatic testing only covers some python versions 3.9 and newer and git 2.3 and newer. par2 is also required if you want fsck to generate recovery information for some types of corruption.
Can bup back up directly to a remote server?
Yes. The README states you can back up directly to a remote bup server without needing large temporary disk space on the machine being backed up, that an interrupted backup resumes where it left off, and that setting up a server just means installing bup on any machine where you have ssh access.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/bup-bup)