# Duplicacy: lock-free deduplication for backups from several machines to one storage

> Duplicacy is a Go command line backup tool that stores every chunk as its own file named after its hash, so several computers can write to one cloud bucket without coordinating. Here is how that design works, how to run a first backup, and where it stops being the right choice.

**gilbertchen/duplicacy** — A new generation cloud backup tool 

- Repository: https://github.com/gilbertchen/duplicacy
- Website: https://duplicacy.com
- Stars: 5,693 · Forks: 357
- Language: Go
- License: NOASSERTION
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/gilbertchen-duplicacy

## The problem Duplicacy solves: many machines, one storage, no chunk database

Most backup tools that deduplicate assume one writer per repository. The README states that Duplicacy is designed so that multiple computers can back up to the same cloud storage, taking advantage of cross-computer deduplication, without direct communication among them. That is the whole pitch, and it is the reason the project exists.

The second half of the design is what makes the first half possible. Chunk-based tools typically group chunks into pack files and keep a chunk database that records which chunk lives in which pack. Duplicacy takes a database-less approach: every chunk is saved independently, using its hash as the file name, so a lookup is a name lookup rather than a query against a central index. The README calls this simpler and less error-prone, and it also says the absence of a central database is what made features such as asymmetric encryption and erasure coding easier to add.

Who this is for: an engineer or a small team with several laptops or servers that share an operating system or a codebase, plus one storage account. If you back up one machine to one disk, the cross-computer argument does not apply to you, and the trade-offs below matter more than the benefit.

## How the chunk model works in practice

Files are split into chunks, each chunk is hashed, and the hash becomes the chunk's file name in the storage. A snapshot is then a list of chunk references, which is why the README describes every incremental backup as a full snapshot in the chunk-based model it credits to Duplicati. Two machines that hold the same bytes produce the same hashes and therefore write the same file names, so the second machine's upload collapses into a check for a name that already exists. No lock is taken, and no machine needs to know the others are running.

The repository layout backs this up. DESIGN.md sits at the top level next to GUIDE.md, and the README links a paper describing the inner workings, accepted by IEEE Transactions on Cloud Computing, with a draft copy stored in the repository as duplicacy_paper.pdf. The Go module in go.mod pulls in github.com/gilbertchen/highwayhash, github.com/minio/highwayhash and github.com/minio/blake2b-simd for hashing, github.com/bkaradzic/go-lz4 and github.com/klauspost/compress for compression, and github.com/klauspost/reedsolomon for the erasure coding feature. Storage backends are ordinary SDK dependencies: aws-sdk-go, the Azure SDK fork, go-dropbox, go-smb2, ncw/swift, pkg/sftp and storj.io/uplink all appear in the same require block.

One consequence is worth stating plainly. Because the storage sees plain file names, any backend that supports a basic set of file operations can host a repository. The README lists local disk, SFTP, Dropbox, Amazon S3, Wasabi, DigitalOcean Spaces, Google Cloud Storage, Microsoft Azure, Backblaze B2, Google Drive, Microsoft OneDrive, Hubic, OpenStack Swift, WebDAV (described as under beta testing), pcloud, Box.com and File Fabric by Storage Made Easy.

## Installing Duplicacy and running a first backup

The README does not print install commands. It points to three wiki pages: a brief introduction under Quick-Start, the command references, and a page titled Installation for building from source. It also says this repository hosts source code, design documents and binary releases of the command line version, and that a Web GUI frontend for Windows, macOS and Linux is available separately from duplicacy.com. So the practical routes are the binary releases attached to this repository, the wiki installation page, or a build from source with the Go toolchain, which go.mod pins at go 1.19.

The README and the repository files give no command line examples, so there is no snippet to copy here. What they do give is the set of storage backends above and a pointer to the wiki command references, which is where the exact syntax for initializing a repository, running a backup and listing snapshots lives. Read that page before your first run; the storage URL format and the available flags are documented there and nowhere in this repository.

Two things are worth deciding before the first run. First, which storage from the README list you will use, since that choice fixes the URL scheme in the initialization command and the credentials the tool needs. Second, whether the machine is one of several that will share the repository, because that is the case the design targets and the case where a second machine's run should upload far fewer chunks than the first. The README does not document a daemon mode, so if you want scheduled backups, the scheduling is the operating system's job, not Duplicacy's.

## Where the lock-free design costs you

A database-less repository trades query power for simplicity. With one file per chunk, listing the chunks of a snapshot is cheap, but there is no index to ask questions of. If you want to know which snapshots reference a given chunk before deleting anything, the tool has to walk snapshots rather than read a table. The README does not describe a garbage collection command or its safety rules, so treat deletion of old backups as something to read up on in the wiki rather than something to guess at.

The comparison section is candid about the other side of this. Duplicacy's own description of duplicity notes that a chain of dependent backups means deleting one backup renders the subsequent ones on that chain useless, and that periodic full backups are required to make earlier ones disposable. Duplicacy's chunk model avoids that chain, which is a real advantage, but it also means the storage holds many small objects. On backends that bill per request or throttle small-object listings, a repository of individually named chunks behaves differently from one built of large pack files, and the README's cloud comparison chart is about running time, not about request costs.

The wrong-tool case is a database server. Nothing in the README promises application-consistent snapshots, and a chunked file-level backup of a running database captures whatever the files looked like while they were being written. Use the database's own dump for that, and let Duplicacy carry the resulting files.

## Duplicacy compared with restic and Kopia

The README compares Duplicacy with duplicity, bup, Duplicati and Attic, not with restic or Kopia, so the difference in approach has to be read off the design rather than off a claim in the repository. Restic and Kopia are also chunk-based deduplicating backup tools, and both are widely used for the same job. The distinction that the README draws for Duplicacy is architectural: chunks are stored as individual hash-named files with no centralized chunk database, and multiple computers can write to one storage without communicating.

If your requirement is several machines sharing one repository, that is the axis to compare on. Ask of any candidate whether it expects a single writer, and whether it keeps an index that must be rebuilt or locked when a second client appears. The README's criticism of Attic is exactly this point: concurrent backups from multiple clients are theoretically possible with locking, but the developer did not recommend them because chunk indices live in a local cache. That is the failure mode Duplicacy was built to avoid.

If instead your requirement is a mature ecosystem of wrapper scripts, a single-writer repository, or a tool whose repository format you already operate, the lock-free argument buys you nothing and you should weigh the smaller communities and thinner documentation around the alternatives accordingly. The README does not make that comparison for you.

## Maintenance, releases and licensing

The repository is not archived, and the last push was on 2026-08-06, so work is happening. Releases are not on a fast cadence: v3.2.3 in October 2023, v3.2.4 in November 2024, and v3.2.5 in May 2025. Between those, users depend on the master branch. The Go module declares go 1.19 and pins a large set of vendored SDK dependencies, several of them forks under github.com/gilbertchen, which means storage-provider API changes arrive through the project rather than through your own dependency updates.

Upgrade cost is mostly the storage SDK surface. Adding a backend is a code change in this repository, not a plugin you install, so a provider that is missing from the list in the README stays missing until it is added upstream. The README does not document a rollback procedure for a repository written by a newer version, and the wiki is the place to look before upgrading a repository you cannot re-create from the source data.

On licensing, the repository metadata reports NOASSERTION, which means no licence identifier was detected, and the authoritative file is LICENSE.md at the top level. Read it yourself for the terms that apply to your use, including whether the command line version and the separately distributed GUI have different terms. The README also mentions a VMware vSphere edition called Vertical Backup distributed from verticalbackup.com, which is a different product from this repository.

## Conclusion

Adopt Duplicacy if several machines with overlapping data should share one cloud storage and you want a database-less repository you can inspect file by file. Do not adopt it if you need point-in-time restore of an application database, because a snapshot is a set of chunks, not a consistent volume image, and do not adopt it if you need a rollback path documented in the README, because the README covers storages and comparisons and leaves upgrade and rollback to the wiki. Before committing, check the storage backends listed in the README against the one you actually use, read the wiki command references for the exact syntax, and check the LICENSE.md file in the repository, since the repository metadata reports NOASSERTION rather than a named licence.

## FAQ

### Is it duplicacy or duplicity?

They are two different backup tools. Duplicacy is this project, a Go command line tool whose README describes lock-free deduplication and a chunk-per-file repository; duplicity is a separate tool the README describes as using the rsync algorithm to upload only differences from previous backups.

### What does duplicacy mean as a word?

The README does not define the word. The project's own name is used for the backup tool, and the README does not discuss the dictionary sense of duplicacy or its spelling.

### How does duplicacy work?

Files are split into chunks, each chunk is saved independently with its hash as the file name, and a snapshot records which chunks it needs. The README states that avoiding a centralized chunk database keeps the implementation simpler and lets multiple computers back up to the same storage without communicating.

### What is duplicacy?

Duplicacy is a cross-platform cloud backup tool written in Go, distributed here as a command line version with source, design documents and binary releases. The README lists local disk, SFTP, Dropbox, Amazon S3, Backblaze B2, Google Drive, OneDrive, Swift, WebDAV and several other storage backends.

### How is duplicacy different from duplicati?

The README says Duplicati splits files into fixed-size chunks, so insertions or deletions of a few bytes defeat deduplication, and that multiple clients cannot back up to the same storage location. Duplicacy's stated design allows multiple computers to share one storage with cross-computer deduplication.

## Sources

- [gilbertchen/duplicacy on GitHub](https://github.com/gilbertchen/duplicacy)
- [Issues](https://github.com/gilbertchen/duplicacy/issues)
- [Project website](https://duplicacy.com)
- [README](https://github.com/gilbertchen/duplicacy/blob/master/README.md)
- [Releases](https://github.com/gilbertchen/duplicacy/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/gilbertchen-duplicacy
