CLI tool
libarchive/libarchive avatar
libarchive/libarchive

libarchive: the streaming C archive library that also ships four command line tools

Multi-format archive and compression library

3,630 stars1,023 forksCNOASSERTION

At a glance

What is it?
Two decades of format handling in portable C, from bsdtar down to WARC and mtree, with a design bet on single-pass streaming rather than random access.
Who is it for?
libarchive is the right answer when the input format is unknown and the archive may be larger than local disk, which is a narrower and more specific claim than support for twenty formats suggests. It is the wrong answer for random access into a ZIP central directory or for extension handling, since both are structural limits of a single-pass design rather than missing features.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly C, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

One library and four command line tools in the same tree

The project is two things at once: a portable C library for reading and writing streaming archives, and a set of command line tools built on top of that library. The README describes the tools as implementations of the common tar, cpio and zcat commands, and the top-level tree backs that up with directories named after each of them, sitting next to `libarchive/` for the core and `libarchive_fe/` for the front-end code.

bash
tar
cpio
cat
unzip

`bsdtar` is a full-featured tar implementation built on the library. `bsdcpio` is described as a different interface to essentially the same functionality, which is the honest way to say that cpio support exists because the same archive machinery already handles it. `bsdcat` is a replacement for zcat, bzcat and xzcat. `bsdunzip` is a replacement for Info-ZIP's unzip.

The examples are worth more attention than they usually get, because this is a callback-driven C API and the shortest way to understand it is to read working code rather than the man pages first. The tree carries `examples/untar.c`, `examples/tarfilter.c` and a self-contained `examples/minitar/`, and the README points readers at exactly these as the places to look for detail `archive.h` comments do not cover. `contrib/` holds items sent in by third parties, which the README is upfront about: you contact the authors, not upstream.

One administrative detail worth settling before you build anything. GitHub reports the license for this repository as NOASSERTION, while the top-level tree does contain a `COPYING` file alongside `SECURITY.md` and `CONTRIBUTING.md`. Read `COPYING` directly rather than inferring terms from the metadata field, and treat the license page on libarchive.org as authoritative for anything the file does not settle.

The format matrix is the actual product, and it is three lists

The Supported Formats section of the README is the most information-dense part of the whole document, but it is easy to misread it because it mixes three separate capabilities. Knowing which list a format appears in tells you what you can actually do with it.

The read list is the widest. It covers Old V7 tar, POSIX ustar, GNU tar including long filenames, long link names and sparse files, Solaris 9 extended tar with ACLs, POSIX pax interchange, four different cpio variants (octet-oriented, SVR4 ASCII, binary in either endianness, and PWB), ISO9660 with Rockridge or Joliet, ZIP including encrypted entries, ZIPX with bzip2, zstd, ppmd8, lzma and xz, GNU and BSD ar, mtree, 7-Zip with zstandard, Microsoft CAB, LHA and LZH, RAR and RAR 5.0, WARC and XAR.

The second list is things detected before the archive proper is parsed. These are wrappers and filters rather than archive formats:

text
  * uuencoded files
  * files with RPM wrapper
  * gzip compression
  * bzip2 compression
  * compress/LZW compression
  * lzma, lzip, and xz compression
  * lz4 compression
  * lzop compression
  * zstandard compression

The create list is shorter and deliberately conservative. It includes ustar, pax interchange, restricted pax that emits ustar except where extensions are genuinely required, old GNU tar, old V7 tar, the cpio family, shar, ZIP, ZIPX, ar, mtree, ISO9660, 7-Zip, WARC and XAR. RAR is absent, and the README is straight about why in the read list: support carries limitations due to RAR's proprietary status. So read RAR, do not plan to write it.

Any output can then be pushed through a filter:

text
  * uuencode
  * base64
  * gzip compression
  * bzip2 compression
  * compress/LZW compression
  * lzma, lzip, and xz compression
  * lz4 compression
  * lzop compression
  * zstandard compression

The practical consequence is that tar, cpio and ISO9660 sit in all three lists, while ZIP sits in two and RAR in one. If your requirement is a symmetric read and write path, the intersection is what you are choosing between.

Streaming is a design bet, and it comes with a stated cost

The README devotes a section to design notes answering what the project says are its most common questions, and the first one sets the whole architecture:

text
* This is a heavily stream-oriented system.  That means that
  it is optimized to read or write the archive in a single
  pass from beginning to end.

The payoff it names is archives larger than available disk, processed on the fly as they arrive from a network or a tape drive. It also names the motivating use case directly: webservers that hand a user an archived view of their own account. That is not a hypothetical workload, and a library that never needs to seek is genuinely simpler to reason about under memory pressure.

The cost is stated just as plainly. In-place modification and random access to the contents of an archive are not directly supported. The README then makes the obvious distinction: for tar.gz this costs nothing, because that format was never designed for random access, while for formats with an index it costs something real. ZIP has a central directory precisely so a reader can jump, and libarchive does not use it.

So the question to ask when evaluating this library is whether your workload is genuinely sequential. If you need to extract one member from a forty gigabyte zip without writing forty gigabytes to disk, libarchive is the wrong tool and a seekable ZIP-specific reader is the right one. If you need to consume an archive from a pipe, a socket or a tape device, or to produce one as bytes leave your application, this is close to ideal, and the `libarchive_fe/` directory exists precisely to support that frontend use.

The C API surface is documented in man pages, not on a website

This is a library whose documentation is a set of manual pages in `doc/`, listed individually in the README. Fourteen of them, covering the whole surface:

text
bsdtar.1
bsdcpio.1
bsdcat.1
libarchive.3
archive_read.3
archive_write.3
archive_write_disk.3
archive_read_disk.3
archive_entry.3
archive_internals.3
libarchive-formats.5
cpio.5
mtree.5
tar.5

The division tells you how the API decomposes. `archive_read.3` and `archive_write.3` are the calling sequences for streaming archive contents. `archive_read_disk.3` and `archive_write_disk.3` are the parallel pair for reading and writing filesystem metadata, which is the mechanism behind ownership, permissions and timestamps surviving a round trip. `archive_entry.3` covers the struct that carries those per-item attributes, and `archive_internals.3` describes internal structure for anyone extending or debugging the library.

The three `.5` pages are the underrated part. `libarchive-formats.5` documents which formats the library supports, `tar.5` and `cpio.5` cover the format families including details the man page tradition tends to skip, and `mtree.5` covers the metadata manifest format. If you have ever wondered why two tar implementations disagree about an edge case in a pax extension, `tar.5` is where the answer lives.

Above all of this sits libarchive.org, which the README calls the home for ongoing development, plus a GitHub wiki. The distinction to keep in mind is that the README itself is an orientation document: it names the manual pages, describes the design intent and lists formats, and it points outward rather than duplicating the API documentation.

Two build systems, one of them generated

The top-level tree carries both a CMake path and an autotools path, which is what a library that has to build on everything from BSD to Linux to Windows tends to accumulate.

text
CMakeLists.txt
configure.ac
Makefile.am
aclocal.m4
Makefile.in
config.h.in

`CMakeLists.txt` is described as input for the cmake build tool, with details in `INSTALL`. The autotools files are the ones the README explicitly marks as maintainer-facing: `configure.ac` feeds autoconf, `Makefile.am` feeds automake, `aclocal.m4` comes from aclocal, and `Makefile.in` and `config.h.in` come from automake and autoheader respectively.

The escape hatch for distribution tarballs is spelled out. If a copy of the source lacks a `configure` script, you can construct one by running the script in `build/autogen.sh`, or use cmake instead. That is a genuinely useful detail because many vendoring pipelines hit exactly this failure when a packaging step strips generated files.

Two other things in the tree say something about how the project is worked on. `CTestConfig.cmake` at the top level registers the suite with CTest, so the tests are drivable from the same tool that drives the build. `test_utils/` exists as its own directory rather than being folded into the library, which suggests the test harness has helpers that are not part of the shipped API. There are no vendored dependency directories in the tree at all, consistent with a library that expects compression backends to come from the system.

What the recent release notes say about where the risk lives

The repository is not archived, GitHub reports the last push as 2026-09-23, and it sits at 3,623 stars with 1,020 forks. Those fork and issue counts are both meaningful signals for a C library: forks in this number usually mean downstream distributions patching it, and the 586 open issues are consistent with a project of this age and scope.

The three most recent releases are all labelled as security or bugfix work, which for a library of this type is the expected shape rather than a worrying one. v3.8.9, published 2026-07-28, adds a Windows port of the unzip tool. v3.8.8, published 2026-06-23, adds reading of encrypted zipx formats and lists a run of core fixes: a heap over-read in pathmatch, an out-of-bounds read in the ACL parser, a double-free in the link resolver, a use-after-free in sparse_reset, a call stack overflow in archive_match, a buffer overrun and wrong output for NULL-name ACL entries, and a set of unchecked memory allocations. v3.8.7, published 2026-04-13, fixes a heap out-of-bounds write in the CAB LZX decoder and a possible heap buffer overflow in iso9660 on 32-bit systems.

Read those lists as a map of the attack surface. The fixes cluster in format parsers (CAB, iso9660, pathmatch, ACLs, sparse handling) rather than in the archive container logic, which is exactly where you would expect exposure in software whose entire job is parsing input it did not create. It is also why the repository ships `SECURITY.md`, routes issue reports through the GitHub tracker, and asks for enhancements as pull requests.

The practical takeaway for anyone embedding this is that the version number is a security variable, not a routine upgrade. Read the ChangeLog between releases rather than waiting for a major version bump.

Editorial conclusion

libarchive is the right answer when the input format is unknown and the archive may be larger than local disk, which is a narrower and more specific claim than support for twenty formats suggests. It is the wrong answer for random access into a ZIP central directory or for extension handling, since both are structural limits of a single-pass design rather than missing features. Start with `examples/untar.c`, then read `libarchive.3` and `archive_read.3` in `doc/` before writing a line against the API, and track the ChangeLog between releases if you vendor the library, because the fix density in the format parsers is where the security exposure lives.

Frequently asked questions

What is libarchive actually used for?

It is used wherever the archive format is unknown or untrusted and the archive may not fit in local storage: ingesting uploads, reading backups, walking WARC crawl data, unpacking ISO images, and generating archives on the fly. The library handles detection itself, so a caller usually does not need to know which format it is holding.

Which archive formats can libarchive create, and which can it only read?

The README lists creation support for ustar, pax interchange, restricted pax, old GNU and V7 tar, the cpio family, shar, ZIP, ZIPX, ar, mtree, ISO9660, 7-Zip, WARC and XAR. Reading adds ISO9660 variants with Rockridge and Joliet, Microsoft CAB, LHA and LZH, RAR and RAR 5.0, and several additional cpio dialects. RAR is read-only because of its proprietary status.

Can libarchive read encrypted ZIP archives?

The read list includes ZIP archives with encrypted entries, and v3.8.8 added reading of encrypted zipx formats using bzip2, lzma, ppmd, xz and zstd compression. Password handling is passed in by the calling application rather than configured inside the library, since the callback-driven API keeps credential policy with the caller.

How do you build libarchive from source?

There are two paths. Use cmake with the top-level CMakeLists.txt, or use the generated configure script. If your copy has no configure script, run the script in build/autogen.sh to construct one first. The autotools inputs listed in the README, including configure.ac and Makefile.am, are maintainer-facing rather than needed by someone just compiling.

Official sources

  1. Issues
  2. libarchive/libarchive on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/libarchive-libarchive.svg)](https://hysenlabs.com/projects/libarchive-libarchive)