Self-hosted service
ArchiveBox/ArchiveBox avatar
ArchiveBox/ArchiveBox

ArchiveBox: Self-Hosted Web Archiving in HTML, PDF, WARC, and Media Formats

🗃 Open source self-hosted web archiving. Takes URLs/browser history/bookmarks/Pocket/Pinboard/etc., saves HTML, JS, PDFs, media, and more...

28,643 stars1,611 forksPythonMIT

At a glance

What is it?
ArchiveBox is an MIT-licensed Python application that preserves web content in multiple redundant formats including HTML, PDF, WARC, screenshots, and downloaded media. It accepts URLs, browser history, bookmarks, RSS feeds, and social media input, and stores everything in standard files and folders without proprietary formats.
Who is it for?
ArchiveBox is worth deploying for anyone who needs to preserve web content under their own control: researchers, legal teams backing up evidence, or individuals archiving bookmarks before link rot removes them. It is not the right tool for someone who only needs single-page offline copies without long-term storage, nor for users who want a cloud service that handles infrastructure.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What ArchiveBox Preserves and Who Uses It

Content on the web disappears: pages go offline, URLs change, and services shut down without warning. ArchiveBox addresses this by creating local copies in formats that remain readable without running the archiving tool again.

The README describes the use cases as saving copies of bookmarks, preserving evidence for legal cases, backing up photos from Facebook, Instagram, and Flickr, downloading media from YouTube and SoundCloud, and saving research papers. The tool accepts both public and private web content.

The output is stored as ordinary files and folders in a standard layout, with no proprietary binary format. The README specifies the output formats as HTML, PNG, PDF, TXT, JSON, WARC, and SQLite, all described as guaranteed to be readable for decades. The SQLite database holds the archive index; the archive content itself lives in subdirectories named by URL hash.

ArchiveBox provides four ways to interact with an archive: a CLI, a self-hosted web interface, a Python API, and direct filesystem access. A REST API and webhooks are also documented for service integrations.

Input Formats: Bookmarks, RSS, Pocket, and a Browser Extension

ArchiveBox accepts URLs one at a time or in bulk from several sources. The README lists browser history exports, bookmark files, social media feeds, RSS feeds, link-saving services like Pocket and Pinboard, and a browser extension as input sources. Regular imports can be scheduled from any of these sources.

For each URL added, ArchiveBox detects content type and applies appropriate extractors: - For general websites: the archive includes original HTML, CSS, and JavaScript, a single-file HTML snapshot, a screenshot PNG, a PDF, a WARC file, the page title, the article text, the favicon, and HTTP headers. - For social media and news: post content, comment text, title, and author. - For YouTube, SoundCloud, and similar media sites: MP3 or MP4 files, subtitles, metadata, and thumbnail. - For GitHub and GitLab repository links: a clone of the source code, the README, and images.

The breadth of output types is what distinguishes ArchiveBox from single-format tools. Each URL produces a folder with multiple redundant copies, so even if one tool (such as wget) misses something, the Chrome-based screenshot and the yt-dlp download capture different aspects of the same page.

Installing ArchiveBox with Docker Compose or pip

The README recommends Docker Compose as the primary installation method:

bash
mkdir -p ~/archivebox/data && cd ~/archivebox
curl -fsSL 'https://docker-compose.archivebox.io' > docker-compose.yml
docker compose pull
docker compose up -d --wait

After that, open http://admin.archivebox.localhost:5797 in a browser to complete setup. The Docker Compose setup initializes a new collection automatically.

For a plain Docker container:

bash
mkdir -p ~/archivebox/data && cd ~/archivebox/data
docker run -d --name archivebox -v "$PWD:/data" -p 5797:5797 archivebox/archivebox:dev

For a pip-based install without Docker:

bash
pip install archivebox>=0.9.0rc0

Brew and apt packages are also mentioned as options in the README's Quickstart section. After any installation, the add command queues URLs for archiving:

bash
docker compose run --rm archivebox add 'https://example.com'

The archive data directory mounted as /data in Docker is the same collection accessed by the CLI, Python API, and web interface, so all four interaction modes work on the same data without duplication.

The Standard Tools Behind Each Archive Format

ArchiveBox does not implement its own content fetching. The README states it uses standard tools: Chrome (for screenshots and JavaScript-rendered pages), wget (for raw HTML and asset downloads), and yt-dlp (for media from video and audio platforms).

This means the archive quality depends on which of those tools are available in the deployment environment. A Chrome instance is needed for screenshot and PDF output. Without yt-dlp, video and audio downloads will not be captured. The Dockerfile comment in the README notes that if the inherited image lacks a downloader dependency, the correct fix is to update the abx-dl package rather than patching the deployed container.

The use of standard tools also means archive output is readable without ArchiveBox. A WARC file can be opened in tools like Replay Web.page or WAIL. An MP4 file is an MP4 file. None of the output requires running ArchiveBox to access after the capture.

Limitations and What ArchiveBox Cannot Reliably Capture

Pages that require authentication cannot be archived without providing session cookies or credentials, and the README does not document an automated credential management flow for that case.

Dynamic single-page applications that load all content through JavaScript may not archive completely via wget. The Chrome-based capture is more reliable for those pages, but capturing dynamic state (logged-in views, paginated content) still depends on the page's behaviour at the time of the request.

Social media content where the platform blocks automated access (CAPTCHA walls, rate limits, or API requirements) will fail silently or produce partial captures. The README acknowledges this implicitly by listing post content and comments as outputs rather than guaranteeing them.

ArchiveBox stores each archive snapshot with the content as it was at capture time. It does not monitor pages for changes and re-archive automatically; re-adding a URL creates a new snapshot rather than updating the old one.

ArchiveBox vs a Read-Later or Bookmark-Only Tool

Wallabag and Linkwarden are alternatives that sit in the bookmark-to-archive category. Wallabag saves a cleaned article view for reading later; Linkwarden stores links with metadata. Neither produces WARC files, media downloads, or multi-format snapshots by default.

The meaningful difference is scope. ArchiveBox is an archiving tool: its output is designed for long-term preservation in multiple redundant formats that remain readable without it. Wallabag and Linkwarden are reading and organization tools where the saved copy is an interface feature, not the primary deliverable.

For someone who wants offline reading with a clean interface, Wallabag is a lighter deployment. For someone who needs court-admissible capture with full WARC output, ArchiveBox is the only self-hosted option in this category. SingleFile, mentioned in the search questions, creates single-file HTML snapshots but does not produce WARC, media downloads, or git clones.

Maintenance Record and Licence

The last push to the repository was on 2026-09-28, with three releases in the preceding five days: v0.9.70 on 2026-09-27, v0.9.64 on 2026-09-26, and v0.9.50 on 2026-09-23. This release cadence indicates active development on the 0.9.x series.

ArchiveBox is licensed under MIT. The MIT licence permits commercial use, modification, and redistribution with no copyleft obligations. Organisations using ArchiveBox in commercial workflows or embedding it in products face no licence restrictions beyond retaining the copyright notice.

The default branch is dev, and the Docker image tag in the README's examples is archivebox/archivebox:dev, which pulls the development channel. Teams requiring stability should pin to a specific release tag rather than using the dev tag in production.

Editorial conclusion

ArchiveBox is worth deploying for anyone who needs to preserve web content under their own control: researchers, legal teams backing up evidence, or individuals archiving bookmarks before link rot removes them. It is not the right tool for someone who only needs single-page offline copies without long-term storage, nor for users who want a cloud service that handles infrastructure. Before deploying, verify that Chrome, wget, and yt-dlp are available or that the Docker image includes them, since those tools drive the actual content capture.

Frequently asked questions

What is ArchiveBox?

ArchiveBox is an open-source self-hosted web archiving application that saves copies of URLs in multiple formats: HTML, PDF, PNG screenshots, WARC files, and media downloads. It is written in Python, licensed under MIT, and stores everything in ordinary files and folders without proprietary formats.

How do I install ArchiveBox?

The recommended method is Docker Compose: create a directory, download the compose file with curl -fsSL 'https://docker-compose.archivebox.io' > docker-compose.yml, then run docker compose pull and docker compose up -d --wait. A pip install (pip install archivebox>=0.9.0rc0) and a plain Docker container are also documented options.

Is ArchiveBox safe to use?

ArchiveBox is MIT-licensed open-source software whose code is publicly auditable. It stores all archived data on your own server in standard file formats. The README notes it uses Chrome, wget, and yt-dlp under the hood, so the same security considerations that apply to running a browser on your server apply here.

How does ArchiveBox compare to the Wayback Machine?

The Wayback Machine is a public service operated by the Internet Archive that archives public web pages. ArchiveBox is self-hosted: you run it on your own server, archive the URLs you choose, and store the data under your control. ArchiveBox can archive private or authenticated content that the Wayback Machine cannot access.

How does ArchiveBox differ from SingleFile?

SingleFile creates a single HTML file containing an entire page's assets inline. ArchiveBox produces multiple archive formats per URL including WARC, PDF, screenshot, media downloads, and a git clone for repository links. SingleFile is a browser extension focused on offline reading; ArchiveBox is a server application focused on long-term preservation in multiple formats.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/archivebox-archivebox.svg)](https://hysenlabs.com/projects/archivebox-archivebox)
Community notes

Community notes