Self-hosted service
tubearchivist/tubearchivist avatar
tubearchivist/tubearchivist

Tube Archivist: a self hosted YouTube media server built on Docker, Elasticsearch and yt-dlp

Your self hosted YouTube media server

8,481 stars426 forksPythonGPL-3.0

At a glance

What is it?
Tube Archivist turns a folder of downloaded YouTube videos into a searchable, playable library. It is for people who already keep their own media and want channel subscriptions, metadata indexing and playback in one stack, at the cost of running three containers and about 2GB of memory.
Who is it for?
Adopt Tube Archivist if you already run Docker, have roughly 2GB of memory to spare for a small setup and want channel subscriptions, metadata indexing and playback in one interface rather than a folder of files. Do not adopt it if you only want to download a handful of videos, or if you cannot give Elasticsearch a permanent volume and a password you are willing to manage.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 33 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What problem Tube Archivist solves, and who it is for

A downloaded video collection stops being useful once it is large. Files land in a folder with names you did not choose, and the only way to find a specific talk or episode is to remember roughly when you saved it. Tube Archivist addresses that by indexing the collection with metadata pulled from YouTube, so the library becomes searchable and playable through a web interface. The README frames the problem directly: once a collection grows, it "becomes hard to search and find a specific video."

The intended user runs a home server. The README asks for Docker, around 2GB of available memory for a small testing setup and around 4GB for a mid to large installation, with a dual core CPU at minimum and a quad core preferred. That is a real floor, not a formality: the stack is three services, and one of them is Elasticsearch. Anyone who wants a single binary that downloads a video and stops there is not the audience. The audience is someone who wants subscriptions to channels, a queue, watched and unwatched tracking, and a search box that works across everything they have saved.

How the three containers fit together

The architecture is visible in docker-compose.yml. There are three services: tubearchivist, archivist-redis and archivist-es. The application container listens on port 8000 and mounts two volumes, media at /youtube and cache at /cache. It receives ES_URL pointing at http://archivist-es:9200 and REDIS_CON pointing at redis://archivist-redis:6379, which tells you the dependency direction: the app talks to both, and both are addressed by container name on the compose network.

The download side is yt-dlp, named in the README as the tool used to download videos. The Dockerfile shows what else ships inside the image: ffmpeg and ffprobe are built in a separate stage and copied to /usr/bin, deno is copied from the denoland/deno image, and atomicparsley, nginx, tini and curl are installed as distribution packages. So transcoding, metadata tagging and the web front end are all inside one container rather than delegated to sidecars.

Elasticsearch is the part that shapes the deployment. The compose file runs it as a single node with xpack.security.enabled=true, a 1GB Java heap via ES_JAVA_OPTS=-Xms1g -Xmx1g, and path.repo pointed at /usr/share/elasticsearch/data/snapshot. The heap setting is why the memory floor exists. The snapshot path is why backups are configured at the Elasticsearch layer rather than inside the app.

Installing Tube Archivist with docker compose

The README states the project requires Docker and points to an example docker-compose.yml, with additional user-provided instructions in the docs for Unraid, Synology and Podman. The compose file below is the shape given in the repository, with the passwords and timezone left as the placeholders the file itself uses.

yaml
services:
  tubearchivist:
    container_name: tubearchivist
    restart: unless-stopped
    image: bbilly1/tubearchivist
    ports:
      - 8000:8000
    volumes:
      - media:/youtube
      - cache:/cache
    environment:
      - ES_URL=http://archivist-es:9200
      - REDIS_CON=redis://archivist-redis:6379
      - HOST_UID=1000
      - HOST_GID=1000
      - TA_HOST=http://tubearchivist.local:8000
      - TA_USERNAME=tubearchivist
      - TA_PASSWORD=verysecret
      - ELASTIC_PASSWORD=verysecret
      - TZ=America/New_York

Three of those variables decide whether the first start works. TA_HOST must be the address you actually browse from, with protocol and port, because the application uses it to build its own URLs. ELASTIC_PASSWORD must match the password set on the archivist-es service, or the app cannot reach the index. TZ drives the scheduler, so a wrong value shifts when downloads run.

Bring the stack up from the directory holding the file, then watch the logs of the application container for the first startup, which is when it creates its index and admin user.

bash
docker compose up -d
docker compose logs -f tubearchivist

The compose file also defines a healthcheck that curls http://localhost:8000/api/health/ every 2 minutes with a 30 second start period, so docker compose ps will report the container as starting until that endpoint answers. When it does, log in at the TA_HOST address with TA_USERNAME and TA_PASSWORD, then add a channel and let the download queue run.

Known limitations and port collisions

The README carries a Known limitations section and a separate Port collisions section, which is a fair signal that the project expects both to bite users. Port collisions are the predictable one. The application publishes 8000:8000, Redis exposes 6379, and Elasticsearch runs on 9200 internally. If anything else on the host already holds 8000, the stack will not start cleanly, and the fix is to change the host side of the mapping rather than the container side.

The Elasticsearch image carries a platform constraint that the compose file states outright: bbilly1/tubearchivist-es is described as "only for amd64, or use official es 8.19.0". On arm64 hardware you are expected to substitute the official Elasticsearch image and reconcile the configuration yourself. That is a real gap for Raspberry Pi and Apple Silicon users, and it is not something the README resolves for you.

The third limitation is operational rather than technical. Elasticsearch is not a lightweight index you can delete and rebuild casually if you have years of watch history and playlists attached to it. The compose file sets path.repo for snapshots, which implies the project expects you to use Elasticsearch snapshot and restore rather than copying the app's data directory. The README does not document a rollback path for a failed upgrade, so treat the snapshot configuration as the recovery mechanism you are responsible for.

Tube Archivist compared with Pinchflat and TubeSync

The two names that come up most often next to Tube Archivist are Pinchflat and TubeSync, and the difference is architectural rather than cosmetic. Tube Archivist is a full application: it runs Elasticsearch for indexing, Redis for the queue, and a Django backend with an nginx front end, and it exposes an API and a health endpoint. The indexing is the product. Search across a large library is the reason the Elasticsearch dependency exists.

Pinchflat and TubeSync approach the same job from the download side. They are built around watching channels and fetching new uploads, and they lean on the filesystem and media servers to make the results usable. If your goal is to have files appear in a Jellyfin or Plex library with sensible names, that is a shorter path, because there is no index to keep in sync and no third container to feed.

Tube Archivist takes the opposite bet: it wants to own the library view. That is why the project ships a browser extension for Firefox and Chrome, a Jellyfin plugin and a Plex plugin, rather than treating the media server as the primary interface. If you want the archive itself to be the thing you browse, the extra containers buy you something. If you want the archive to be invisible plumbing behind a media server, they cost you memory and maintenance for a view you will not use.

Maintenance, upgrades and the GPL-3.0 licence

Tube Archivist is licensed GPL-3.0. In practical terms for a self-hoster, that means the software can be run, modified and redistributed under the same licence, and that if you distribute a modified version you must make the source available under GPL-3.0 as well. Running it on your own server for yourself does not trigger distribution obligations. This is a description of the licence, not legal advice, and anyone planning to bundle the project into something they ship should read the LICENSE file in the repository rather than rely on a summary.

The maintenance picture is active. The last push to the develop branch was on 2026-08-28, and the most recent release listed is v0.5.12 from 2026-08-25, following v0.5.11 on 2026-08-15 and v0.5.10 on 2026-03-28. The gap between v0.5.10 and v0.5.11 is roughly four and a half months, so the release cadence is not uniform, but the recent pair of releases in August 2026 shows work landing on the current branch.

Upgrade cost is dominated by two moving parts. The first is Elasticsearch: a major version change there is the kind of upgrade that needs a snapshot taken first, which is what ES_SNAPSHOT_DIR and the path.repo setting are for. The second is yt-dlp, which breaks whenever YouTube changes something. The README exposes TA_AUTO_UPDATE_YTDLP to install the latest yt-dlp on container start, which trades a manual step for the risk of picking up a regression automatically. There is also a TA_LOGIN_AUTH_MODE variable that selects between single, local, ldap, forwardauth and ldap_local, so authentication changes are a configuration decision rather than a rebuild.

Where the documentation is thin

The README covers environment variables well and links to the docs site for full explanations, but several things a new operator needs are not resolved in the repository itself. There is no documented rollback procedure for a failed upgrade, which matters because Elasticsearch index migrations are the least reversible part of the stack. The README also does not state what happens to in-flight downloads when the application container restarts, despite restart: unless-stopped being set on all three services.

The arm64 path is the other soft spot. Substituting the official Elasticsearch 8.19.0 image is mentioned in a compose comment, but the compose file as written still points at bbilly1/tubearchivist-es, so anyone on non-amd64 hardware is adapting a configuration the project does not ship. The docs site is the place to check before assuming a copy-paste deployment will work on your hardware.

Editorial conclusion

Adopt Tube Archivist if you already run Docker, have roughly 2GB of memory to spare for a small setup and want channel subscriptions, metadata indexing and playback in one interface rather than a folder of files. Do not adopt it if you only want to download a handful of videos, or if you cannot give Elasticsearch a permanent volume and a password you are willing to manage. Before committing, verify the port collision situation on your host, confirm that TA_HOST matches the address you will actually browse from, and check whether your platform can use the bbilly1/tubearchivist-es image or needs the official Elasticsearch 8.19.0 build instead.

Frequently asked questions

How do I install Tube Archivist?

The project requires Docker. The README points to an example docker-compose.yml with three services (tubearchivist, archivist-redis and archivist-es), and you set TA_HOST, TA_USERNAME, TA_PASSWORD, ELASTIC_PASSWORD, REDIS_CON and TZ before starting it. The docs also carry user-provided instructions for Unraid, Synology and Podman.

How do I use Tube Archivist?

Once the stack is running you log in with TA_USERNAME and TA_PASSWORD at the TA_HOST address, subscribe to YouTube channels, and let yt-dlp download the videos. The application indexes them with metadata from YouTube so you can search, play and track watched and unviewed videos offline through the web interface.

How do I update Tube Archivist?

The README does not give an upgrade procedure. It does expose TA_AUTO_UPDATE_YTDLP, which installs the latest yt-dlp when the container starts, and the compose file sets path.repo for Elasticsearch snapshots, which is the mechanism available for backing up the index before a change.

Is Tube Archivist safe?

The repository does not make a security claim, so the honest answer is that safety depends on your deployment. The compose file runs Elasticsearch with xpack.security.enabled=true and requires an ELASTIC_PASSWORD, and the app supports reverse proxy authentication through TA_ENABLE_AUTH_PROXY and LDAP through TA_LDAP.

What is Tube Archivist?

It is a self hosted YouTube media server. It subscribes to channels, downloads videos with yt-dlp, indexes them with metadata from YouTube, and lets you search, play and track watched and unviewed videos offline through a web interface.

Is there a YouTube channel archiver available?

Tube Archivist is one. The README lists subscribing to YouTube channels as core functionality, with yt-dlp handling the downloads and Elasticsearch making the resulting collection searchable.

Official sources

  1. License: GPL-3.0
  2. Project website
  3. README
  4. Releases
  5. tubearchivist/tubearchivist on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tubearchivist-tubearchivist.svg)](https://hysenlabs.com/projects/tubearchivist-tubearchivist)