# Owez/yark: YouTube channel archiving with a local viewer and change reports

> A Python wrapper around yt-dlp that keeps a channel as a directory of files, tracks deletions, and serves an offline viewer. Worth understanding that the PyPI install line and the README's own note about PyPI do not agree.

**Owez/yark** — OSINT for YouTube made simple.

- Repository: https://github.com/Owez/yark
- Website: https://pypi.org/project/yark/
- Stars: 2,191 · Forks: 71
- Language: Python
- License: MIT
- Published: 2026-10-08 · Updated: 2026-10-08 · Language: en
- Canonical page: https://hysenlabs.com/projects/owez-yark

## New, refresh, view: the whole interface is three commands

The README's walkthrough is three commands, and the design is legible enough to explain in a paragraph. You name an archive and hand it a channel URL:

```shell
$ yark new foobar https://www.youtube.com/channel/UCSMdm6bUYIBN0KfS2CVuEPA
```

Then you pull everything down with a refresh, which is where the downloading happens:

```shell
$ yark refresh foobar
```

And when you want to watch what you have, you serve it locally:

```shell
$ yark view foobar
```

That third command opens an offline website in your browser. The repository tree is small and matches this: a `yark/` package directory, a `pyproject.toml`, a `poetry.lock`, an `examples/` directory holding `madness.py` and an images folder, a LICENSE and a README. The console entry point is declared in `pyproject.toml` as `yark = "yark.cli:_cli"`, so the three verbs are all funnelled through one function. Dependencies are Flask, requests, colorama, yt-dlp and progress, which tells you the shape immediately: yt-dlp does the fetching, Flask serves the viewer, colorama colours the terminal report.

## The archive is a directory you can read without the tool

The archive format is documented in enough detail to be the strongest part of this project. Each archive is a self-contained directory named after whatever you called it at `yark new` time. Inside it, `yark.json` holds all the metadata, `yark.bak` is a backup copy, `videos/` contains one file per known video named by its ID, and `thumbnails/` contains PNG files named by hash.

```
`[name]/` – Your self-contained archive
  - `yark.json` – Archive file with all metadata
  - `yark.bak` – Backup archive file to protect against data damage
  - `videos/` – Directory containing all known videos
    - `[id].*` – Files containing video data for YouTube videos
  - `thumbnails/` – Directory containing all known thumbnails
    - `[hash].png` – Files containing thumbnails with its hash
```

Two things follow from that layout. First, because everything is ordinary files, an archive survives the project being abandoned, which is a property worth more than it sounds for content you cannot re-download later. Second, the backup file and the migrator described in the v1.1 release notes are a deliberate answer to the risk of a migration corrupting `yark.json`.

The README's own advice here is to spend a few minutes reading files inside an archive, and it is not wrong. Someone who knows the format can write their own queries against the metadata without Yark in the loop.

## Archives are additive, so a deleted video stays on your disk

The single most consequential behaviour is stated in the README's Details section as a deliberate rule rather than a bug. If a video is deleted or made private on the YouTube side, Yark keeps the file and tags it `deleted`. The videos remain in the local archive.

That decision has consequences in both directions. For someone keeping evidence of a channel's history, which is the archival use the project name suggests, retaining the file is the entire point and a takedown on the platform does not propagate. For someone expecting their mirror to reflect the live channel, it is the wrong behaviour, and there is no configuration in the README for changing it.

The same section sets two more expectations. Do not create a second archive for the same channel, because Yark accumulates new metadata for a given archive using timestamps, so a duplicate archive means duplicated work and a split history. And scheduling is not implemented: the README explicitly says to use cron or something similar for that.

Deletion tracking itself was one of the headline items in the v1.1 release, which describes it as working through metadata and download trial and error with an issue number attached for future support. That phrasing is a fair signal that detection is heuristic rather than definitive.

## Version 1.2 dated January 2023 against a dependency set that looks newer

The published release history has two entries, both in the first week of January 2023: v1.1 on 2023-01-02 and v1.2 on 2023-01-05. The v1.2 notes describe adding the `report` command, the ability to skip metadata fetching or downloading during a refresh, easier navigation back to channel homepages, and describe the release as stable apart from the absence of formal tests.

Against that, `pyproject.toml` declares `version = "1.2.12"`, which is a patch line past the v1.2 tag rather than at it. It also pins `yt-dlp` to the exact revision `2024.10.07`, not a range, while Flask, requests, colorama and progress all use caret ranges. For a project whose entire job is talking to YouTube, a single pinned downloader version is the dependency to look at first. yt-dlp is the component that breaks when the platform changes, and a hard pin means it does not move when an upstream fix lands.

The last push was on 2026-07-30, well after both releases, so commits are landing even though no new tag has been published. The repository also still has a README commented-out banner about a rewrite on a `v1.3-rewrite` branch with the note that the rewrite is delayed, and image links in the README point at a `1.2-support` branch. The default branch is `main`.

A usable rule for reading this: treat `main` as the current line, treat the v1.2 release notes as the last published changelog, and treat `pyproject.toml` as the authoritative statement of what the installed code actually depends on.

## The install line and the README's PyPI note contradict each other

The README's installation section is a single command with two prerequisites stated in prose: Python 3.9 or newer, and FFmpeg marked optional.

```shell
$ pip3 install yark
```

The same file also carries a commented-out block, currently inactive, that says if you are reading it you are probably trying to install Yark via PyPI, that this route has been removed in newer versions, and that a modern version should be downloaded from the GitHub repository instead. The block is marked for uncommenting when a new interface ships, and it embeds a transition image from the `1.2-support` branch.

Those two statements are about the same install route and they cannot both describe the current state. The visible command is the one the project wants you to run; the hidden note says that path has been closed. The project homepage field also points at the PyPI project page, which is where a listing described as removed would have lived. Rather than pick one, the practical move is to run the install and see what you get: a working package means the note is stale, and a failure or an unexpected version means the note is the accurate one and the source checkout is the route.

The `pyproject.toml` supports a source install cleanly enough that this is not a hard problem. It declares Poetry as the build backend with `poetry-core>=1.0.0` as the requirement, so cloning the repository and building it locally is a normal path rather than an unsupported one.

## Where yark sits against plain yt-dlp

The honest comparison is against yt-dlp on its own, since Yark wraps it and does not replace it. Run yt-dlp directly and you get the same downloads with more control over formats, cookies and post-processing, and no persistent history. Yark adds the thing yt-dlp has no concept of: a time-series record of a channel, with per-video reports, timelines and graphs in the viewer, plus timestamped and permalinked comments.

That history is the actual product. The viewer gives each video a rich history report, and colour-coded report information in the terminal, which is what makes it useful for tracking how a channel behaves over time rather than for downloading a single video. Adding a channel and refreshing it is the whole workflow; there is no config file to write and no database to provision.

What Yark does not add is anything for the hard cases. There is no documented authentication flow for private or members-only content, no incremental partial download, no scheduling, and no resumable large-batch behaviour described anywhere in the README. Twenty-one open issues sit on the repository, which for a project of this size is a reasonable number of open threads but more than a quiet archive would carry.

If you need to prove what a channel published and when, Yark is more useful than yt-dlp. If you need one specific video in the best available quality today, yt-dlp alone is the shorter path and Yark adds nothing to that task.

## Conclusion

Yark earns a place on your disk if you need to keep a specific channel's videos and the history of when each one was published, and you would rather have plain files you can read than a database. The archive format is a directory with a JSON sidecar, so you can inspect it without the tool running. Two things to check before planning around it: confirm that `pip3 install yark` still resolves to the version you expect, because the README carries a note saying the PyPI listing has been removed in newer versions, and look at the yt-dlp pin in `pyproject.toml`, which is fixed at 2024.10.07 rather than tracked by a range. If your requirement is scheduled unattended capture, the project does not do it and its README points you at cron.

## FAQ

### What is Owez/yark used for?

Archiving a YouTube channel locally. You create a named archive for a channel URL, refresh it to download videos and metadata, and then view the result in an offline Flask site. Each archive is a directory holding a yark.json metadata file plus videos and thumbnails as plain files.

### How do I install yark?

The README gives `pip3 install yark` and lists Python 3.9 or newer as the requirement, with FFmpeg optional. The same README carries a commented-out note saying the PyPI route has been removed in newer versions and pointing to the GitHub repository, so the two disagree and the install line is worth verifying before you rely on it.

### Does yark keep videos that get deleted on YouTube?

Yes. The README states archives are always additive: if a video is deleted or marked private on the YouTube side, Yark keeps the local file and tags it `deleted`. There is no configuration described for changing this behaviour.

### Can yark run on a schedule automatically?

Not natively. The README says scheduling is not a feature yet and points you at cron or a similar tool. Everything else, creating the archive, refreshing it and viewing it, is manual or driven by whatever you wire up around it.

## Sources

- [License: MIT](https://github.com/Owez/yark/blob/main/LICENSE)
- [Owez/yark on GitHub](https://github.com/Owez/yark)
- [Project website](https://pypi.org/project/yark/)
- [README](https://github.com/Owez/yark/blob/main/README.md)
- [Releases](https://github.com/Owez/yark/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/owez-yark
