# Karakeep: a self-hosted bookmark-everything app with AI tagging

> Karakeep stores links, notes, images and PDFs in one self-hosted pile, then crawls, archives and tags them automatically. It is built for homelab owners who want full-text search over their own hoard, and it costs you a stack of containers to run.

**karakeep-app/karakeep** — A self-hostable bookmark-everything app (links, notes and images) with AI-based automatic tagging and full text search

- Repository: https://github.com/karakeep-app/karakeep
- Website: https://karakeep.app
- Stars: 29,276 · Forks: 1,533
- Language: TypeScript
- License: AGPL-3.0
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/karakeep-app-karakeep

## The problem Karakeep solves for self-hosters

Most people who save links end up with three separate piles: a read-it-later queue, a notes app, and a folder of screenshots. The README describes the author's own path through this, starting with Pocket, then moving to memos for quick notes, and finding that memos did not preview or archive the links dumped into it. Karakeep is the answer to that gap. It treats a bookmark as an object with a title, a description, a preview image, an archived copy of the page, extracted text, and tags, rather than a URL in a list.

The target user is specific. Someone running a home server who wants their saved material on hardware they control, and who is willing to accept operational work in exchange. The README is explicit that this is a self-hosting first app, and the managed cloud at cloud.karakeep.app exists mainly for people who are not comfortable with that trade. If you just want a bookmark bar that syncs, this is far more machinery than the job requires.

## How crawling, archiving and tagging actually flow

The stack section of the README names the moving parts. NextJS runs the web app on the app router. Drizzle handles the database and its migrations. NextAuth does authentication, tRPC carries client to server calls, and Puppeteer crawls bookmarks. Meilisearch backs the full content search, and OpenAI is listed for the LLM work, with the README also stating support for local models through ollama.

So the path of a saved URL is roughly: you submit it, a worker picks it up, Puppeteer loads the page to pull the title, description and preview image, the page is archived with monolith for full page archival, and the text lands in Meilisearch for full text and semantic search. Tagging and summarization run through the LLM, which is why the README can offer both a hosted model and a local ollama instance. Video archiving is handled by yt-dlp.

That is a lot of processes for a bookmark app, and the repository layout reflects it: apps/ and packages/ sit under a pnpm and Turbo workspace, with separate scripts for the web app and the workers. The root package.json exposes db:migrate, workers and web as distinct commands, so the web tier and the background tier are meant to be run and restarted separately. If the worker is down, bookmarks still save but previews, archives and tags do not appear.

## Installing Karakeep with Docker and saving a first link

The README points installation at the Docker page of the docs, and the repository carries docker/, charts/, kubernetes/ and a karakeep-linux.sh script, so container deployment is the intended route. The repository's own .env.sample is minimal and names only two variables, which tells you the sample is a starting point rather than a complete configuration.

If you would rather run it from source, the workspace is a pnpm and Turbo monorepo, and the root package.json defines the commands the project itself uses. Install dependencies and start the web app with:

```bash
pnpm install
pnpm web
```

Before the web app is useful you need the database schema applied, which the root package.json exposes as a separate script:

```bash
pnpm db:migrate
```

The background work does not run inside the web process. Preview fetching, archiving and tagging come from the workers package, started with its own command:

```bash
pnpm workers
```

Configuration is read from environment variables. The repository's .env.sample shows the shape but only names two keys:

```bash
DATA_DIR=<path>
NEXTAUTH_SECRET=<secret>
```

Set DATA_DIR to the directory that will hold your stored files and NEXTAUTH_SECRET to a generated secret, then consult docs.karakeep.app/configuration for the rest. Once the web app is up, the README's demo at https://try.karakeep.app gives you a read-only look at a seeded instance, which is the fastest way to judge the tagging and search before you commit disk space to it.

## Where Karakeep gets expensive to run

The honest cost here is not the licence, it is the footprint. Puppeteer means a headless browser is part of your deployment, and headless browsers are memory hungry and need to be kept current with the sites they render. Meilisearch is a separate search service with its own index storage. The database and the archive of full pages both grow with every bookmark, and full page archival with monolith means you are storing pages, not links.

The AI layer is the second cost. The README lists OpenAI in the stack and separately notes support for local models using ollama. Those are two very different bills: an API key that scales with every tag and summary, or a local model that needs GPU or at least serious CPU and RAM on the same box. Neither is wrong, but the choice is not cosmetic, and the README does not present it as a decision you can defer.

A third limit is scope. Karakeep is a hoarding tool, not a reading tool. There is no indication in the README of a reader mode or an annotation workflow beyond marking and storing highlights from hoarded content. If your actual habit is to read long articles end to end and keep notes inside them, Wallabag, which the README describes as a well-established open source read-it-later app written in php, is closer to that shape.

## Karakeep vs Raindrop, Linkwarden and memos

The README names its own alternatives, which makes comparison straightforward. Raindrop is described as a polished open source bookmark manager supporting links, images and files, with the key difference that it is not self-hostable. If you want zero servers, Raindrop wins on that single axis and Karakeep cannot answer it. Linkwarden is described as an open source self-hostable bookmark manager focused mostly on links, with collaborative collections. The difference is emphasis: Linkwarden stays close to links, while Karakeep pulls notes, images, PDFs, OCR and LLM tagging into the same object model.

memos gets the most affectionate treatment in the README, and the comparison is the sharpest. The author ran memos, liked it, and found it did not archive or preview shared links and did not tag automatically. That is the exact gap Karakeep was built to fill. If you already run memos and are happy with it as a notes tool, Karakeep is not a replacement so much as a second service for the link side, which is more infrastructure, not less.

The README also lists mymind as the closest alternative overall and the source of much of the inspiration, while noting it is a commercial product. Between a commercial hosted service and a self-hosted container stack, the deciding question is whether you want the data on your own disk badly enough to run the containers.

## Clients, extensions and the API surface

Karakeep does not stop at a web UI. The README lists a Chrome plugin, a Firefox addon, a Safari extension, an iOS app and an Android app, plus mobile offline reading. There is a REST API and, per the README, a CLI and official skills aimed at LLM agents. Auto hoarding from RSS feeds and sync with browser bookmarks through floccus are also listed.

The practical consequence is that the server becomes the single store and the clients are thin. That is good for consistency and bad for anyone who wanted a lightweight client that works without the server. Mobile offline reading is the exception the README calls out, and it is worth testing on your own device before trusting it, since offline behaviour is exactly the kind of thing that varies by platform and is not detailed in the README. Importers exist for Chrome, Pocket, Linkwarden, Omnivore and Tab Session Manager, which matters if you are migrating rather than starting fresh.

## Licence, releases and upgrade cost

Karakeep is licensed under AGPL-3.0. The practical implication for most self-hosters is nil, since running it on your own server for yourself does not trigger the network copyleft clause in a way that obliges you to publish anything. The obligation matters if you modify Karakeep and expose it to other users over a network, because the AGPL's source-availability requirement is aimed at exactly that case. This is a description of the licence, not legal advice; read the LICENSE file and get proper counsel if you plan to build a service on top of it.

Release cadence is visible in the tags: v0.32.0 in May 2026, then v0.33.1 in early August 2026 and v0.33.2 on 2026-08-11. The project is still on 0.x, which means no compatibility promise across minor versions. Database migrations are handled by Drizzle through the db:migrate script in the root package.json, so an upgrade is not just pulling a new image; you need to run the migration step against your database, and you should have a backup of DATA_DIR before you do. The README does not document rollback, so treat the backup as your rollback.

## Conclusion

Adopt Karakeep if you already run Docker on a home server and want link previews, full page archival and automatic tagging over a personal hoard of links, notes and images. Skip it if you only need a plain read-it-later list, if you cannot give Puppeteer and Meilisearch a few gigabytes of headroom, or if you refuse to send content to an LLM and do not want to run ollama locally. Before committing, verify on the demo at try.karakeep.app that the tagging quality matches how you file things, then check the installation and configuration pages at docs.karakeep.app for the current image names and required environment variables, because the repository's .env.sample lists only DATA_DIR and NEXTAUTH_SECRET.

## FAQ

### Is Hoarder now Karakeep?

Yes. The README states that Karakeep was previously called Hoarder. Some older links and identifiers still carry the old name, such as the Android package id and the Weblate project URL.

### What is Karakeep and what does it do?

It is a self-hostable app for bookmarking links, notes, images and PDFs. It fetches titles, descriptions and preview images automatically, archives full pages with monolith, and adds LLM-based tagging and summarization.

### How do I install and set up Karakeep?

The README links installation to the Docker page of docs.karakeep.app, and the repository ships docker/, kubernetes/ and charts/ directories plus a karakeep-linux.sh script. Configuration is documented at docs.karakeep.app/configuration; the repository's .env.sample only lists DATA_DIR and NEXTAUTH_SECRET.

### Is Karakeep free and open source?

The repository is licensed under AGPL-3.0, so the code is open source. The README also mentions a managed Karakeep cloud at cloud.karakeep.app for people who do not want to self-host, with subscriptions that support development.

### How does Karakeep compare with Raindrop?

The README describes Raindrop as a polished open source bookmark manager for links, images and files, and notes that it is not self-hostable. Karakeep is self-hosting first, and adds notes, OCR, full page archival and LLM tagging to the same object model.

## Sources

- [Official documentation](https://karakeep.app)
- [Official README](https://github.com/karakeep-app/karakeep#readme)
- [Project repository](https://github.com/karakeep-app/karakeep)
- [Release notes](https://github.com/karakeep-app/karakeep/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/karakeep-app-karakeep
