Self-hosted service
lcomplete/huntly avatar
lcomplete/huntly

Huntly: a self-hosted reading list with Lucene search and a capture extension

Your Personal AI-Powered Information Hub

2,347 stars197 forksTypeScriptApache-2.0

At a glance

What is it?
A browser extension and server pair that saves pages, RSS feeds and tweets into your own SQLite database, then makes them searchable with Apache Lucene and reachable by an AI agent over MCP.
Who is it for?
Huntly is a personal information system rather than a team tool: one server, one administrator account, a database you own, and an extension that feeds it. The parts that make it more than a bookmark manager are Lucene search with Chinese tokenization, Twitter thread reconstruction, and an MCP server that lets an assistant query the same index.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 130 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.

Editorial analysis

One server, three clients: web, extension and desktop

Huntly is built as a server with clients attached, and the feature table states the architecture more precisely than the tagline does. The storage layer is SQLite, deployment is Docker, and the README is explicit that data ownership is complete and self-hosting is 100 percent.

Three clients talk to that server. There is a web app, a Chrome extension on Manifest V3, and desktop apps built with Tauri. The extension is the component that does the capturing, and the README is direct about it: the extension is essential for automatically saving web pages and Twitter content. Everything else is a way of reading and searching what the extension collected.

That split is what makes the project more than a read-later service. The server is the archive; the extension is the sensor. Because the extension is described as sending only relevant browsing data such as saved pages and tweets to the server, rather than everything you visit, the privacy story depends on configuration rather than on trust in a third party.

For remote access the README recommends HTTPS for privacy, which is the right advice and also a real operational step. The extension has a server address setting, and the web app handles first-run administrator registration, after which you are logged in automatically.

Docker Compose, with a Watchtower label already in the file

The recommended install is a Docker Compose file, and the example is complete enough to copy. Note that it already carries the Watchtower label, so the README's advice about automatic updates is built into the file rather than left to you.

yaml
version: '3.8'

services:
  huntly:
    image: lcomplete/huntly
    container_name: huntly
    restart: always
    ports:
      - "8088:80"
    volumes:
      - ~/data/huntly:/data
    labels:
      - "com.centurylinklabs.watchtower.enable=true"

Then docker-compose up -d brings it up. The port mapping is 8088 on the host to 80 inside the container, and the volume mount is what holds your data, so the single path in that file is the thing to back up.

The Dockerfile in the repository shows what the container actually is. It builds from adoptopenjdk/openjdk11:latest, declares a volume at /data, creates /data and /data/lucene, and takes the application as a single JAR through build arguments. The entrypoint passes JVM memory settings, a hardcoded GMT+08 timezone, a Spring profile, a port, and both huntly.dataDir and huntly.luceneDir pointing into the mounted volume.

Two things in that Dockerfile are worth flagging. A timezone of GMT+08 baked into the image is a decision that will be wrong for most readers and cannot be changed through the exposed APP_ARGS. And openjdk11 is not the current long-term release of Java, so an image tracking latest on that base has aged with the project rather than with the platform.

Lucene with a Chinese analyzer, and search that understands structure

The search layer is the part of Huntly that goes beyond storing and filtering. It uses Apache Lucene with the IK Analyzer, named specifically for Chinese text tokenization, alongside boolean operators and fuzzy search. For a project whose user base skews toward Chinese-language material, picking a Chinese analyzer as the default rather than an add-on is a real design decision.

The releases show how much of the work has been in making search behave predictably. Version 0.6.4 added date-range filtering to the MCP search tool through start_date, end_date and date_field parameters, so an agent can retrieve content from a specific time window rather than paging through everything.

That same release made structured filters first-class fields on SearchQuery instead of being encoded in a legacy queryOptions string, and it added a cap of 20 results per page tied to Lucene's MAX_LUCENE_RESULT_WINDOW limit. A cap set by the engine's own constraint rather than by preference is worth knowing about if you plan to script against the search API.

Quote support arrived in the same release, so search tokens containing spaces can be wrapped in double quotes, as in collection:"my collection". The collection list in the interface auto-quotes names containing spaces when it pre-fills the search box, which is a small detail that tells you the author has actually hit this bug.

Version 0.6.6 added time range search to the web interface and a scheduled backup feature built on SQLite VACUUM INTO, with an API to list and download snapshots, a settings toggle, a configurable path, a keep-count and a file list.

The MCP server, Agent Skills, and what an assistant can reach

The feature table lists MCP and Agent Skills together, describing an MCP server plus Agent Skills that let AI assistants search your knowledge base, RSS feeds, tweets and highlights, installed with npx skills add lcomplete/huntly. This is the feature that separates Huntly from a conventional read-later service, and it is also the one that deserves the most careful thought before enabling.

What an agent can reach is your archive: everything you saved, every feed you follow, every tweet you bookmarked. The release history shows the surface growing in a specific order. Version 0.6.4 added ListCollections, ListCollectionContent and GetXSaveRules as new MCP tools, described as enabling richer agent workflows, and 0.6.6 added time range search so retrieval can be bounded.

Version 0.6.5 added checksum verification and JWT secret management for the embedded server, described as improvements to embedded server security. Those are the right things to fix, and they are worth understanding in context: the desktop application embeds the server, so a desktop install is not simply a client of a remote deployment.

The same release added a backend-driven update check, with a new Tauri command that fetches the latest.json manifest from GitHub releases and hands off to the updater plugin, removing a dependency on a hardcoded update endpoint. It also replaced a one-shot auto-update flag with a 24-hour interval scheduler persisted through localStorage, with correct cleanup on unmount and toggle. Small engineering, but the kind that separates a desktop app people keep from one they uninstall.

A roadmap entry that contradicts the feature table

The README contains a direct contradiction, and it is more useful than a clean feature list because it shows where the project's centre of gravity is right now.

The feature table lists AI content processing for summarization, translation, browser extension chat and intelligent content analysis with custom shortcuts, presented as a delivered capability. The roadmap section below it shows three items checked off: export all saved content to Markdown, flexible organization through collections, and an enhanced extension with standalone AI processing that requires no server. The fourth item, built-in browser extension chat with page context, attachments and AI shortcuts, is unchecked.

So browser extension chat appears in the feature table as something the product does and in the roadmap as something not yet finished. The reconciliation is probably that standalone AI processing works inside the extension while page-context chat with attachments is the unfinished part, but the README does not say so, and the screenshots are named extension_shortcuts rather than extension_chat.

Other feature-table entries are specific enough to be checkable. Web archiving uses Defuddle and Mozilla Readability for content extraction. RSS management includes intelligent categorization, OPML import and export, and full-text search. Social media integration has special handling for Twitter and X, with automatic tweet thread reconstruction and media preservation. GitHub integration syncs starred repositories with metadata and README extraction.

The repository was last pushed on 2026-05-30 and the most recent release is v0.6.6 from May 2026, on a 2,347-star project with only 10 open issues. The tree includes AGENTS.md and CLAUDE.md at the root, along with a skills directory, which means the project ships instructions for automated coding agents alongside the features that give those agents something to query.

Where the README stops and the wiki starts

The README covers the Docker Compose path, the desktop client download, extension installation, and three configuration steps. Beyond that it defers: the Run the Server wiki page is linked as the place for other options, and the project website carries separate documentation and download pages.

One instruction deserves a caveat rather than a transcription. If you hit the error that Huntly.app is damaged and can't be opened on macOS, the README tells you to run sudo xattr -r -d com.apple.quarantine on the application path. That command removes the quarantine attribute macOS applies to downloaded applications, and it is the standard fix for this class of error. It also bypasses a Gatekeeper check that exists precisely to warn you about applications you did not build. Running a Tauri desktop client inside a virtual machine is a way to evaluate Huntly without that tradeoff, and the MCP surface is worth sizing up before you give a desktop binary the benefit of the doubt.

The repository tree is small for a project of this scope: a Dockerfile, a Dockerfile.multistage, a docker-compose.yml, app, skills and static directories, plus two README files and licence. The absence of a visible source tree for the server suggests the Java server is built elsewhere and delivered as the JAR the Dockerfile expects, which matches the build arguments in that file.

Editorial conclusion

Huntly is a personal information system rather than a team tool: one server, one administrator account, a database you own, and an extension that feeds it. The parts that make it more than a bookmark manager are Lucene search with Chinese tokenization, Twitter thread reconstruction, and an MCP server that lets an assistant query the same index. Start with the Docker Compose path, since the desktop client and the container both front the same server, and set up HTTPS before pointing the extension at anything remote.

Frequently asked questions

What is Huntly and what does it do with my browsing data?

Huntly is a self-hosted server with a Chrome extension, web app and desktop client. The extension captures pages, RSS items and tweets into your own SQLite database, which you can then search locally. The extension is described as sending only relevant data such as saved pages and tweets rather than a complete browsing history.

Do I need Docker to run Huntly?

No, Docker Compose is the recommended path but not the only one. Desktop installation packages are published on the releases page for each operating system and bundle a server you can run locally. The wiki page linked from the README covers further server options.

What search engine does Huntly use?

Apache Lucene with the IK Analyzer for Chinese text tokenization, boolean operators and fuzzy search. Recent releases added time range filtering, quoted phrase support for values containing spaces, and a 20-result page cap set by Lucene's own max result window.

What can an AI assistant do with my Huntly archive?

Huntly ships an MCP server and Agent Skills, installable with npx skills add lcomplete/huntly, that let an assistant search your saved content, RSS feeds, tweets and highlights. Tools include search with date filtering plus ListCollections, ListCollectionContent and GetXSaveRules, so the assistant can reach both your content and your organisation rules.

How do I keep my Huntly data safe?

Everything lives in the SQLite database and Lucene index under the mounted /data path, so backing up that directory is the whole job. Version 0.6.6 added scheduled VACUUM INTO snapshots with a configurable schedule, path, keep-count and a download list, and Huntly recommends HTTPS if you reach the server remotely.

Official sources

  1. lcomplete/huntly on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/lcomplete-huntly.svg)](https://hysenlabs.com/projects/lcomplete-huntly)