CLI tool
kenn-io/msgvault avatar
kenn-io/msgvault

msgvault: a local-first archive for a lifetime of email and chat

Archive a lifetime of email and chat. Offline search, analytics, and AI query over your full message history. Powered by SQLite and DuckDB

2,066 stars156 forksGoMIT

At a glance

What is it?
A Go binary that syncs mail, chat, meetings and calendars into a SQLite and DuckDB archive you can search, query with SQL, and hand to an agent over MCP.
Who is it for?
msgvault is doing something that most mail clients deliberately do not: treating your history as an archive you own rather than a synchronized view of a provider's database. The four command quick start gets from nothing to a served archive in about a minute, and everything after that, SQL, the web UI, the TUI, the MCP server, is built on the same local store.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 18 days ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The short version: four commands to a served archive

The README's Gmail path is four commands long, and the order matters. The first creates the database, the second registers a provider account, the third pulls a bounded batch of history, and the fourth starts the server that hosts the browser application.

bash
msgvault init-db
msgvault add-account [email protected]
msgvault sync-full [email protected] --limit 100
msgvault serve

Two details are worth pausing on. Gmail needs an OAuth credential created beforehand, and the guide for that lives on the docs site rather than in the repository. And the sync is deliberately limited to 100 messages, which is a sane default for a first run against a mailbox with years in it.

The last command prints an API server URL. The release binary carries the compiled browser application inside itself, so nothing else needs to be installed to view it, and msgvault tui opens the terminal interface for the same data.

Sources, and what each one costs you

The source guide is where the design becomes visible. msgvault reads mail over IMAP and JMAP, pulls PST files through a Go library built for them, syncs Slack, and imports meeting notes from Circleback. Release notes show the seams: v0.19.2 added support for Fastmail pod-scoped JMAP API URLs, and v0.19.3 added rate limit handling for Circleback syncs plus repairs for dangling message recipients during database upgrades.

The privacy story is stated unusually precisely for this category of tool. Keyword search and analytics read the archive without contacting the source services. Sync, by definition, needs access to those services. And the optional AI features, meaning embeddings, semantic search and profile automation, send selected data to whatever endpoint you configure, with local embedding servers supported as an alternative. The README also notes that profile automation and conversation briefs require separate consent.

That layering is the sensible design. The default path stays entirely local, and the network-touching features are opt-in and separately documented under a recommended configuration page.

SQLite, DuckDB, Postgres, and why all three

The project description says it is powered by SQLite and DuckDB. The dependency list says a bit more. The go.mod pulls in the SQLite driver, the DuckDB driver, and pgx for PostgreSQL, plus sqlite-vec bindings for vector search. The Makefile then sets build tags of fts5 and sqlite_vec on every compile, so the full text search extension and the vector extension are compiled in rather than loaded at runtime.

The Makefile is also unusually candid about why PostgreSQL support exists. A long comment explains that go test defaults its parallelism to the host CPU count, that every PostgreSQL-backed test binary opens its own connections including a pinned admin handle, and that on a wide runner the sum exceeds a stock server's 100 connection limit. The cap of four binaries is described as the GitHub-hosted profile the lanes were tuned on.

That comment is a small signal about how the project is maintained. Build flags, test shards and the reason behind each one are written down in the file rather than living in someone's head.

The SQL surface matters for anyone planning to build on this. The README lists querying with SQL, exporting messages and attachments, deduplicating copies and reviewing staged mail deletions before removing anything upstream. Remote deletion preserves the archived copy, which is the behaviour you want from an archive and the opposite of what the provider does.

Four ways in: browser, terminal, HTTP and MCP

The same archive is reachable four ways, and the README treats them as peers rather than as one interface plus extras. There is a web application, a TUI built on Bubble Tea and Lip Gloss, an HTTP API described by an OpenAPI file in the api directory, and an MCP server for connecting an agent.

The MCP endpoint is the part that changes how people use this. An archive you can query in SQL is useful; an archive an assistant can search on your behalf is useful in a different way. The project leans into it, with a skills directory in the tree and a separate AGENTS.md for agents working on the codebase itself, which is a different concern from agents querying your mail.

There is also a browse and group axis: messages grouped by people, domains, time, source and type, with saved views. Contact activity is tracked, identities across different addresses and handles can be connected into one person, and contacts sync over CardDAV. For anyone who has lost track of who someone was, that identity graph is probably the feature that justifies the whole project more than the search does.

Building it yourself is heavier than running it

Installing the release binary is a single script that detects your OS and architecture, downloads from GitHub Releases, verifies a SHA-256 checksum, and installs. The scripts are linked so you can read them first, and Homebrew and conda-forge packages exist alongside them.

Building from source is a different proposition. The README lists Go 1.27 or newer, Bun 1.3.14 or newer, Node 20.19 or newer on the 20 line, 22.13 or newer on 22, or 24 or newer, and a C and C++ compiler for CGO, since DuckDB is statically linked. On Debian and Ubuntu you also need libsqlite3-dev for the sqlite3.h header used by the sqlite_vec build.

bash
git clone https://github.com/kenn-io/msgvault.git
cd msgvault
make install

The Dockerfile explains the build layout in stages: a Bun stage builds the web frontend, and the comment at the top of that stage notes that Bun exists only there, because the release image contains neither a JavaScript runtime nor a filesystem web distribution. The Go stage installs gcc, g++, make, git and libsqlite3-dev, then builds with CGO_ENABLED=1 and the fts5 and sqlite_vec tags. There is a deliberate line that deletes anything already sitting in the web dist directory before copying the frontend build in, with a comment about not trusting an ambient or host-staged bundle.

The result is a single binary that serves its own UI, which is why the runtime requirements are much lighter than the build requirements.

Alpha software with fifty open issues

The README opens its main section with a warning that is worth taking at face value: APIs, the storage format and CLI flags may change, and it tells you to back up your data. It also notes that the document describes current main rather than a release, and points at a changelog section for unreleased features.

The numbers behind that caution are 2066 stars, 156 forks and 50 open issues. Fifty open issues against that star count is high, and some of the recent release notes explain why the project needs a release cadence rather than a single announcement: v0.19.3 starts the API immediately while the analytics cache builds in the background, v0.19.2 bounds relationship index builds to improve resource usage, and v0.19.1 stops Windows daemons cleanly during releases and recovers interrupted releases from existing tags.

Those are the kinds of fixes that only appear once real people run the thing for real. The repository tree reflects the shape of the work: an api directory for the HTTP surface, cmd and internal for the binary, pkg for shared pieces, web for the frontend, website for the docs, testdata for fixtures, and a tools directory alongside scripts.

Alpha status here means the storage format may move, not that the tool is a prototype. The dependency list is long and specific, with pgx, huma for the API layer, bubbletea, cobra for the CLI, cron for scheduling, prometheus-style observability libraries and OIDC for auth. It is a real codebase with a young interface.

Editorial conclusion

msgvault is doing something that most mail clients deliberately do not: treating your history as an archive you own rather than a synchronized view of a provider's database. The four command quick start gets from nothing to a served archive in about a minute, and everything after that, SQL, the web UI, the TUI, the MCP server, is built on the same local store. It is still labelled alpha with 50 open issues, so treat the first sync as something to back up. Start with msgvault init-db and a limited sync on one account, then read docs/usage/recommended-configuration.md before turning on any feature that sends data to a model endpoint.

Frequently asked questions

Where does msgvault store my messages?

On your own hardware, in a local database. The project ships drivers for SQLite and DuckDB and also depends on pgx for PostgreSQL, and full text search plus vector search extensions are compiled into the binary through the fts5 and sqlite_vec build tags. The README states that keyword search and analytics read the archive without contacting the source services.

Which chat and mail services does msgvault sync?

Mail over IMAP and JMAP, including Fastmail pod-scoped JMAP URLs, PST files, Slack, and meeting notes from Circleback. The source guide on the docs site is the authoritative list. Releases in August 2026 added Fastmail URL handling and Circleback rate limit respect, which tells you where the recent provider work has gone.

Can I connect an AI assistant to my msgvault archive?

Yes, through the bundled MCP server. The same archive is also reachable from a browser application, a terminal interface and an HTTP API described by an OpenAPI file. The optional model and enrichment features send selected data to the endpoints you configure, and local embedding servers are supported as an alternative.

How long does it take to set msgvault up?

Four commands, once you have an OAuth credential for Gmail. Run msgvault init-db, msgvault add-account with the address, a bounded sync such as msgvault sync-full with a limit, then msgvault serve. Building from source is much heavier than installing the binary, since it needs Go 1.27, Bun, Node and a C compiler for CGO.

Does deleting an email in msgvault delete it at the provider?

The archive stages deletions for review before anything is removed upstream, and remote deletion preserves the archived message and its attachments. That is deliberate, and it is the opposite behaviour from the provider, which is the point of keeping an archive.

Official sources

  1. kenn-io/msgvault on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/kenn-io-msgvault.svg)](https://hysenlabs.com/projects/kenn-io-msgvault)