Self-hosted service
groupultra/telegram-search avatar
groupultra/telegram-search

groupultra/telegram-search: export and fuzzy search your own Telegram history

🔍 导出并模糊搜索 Telegram 聊天记录 | Export and fuzzy search your Telegram chat history

4,109 stars274 forksTypeScriptAGPL-3.0

At a glance

What is it?
A self-hosted TypeScript tool that backs up personal Telegram chats into PGlite or PostgreSQL and searches them with tokenisation, vector matching and an optional bot. It is a personal archive, not a directory of other people's channels.
Who is it for?
Adopt it if you need to search your own Telegram history in a language Telegram's own search handles badly, and you accept running a database and a media store yourself. Do not adopt it if you are looking for a directory of public channels or a way to find strangers, because the README frames the project as being for exporting and retrieving your own chat records and warns against unlawful use.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap telegram-search fills

Telegram's built-in search is weak on Chinese and other languages that need word segmentation, and it has no way to match a sentence by meaning. The README opens with exactly that complaint, then adds a second one: important messages get buried under volume. The project's answer is to pull your own chat history out of Telegram, store it in a database you control, and index it twice, once with a tokeniser and once with embeddings.

The audience is narrow on purpose. This is for people who want their own archive searchable, not for people looking for public channels or strangers. The README carries an explicit warning that the software is for exporting and retrieving personal chat records and must not be used for anything unlawful, and it also warns that the project has issued no cryptocurrency. That framing matters when you read the feature list, because almost every capability is scoped to conversations your account can already see.

How the pipeline works: Telegram, tokeniser, vectors, storage

Messages arrive over Telegram's API and are written into a database chosen by DATABASE_TYPE, which the README lists as either postgres or pglite. PGlite is the embedded option, and the CLI documentation states that sync results go into a profile-scoped PGlite instance, so a CLI user can be productive without standing up a server. Media is separate: if MinIO parameters are configured, photos and stickers go to object storage, and the Docker image README says that without MinIO configuration, media files default to a local data/media directory.

Indexing happens during export, not as a later batch job. The feature list says messages are embedded and tokenised as they are exported, which is why search can mix fuzzy matching with vector similarity. The same embedding step is what makes image search possible: images get embeddings too, so a text query can retrieve a picture. On top of retrieval sits a RAG question-answering path and unread-message summarisation, and the README notes that AI embedding and LLM settings are now configured per account inside the app under Settings, then API.

The bot is a second surface over the same data. It can search and export messages and produce deep links back to the original conversation, which is the part that makes search results actionable rather than just readable.

Installing telegram-search with Docker Compose and running a first search

The README's quickest path is Docker Compose. Create an empty directory, download three files into it (the compose file, the .env example and init.sql), then bring the stack up. The compose file starts the surrounding services, which the README describes as the database and MinIO among others.

bash
mkdir telegram-search
cd telegram-search
curl -L https://raw.githubusercontent.com/groupultra/telegram-search/refs/heads/main/docker/docker-compose.yml -o docker-compose.yml
curl -L https://raw.githubusercontent.com/groupultra/telegram-search/refs/heads/main/docker/.env.example -o .env
curl -L https://raw.githubusercontent.com/groupultra/telegram-search/refs/heads/main/docker/init.sql -o init.sql
docker compose -f docker-compose.yml up -d

After that, the README says to open http://localhost:3333. You should see the web interface; the login step is where you supply your Telegram credentials. If you edit .env, the README is explicit that you must run the compose up command again for the change to take effect.

The single-container route skips Compose. This is the right choice if you want the media files on disk next to the container rather than in MinIO.

bash
docker run -d --name telegram-search -p 3333:3333 ghcr.io/groupultra/telegram-search:latest

Credentials come from my.telegram.org. The environment table lists TELEGRAM_API_ID and TELEGRAM_API_HASH, and the .env.example also defines VITE_TELEGRAM_API_ID and VITE_TELEGRAM_API_HASH, which the comments say the frontend needs for browser-only mode and which should match the server values. TELEGRAM_BOT_TOKEN, obtained from BotFather, is what enables the bot surface. Every variable in that table is described as optional with defaults, but without the Telegram API pair the tool has nothing to authenticate against.

There is also a local-first CLI, which the README presents as the agent-friendly path. A named profile holds the login, and the agent decides what to pull.

bash
pnpm run build:packages
pnpm -F @tg-search/cli build

pnpm cli --profile work profile configure --apiId 123456 --apiHash abcdef
pnpm cli --profile work auth login
pnpm cli --profile work chats list --json
pnpm cli --profile work sync --takeout --chat 123456 --from 2026-01-01 --to 2026-12-31
pnpm cli --profile work export --from 2026-01-01 --to 2026-12-31 --output ./telegram-2026

The README states that chats list and messages list are bounded remote reads that do not persist messages, and that bulk sync requires explicit user consent with --takeout added to the sync command. If consent is withheld, or Takeout initialisation fails, the run stops rather than falling back to ordinary GetHistory. Text sync requests only the selected chat categories, not contacts or file export permissions. Standard output carries JSON only; Telegram logs, prompts and progress go to stderr, which is what makes the CLI composable. Exports are monthly JSONL plus a checksum manifest, and the README says they exclude media binaries, login sessions, vectors and keys.

Where telegram-search stops being the right tool

The scope boundary is the first limitation, and it is deliberate. Nothing in the README describes discovering channels you are not a member of, resolving a username to a person, or searching anyone else's messages. If that is what you want, this project will not do it, and the README's warning about lawful use suggests the authors do not want it used that way either.

The second limitation is operational. This is not a hosted service you point at an account. You run a Node 24.13.0 or newer toolchain, or you run containers, and you keep a database alive. Choosing pglite removes the database server but the CLI documentation ties it to a profile, so multi-device access means either a server deployment or repeated local syncs.

The third is what the export does not contain. The CLI README is explicit that exports omit media binaries, login sessions, vectors and keys, and that the CLI will not run AI summaries on your behalf. So a JSONL export is a text archive, not a complete restore point, and semantic search over that export is something you would have to rebuild. The roadmap confirms the shape of the gap: automatic conversation summaries, a knowledge graph over people and events, link and image deep indexing with OCR, and support for Discord and other platforms are all listed as unchecked items. The roadmap is a plan, not a description of what ships in v1.3.1.

telegram-search compared with keeping everything in Telegram

The obvious alternative is Telegram itself: rely on the in-app search bar, saved messages and folders, and skip the infrastructure. The difference in approach is where the index lives. Telegram searches its servers with its own tokeniser and returns matches from your chats; telegram-search copies the messages into your database and builds two indexes you can inspect, one lexical and one vector. That is the whole trade: you gain language-aware tokenisation, sentence-level semantic matching, image search and summarisation, and you take on storage, backups and the job of keeping the copy in sync.

A second alternative is a general-purpose export tool that dumps your Telegram data to JSON and stops there. That gets you a portable file with no database, and it is the better answer if your goal is archival rather than retrieval. The difference is that a dump has no query layer: no tokeniser, no embeddings, no bot, no date-range filtering in a UI. telegram-search's CLI export is closer to that model, but the export command exists alongside a sync path that writes into PGlite, so the project assumes you will search the synced copy rather than the JSONL.

If you already run PostgreSQL with pgvector, the deployment question is mostly settled, because DATABASE_TYPE=postgres with a DATABASE_URL is the documented production configuration. If you do not, PGlite is the lower-friction entry point.

Maintenance, licensing and what you are taking on

The repository is not archived, and the last push was on 2026-09-01. Releases have been reasonably paced: v1.3.1 on 2026-08-18, after v1.2.8 and v1.2.7 in late June 2026. The project ships a Docker release workflow and a CI workflow, and the workspace is a pnpm monorepo with apps/ and packages/ directories, a Turbo build, Drizzle migrations and Vitest. That layout tells you upgrade cost is real: the root package.json pins pnpm 10.34.5 and requires Node 24.13.0 or newer, and the postinstall hook runs the package build, so a version bump can mean a dependency refresh plus a migration.

Upgrade mechanics are where the documentation is thinnest. The README documents re-running docker compose -f docker-compose.yml up -d after editing .env, and the repository has a drizzle/ directory and a db:generate script, which implies schema migrations exist. The README does not document rollback, and it does not describe a backup procedure for the database before an upgrade. Treat that as something you design yourself.

The licence is AGPL-3.0. That is a copyleft licence with a network clause: if you modify the software and let users interact with it over a network, the licence's terms reach that deployment. Running it privately for your own chat archive is a different situation from offering it as a service to others. This is a description of the licence, not legal advice; if you plan to host it for anyone but yourself, read the licence text or ask someone qualified.

Editorial conclusion

Adopt it if you need to search your own Telegram history in a language Telegram's own search handles badly, and you accept running a database and a media store yourself. Do not adopt it if you are looking for a directory of public channels or a way to find strangers, because the README frames the project as being for exporting and retrieving your own chat records and warns against unlawful use. Before committing, check the CLI README for the exact sync consent flow, confirm whether you want DATABASE_TYPE=postgres or pglite, and decide where media lands: MinIO if MINIO_URL is set, otherwise data/media.

Frequently asked questions

What is groupultra/telegram-search?

It is a self-hosted tool that exports your own Telegram chat history into PGlite or PostgreSQL and indexes it for fuzzy and vector search. The README describes it as being for exporting and retrieving personal chat records, with optional media backup to MinIO.

How do I install and run groupultra/telegram-search?

The README's quick start downloads docker-compose.yml, .env and init.sql into an empty directory and runs docker compose -f docker-compose.yml up -d, after which the interface is at http://localhost:3333. A single docker run of ghcr.io/groupultra/telegram-search:latest on port 3333 is also documented, and a local CLI path builds the packages and configures a named profile.

How do I use the groupultra/telegram-search bot?

A bot token from BotFather goes into TELEGRAM_BOT_TOKEN, and the README lists searching and exporting messages through the bot plus deep links that jump to the original conversation. The README does not document the bot's exact command syntax, so check the repository for that.

Can groupultra/telegram-search find people or public channels I am not in?

No. The README scopes the project to exporting and retrieving your own chat records and warns against unlawful use, and no feature in the documentation covers discovering channels or resolving usernames to people.

Is groupultra/telegram-search safe to run?

The README states that the CLI's chats list and messages list are bounded remote reads that do not persist messages, and that bulk sync requires explicit consent with --takeout and stops if Takeout initialisation fails. Exports are documented as excluding media binaries, login sessions, vectors and keys. The project also warns that it has issued no cryptocurrency.

How do I access my groupultra/telegram-search history after exporting?

The README says exports are monthly JSONL files plus a checksum manifest, and that they exclude media binaries, login sessions, vectors and keys. The searchable copy lives in the synced database, which for the CLI is a profile-scoped PGlite instance, so the JSONL itself is an archive rather than a query surface.

Official sources

  1. groupultra/telegram-search on GitHub
  2. License: AGPL-3.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/groupultra-telegram-search.svg)](https://hysenlabs.com/projects/groupultra-telegram-search)