# trace.moe: self-hosting the anime scene search engine

> trace.moe is an index page for a distributed anime screenshot search system, not a single application. Self-hosting it means running six containers, feeding it video files named by AniList ID, and waiting on a Milvus index that can take days to settle.

**soruly/trace.moe** — Timestamp Retrieval for Anime Clips Everywhere

- Repository: https://github.com/soruly/trace.moe
- Website: https://trace.moe
- Stars: 5,035 · Forks: 261
- Language: Unknown
- License: MIT
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/soruly-trace-moe

## The gap trace.moe fills: a screenshot with no filename

A cropped anime frame carries no metadata. There is no filename, no episode number, no timestamp burned into the pixels. The README states the goal plainly: it tells you which anime, which episode, and the exact moment this scene appears. That is a different problem from reverse image search on stills, because the answer is a position inside a video file, not a matching picture. The audience is narrow and specific. People who build bots, browser extensions and chat integrations around anime screenshots, and people who want that lookup running on their own machine rather than through the hosted site. The repository is an index page for the whole system, and it lists the parts by name: an API server, a web server, a Telegram bot, an MCP server for AI agents, a browser extension, and a library that generates the image vectors. If you want the hosted service, you use trace.moe directly. If you want the machinery, you clone this repo and read compose.yml.

## How the pieces fit: Postgres for files, Milvus for vectors

The architecture visible in compose.yml is a two-store split. Postgres 18 holds the file records and indexing status. Milvus holds the vectors, and it does not run alone: the compose file also starts etcd for metadata and MinIO for object storage, which is the standard Milvus deployment shape. The api service depends_on both postgres and milvus, so ordering is handled for you. The api container mounts your video directory read-only at /app/video/, and the README says it scans VIDEO_PATH every minute for new files, accepting .mp4, .mkv and .webm only. Everything else is ignored. Hashing is parallelised by MAX_WORKER, set to 4 in compose.yml, and the README notes you can raise it to make hashing faster. The www container talks to the API through NEXT_PUBLIC_API_ENDPOINT, which compose.yml points at http://localhost:3001, the host-side port that maps to the API's internal 3000. That port split is worth noticing before you start changing things: the API listens on 3000 inside its container and is published on 3001 on your machine.

## Installing trace.moe with docker compose and a first search

The README requires docker compose, with Windows supported through WSL2. The video layout is the part people get wrong. Each show lives in a folder named after its AniList ID, and the full path should look like /mnt/c/trace.moe/video/{anilist_ID}/foo.mp4. Create that directory, copy .env.example to .env, and point VIDEO_PATH at it.

```bash
cp .env.example .env
# .env
VIDEO_PATH=/mnt/c/trace.moe/video/
```

With that file in place, the README's next step is a single command. The api container will begin scanning the mounted directory within a minute of starting, and the README says you can watch the indexing process in the api server logs.

```bash
docker compose up -d
```

Once the containers are up, open http://localhost:3000 and search. The www service is published on 3000 and the API on 3001, so the page you load is the web front end, not the API. If you want to reach the API directly, it is on 3001. Adminer is also started on 8080 with ADMINER_DEFAULT_SERVER set to postgres, which gives you a browser view of the file table while indexing runs.

## Loading the pre-hashed dump instead of hashing your own library

Hashing a library takes time proportional to its size, so the README offers a shortcut: a database dump published on Hugging Face. The trade is memory. The README states that loading all 114,514 files into memory requires about 140GB RAM, which is the single most important number on the page. The procedure starts with zstd, which the README installs per distribution.

```bash
sudo apt install zstd   # Ubuntu / Debian
```

Then start the containers, restore the dump into Postgres, and reset the loaded flag so the API picks the hashes up. The README says to ignore errors like ERROR: relation "xxx" already exists during the restore.

```bash
docker exec -i tracemoe-postgres-1 psql -U postgres postgres < <(zstdcat dump.sql.zst)
docker exec -i tracemoe-postgres-1 psql -U postgres postgres < <(echo "UPDATE files SET loaded=NULL")
```

The second step is the slow one. The README states this process may take 24 hours, and gives a query to watch it, grouping the files table by status. Once every row reads LOADED, search works, but background optimisation inside Milvus may take a few days to complete. That last sentence matters more than it looks: the system answers queries before it is done settling.

## What self-hosting trace.moe actually costs you

The README is unusually direct about the resource profile, and the numbers are worth repeating without softening. About 140GB of RAM for the full pre-hashed set. A restore measured in minutes and a hash load measured in up to 24 hours. A Milvus optimisation phase measured in days. The README does not document a rollback path for a partially loaded database, and it does not describe how to remove a single show once it has been indexed; the cleanup instructions cover the whole system, not one entry. docker compose down -v removes all containers and deletes all database volumes, and the README then tells you to clean up VIDEO_PATH manually. That is an all-or-nothing reset. There is also a security surface the README does not discuss: compose.yml ships TRACE_API_SALT=SALT, DB_PASS=postgres, MILVUS_TOKEN=root:Milvus and a Postgres password of postgres, all as literal defaults. Those are fine on a laptop behind a firewall and not fine on anything reachable from the internet. Changing them is not documented in the README, so you are reading compose.yml and making your own calls.

## trace.moe against a general reverse image search

The obvious alternative is a general reverse image search engine, which takes a picture and returns pages where that picture or a near copy appears. The difference in approach is structural. A general engine matches against crawled web images and returns documents. trace.moe matches a query frame against vectors extracted from video files you control, and returns a position: title, episode, timestamp. That is why the answer format is a timestamp and not a link, and why the indexing input is a folder tree named by AniList ID rather than a crawl. The trade-off runs both ways. A general engine needs no local storage and no Milvus cluster, but it cannot tell you the second a frame appears, because it was never indexing seconds. trace.moe needs your disk, your RAM and your patience, and in exchange it answers a question a web image index cannot represent. If your query images are manga pages, live action stills or photographs, neither the vector library nor the folder convention is aimed at you.

## Licence, upgrades and the maintenance picture

The repository is MIT licensed, which is permissive and places few conditions on reuse beyond preserving the notice; this is a description of the licence text, not legal advice, and the LICENSE file at the repository root is the authority. Note that the licence covers this index repository, and the README points to separate repositories for the API, the web server, the Telegram bot, the MCP server, the browser extension and the image vector library. If you plan to redistribute or modify a component, check that component's own licence rather than assuming MIT propagates. On maintenance, the last push to this repository was on 2026-09-13. The repository is not archived. Upgrades are effectively image upgrades: compose.yml pins ghcr.io/soruly/trace.moe-www:latest and ghcr.io/soruly/trace.moe-api:latest, so a docker compose pull followed by docker compose up -d moves you forward, and there is no version number in the file to tell you what you moved to. Postgres is pinned to 18, but the Milvus-related images and the rest of the truncated compose file are not something the README explains, so read the whole file before you pull.

## Conclusion

Adopt trace.moe if you already hold a local anime video library and want screenshot-to-timestamp search on your own hardware, and if you can spare the RAM the README quotes for the pre-hashed route (about 140GB for 114,514 files) plus a Milvus optimisation window measured in days. Do not adopt it as a general visual search engine for manga, live action or photographs; the pipeline is built around video frames and AniList IDs, and the README's own example image is an anime screenshot. Before committing, verify the disk and memory headroom on your host, confirm your files sit under VIDEO_PATH in {anilist_ID}/name.mp4 layout, and decide whether you will hash your own library or load the published dump, because those two paths have very different waiting times.

## FAQ

### What is trace.moe?

It is an anime scene search engine: you give it an anime screenshot and it tells you which anime, which episode and the exact moment the scene appears. The repository is the index page for the whole system, which is split across separate API, web, bot and vector-generation projects.

### What anime does trace.moe find?

It searches the video files you index, which the README says must be placed under VIDEO_PATH in folders named by AniList ID, using .mp4, .mkv or .webm files. There is no separate catalogue in the repository; what it can find is what you feed it.

### Is trace.moe safe?

The README does not address safety or privacy, and the repository contains a SECURITY.md file but no security guidance in the documentation. What the README does show is that compose.yml ships default credentials for Postgres, Milvus and the API salt, so a self-hosted instance exposed to a network should not keep those values.

### What is the best anime scene search engine?

This repository does not compare itself to other engines, so it offers no basis for ranking them. The README does describe what trace.moe returns, which is the anime, the episode and the timestamp of the scene, and that is the concrete capability to compare against.

## Sources

- [Issues](https://github.com/soruly/trace.moe/issues)
- [License: MIT](https://github.com/soruly/trace.moe/blob/master/LICENSE)
- [Project website](https://trace.moe)
- [README](https://github.com/soruly/trace.moe/blob/master/README.md)
- [soruly/trace.moe on GitHub](https://github.com/soruly/trace.moe)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/soruly-trace-moe
