# meme-search/meme-search: a self-hosted semantic meme search engine

> Meme Search indexes your own image collection with local vision models and vector search, then serves it from a loopback-only web app. It is a homelab tool for people who already have thousands of memes and no way to find the right one.

**meme-search/meme-search** — The open source Meme Search Engine and Finder.  Free and built to self-host locally with Python, Ruby, and Docker.

- Repository: https://github.com/meme-search/meme-search
- Website: https://meme-search.neonwatty.com/
- Stars: 753 · Forks: 27
- Language: Ruby
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/meme-search-meme-search

## What problem meme-search/meme-search actually solves

A folder of a few thousand memes is effectively unsearchable. Filenames are noise, and the text you remember is inside the image, not in the path. Meme Search targets that specific gap: it reads each image, produces a description, embeds that description into a vector, and lets you query it semantically. The README frames the goal as indexing "your memes by their content and text, making them easily retrievable." The audience is narrow and clearly stated by the project itself: people who self-host. There is no hosted tier described, and the homepage is a feature overview plus a demo, not a signup page. The topics list confirms the positioning: homelab, self-hosted, semantic-search, vector-database. If your memes live in a cloud photo service you do not control, this tool has nothing to point at.

## How the indexing pipeline is put together

The repository ships three cooperating pieces rather than one monolith. The Ruby on Rails application (the meme_search directory, plus the ghcr.io/neonwatty/meme_search image) serves the web UI and the Search API v1. A separate meme_search_jobs container runs ./bin/jobs, which is where background work such as description generation happens, so a long generation run does not block the request path. Postgres sits behind both, addressed as postgres://postgres:postgres@meme-search-db:5432/meme_search in the compose file. Description generation is local by default, governed by IMAGE_DESCRIPTION_PROVIDER, and the README lists the vision models you can choose from: Florence-2-base as the default, Florence-2-large, SmolVLM-256, SmolVLM-500, Moondream2, and Moondream2-INT8. The parameter counts matter more than the names. Florence-2-base is roughly 250 million parameters; Moondream2 is about 2 billion, and the README notes the INT8 variant drops memory from roughly 5GB to 1.5-2GB. That is the axis you are really choosing on. The README also states that embeddings and search stay local even when you switch description generation to an OpenAI-compatible vision API.

## Installing it and indexing a first batch of memes

The README's quick start is three commands and no build step, because the compose file pulls a prebuilt image. Run this from a directory where you are happy to create a meme_search folder:

```bash
git clone https://github.com/meme-search/meme-search.git
cd meme-search
docker compose up
```

The compose file mounts ./meme_search/memes/ to /rails/public/memes inside the container, and that mount is the one that matters. The comment in docker-compose.yml is explicit: any additional meme directory must be mounted there and placed under /rails/public/memes in the container. Adding a second collection means uncommenting the example line and pointing it at your own path.

```yaml
volumes:
  - ./meme_search/memes/:/rails/public/memes
  - ${MEME_SEARCH_DIRECT_UPLOADS_PATH:-./meme_search/direct-uploads}:/rails/public/memes/direct-uploads
  # -  /route/to/my/personal/additional_memes/:/rails/public/memes/additional_memes
```

Once it is up, open http://localhost:3000 and drag, drop, or paste images on the upload page. Two defaults are worth setting before you start. The .env.example shows APP_BIND_ADDRESS=127.0.0.1 and APP_PORT=3000, and the compose port mapping reads ${APP_BIND_ADDRESS:-127.0.0.1}:${APP_PORT:-3000}:3000. The README warns that the first local description generation downloads the selected model, so expect the first run to take noticeably longer than subsequent ones. If you would rather not run a local vision model, the openai compose variant exists and the .env.example documents the three keys it needs: OPENAI_API_BASE_URL, OPENAI_API_KEY, and OPENAI_VISION_MODEL, with gpt-4o-mini shown as the example model. The README states that selected images are sent to that endpoint when the provider is set to openai.

## The loopback-only boundary is a design decision, not a gap

Meme Search does not authenticate its web UI, and the project says so repeatedly rather than burying it. The README states the UI binds to 127.0.0.1 by default "because it does not include user authentication," and the .env.example adds that a trusted LAN is not an authentication boundary. The API is different: Search API v1 uses bearer tokens, but the README is precise that those tokens protect only the versioned integration endpoints and not the web UI or settings. MEME_SEARCH_ALLOWED_HOSTS exists for reverse-proxy setups and accepts a comma-separated list of exact Host values, with wildcards, URLs, ports, and CIDR ranges rejected. The .env.example is blunt that this does not authenticate the UI or make remote exposure safe. Read that as a deliberate scope limit. The project supports the loopback case well and declares everything else advanced and unsupported, which is more honest than shipping a half-built auth layer. If you need multi-user access with real accounts, this is the wrong tool and no configuration flag changes that.

## Where the local-first model choice bites

The default Florence-2-base keeps the footprint small, but the README's own model list shows the cost of scaling up. Moondream2 at roughly 2 billion parameters is the quality option, and the INT8 build is described as reducing memory from about 5GB to 1.5-2GB "with minimal quality loss" for memory-constrained hardware. That phrasing is the project's, and it is the only quality guidance given: there is no published comparison of description accuracy across the six models. You are choosing blind on quality and informed only on resource cost. The second constraint is the mount model. Because meme directories must land under /rails/public/memes in the container, adding a collection that lives somewhere else on the host is a compose edit, not a UI action. The README does not document rollback for a generation run, and it does not describe what happens to existing descriptions if you switch models after indexing a library. Anyone with a large collection should treat a model change as an open question rather than a reversible toggle.

## Self-hosted indexing versus hosted reverse image search

The obvious alternative is a hosted reverse image search: upload a picture, get visually similar results back from someone else's index of the public web. That is a different product with a different data flow. Hosted reverse image search matches your query against images it already crawled, which is useful for finding the source of a meme you did not create. Meme Search matches against images you own, using descriptions generated from their content, which is useful when you remember the joke but not the file. Neither substitutes for the other. The trade is straightforward: hosted tools require uploading your images and give you no control over the index, while Meme Search keeps processing local by default and gives the API a stable boundary that, in the README's words, "does not expose your database." The matching mechanism also differs. Semantic vector search over generated descriptions will find a meme described as a smug cat even if the query never says cat, but it depends entirely on the quality of the generated description. Keyword search over the same descriptions is available alongside it in the web app, which is the fallback when the embedding misses.

## Maintenance, releases and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-09-10. Releases are frequent and versioned: v2.3.0 on 2026-07-30, v2.3.1 the same week, v2.3.2 on 2026-08-05. The version in package.json (2.0.2) does not match the release tags, which is a small inconsistency worth knowing before you script anything against it. The compose file pins ghcr.io/neonwatty/meme_search:latest, so a plain docker compose up will follow new releases without asking. If you want reproducible upgrades, that tag is the thing to change. Upgrading also means re-pulling model weights when the selected model changes, which the README flags as the slow first-run path. The licence is Apache-2.0, a permissive licence that permits commercial use and modification and includes an explicit patent grant; the repository ships a LICENSE file. That is a description of the licence text, not legal advice, and if you plan to redistribute a modified build you should read the file yourself, including any NOTICE requirements it carries.

## Conclusion

Adopt it if you already keep a large local meme folder, can run Docker on a machine with enough RAM for the model you pick, and accept that the web UI has no authentication and is meant for loopback only. Skip it if you want a hosted service, a mobile app, or a search box over the public web rather than your own files. Before committing, check which image-to-text model you can afford (Moondream2-INT8 is the documented option for CPU-only machines) and confirm how the compose file mounts your meme directory, since only paths placed under /rails/public/memes inside the container are visible to the indexer.

## FAQ

### What exactly is meme-search/meme-search?

It is an open source meme search engine and finder that indexes your own images by their content and text so you can retrieve them later. The README describes it as built to self-host locally with Python, Ruby, and Docker.

### How do I install meme-search/meme-search?

The README's quick start clones the repository and runs docker compose up, which pulls the prebuilt image and starts the app. You then open http://localhost:3000 and upload images on the upload page.

### Does meme-search/meme-search send my memes to a cloud service?

By default, processing from image-to-text extraction to vector embedding to search is performed locally. You can optionally set IMAGE_DESCRIPTION_PROVIDER to openai and point it at an OpenAI-compatible vision endpoint, in which case selected images are sent there while embeddings and search stay local.

### Can I reach the meme-search/meme-search web app from another machine?

The web UI binds to 127.0.0.1 by default because it has no user authentication, and the README states that direct public exposure is unsupported. Proxy and VPN operation is described as advanced and unsupported, with authentication, TLS, and network controls left entirely to the operator.

### Which image-to-text model should I pick in meme-search/meme-search?

The default is Florence-2-base at roughly 250 million parameters, and the README also lists Florence-2-large, SmolVLM-256, SmolVLM-500, Moondream2, and Moondream2-INT8. The INT8 variant is documented as reducing memory from about 5GB to 1.5-2GB for memory-constrained hardware, ideal for CPU-only machines.

## Sources

- [License: Apache-2.0](https://github.com/meme-search/meme-search/blob/main/LICENSE)
- [meme-search/meme-search on GitHub](https://github.com/meme-search/meme-search)
- [Project website](https://meme-search.neonwatty.com/)
- [README](https://github.com/meme-search/meme-search/blob/main/README.md)
- [Releases](https://github.com/meme-search/meme-search/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/meme-search-meme-search
