# clip-retrieval: a CLIP embedding and retrieval stack you run yourself

> clip-retrieval turns CLIP embeddings into a searchable index with a Flask backend and a browser front end. It is a pipeline of separate commands, not a hosted service, and that shape decides who gets value from it.

**rom1504/clip-retrieval** — Easily compute clip embeddings and build a clip retrieval system with them

- Repository: https://github.com/rom1504/clip-retrieval
- Website: https://rom1504.github.io/clip-retrieval/
- Stars: 2,798 · Forks: 237
- Language: Jupyter Notebook
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/rom1504-clip-retrieval

## What clip-retrieval is for, and who it is not for

The project targets a narrow but awkward job: taking a pile of image URLs and captions and ending up with a search box that answers text queries with similar images. Doing that by hand means picking a CLIP checkpoint, batching inference over the images, choosing a FAISS index type, writing an HTTP layer, and building a UI. clip-retrieval packages all five steps as separate commands, so you can run them in sequence or stop after any one of them.

The intended user is someone with a dataset already in hand, most likely produced by img2dataset, the sibling project the README links. The README's own examples are dataset-scale: it states that 100M text plus image embeddings can be processed in 20 hours on a 3080, and that inference runs at roughly 1500 samples per second on the same card. Those are the author's figures, not measurements of your data, and they say nothing about how long the index build takes on a given machine.

If you want a hosted search API and nothing else, the project is the wrong shape. It is also a poor fit if you need incremental updates, because the pipeline is built around computing embeddings for a whole dataset and then building an index over them. There is a clip filter stage for narrowing data using the index, but the README does not describe a supported path for adding a few thousand new images to an existing index without rebuilding.

## The five commands and how data moves between them

The architecture is a straight line. clip inference reads images and produces embeddings plus metadata, clip index turns those embeddings into a FAISS index, clip back loads the index and serves queries over HTTP, and clip front is a browser UI that talks to the back. clip client is the Python-side counterpart to the front end: it sends a query to a backend and returns results as dictionaries.

The index is the center of the system. autofaiss is the dependency that chooses an index structure, and the README lists faiss-cpu in requirements.txt, so the default install runs FAISS on CPU. The backend is a Flask service with flask_restful and flask_cors, and it exposes Prometheus metrics through prometheus-client, which is the one piece of operational plumbing the dependency list reveals.

Queries can arrive as text, as an image path or URL, or as a precomputed embedding. On the way out, the backend can apply an aesthetic score filter, an optional multilingual CLIP variant, and two safety filters. Those are backend concerns, not client concerns, which is why the client exposes them as constructor arguments rather than per-query options.

## Installing clip-retrieval and running the end2end example

The README gives one install line. It pulls a large dependency set including torch, torchvision, faiss-cpu, autofaiss, flask and open-clip-torch, so expect a heavy environment rather than a small library.

```bash
pip install clip-retrieval
```

For local development the Makefile does the same thing with an editable install, which is what you want if you plan to read the source alongside the docs.

```bash
python -m pip install -U pip
python -m pip install -e .
```

The fastest way to see the whole pipeline work is the end2end command. It downloads a small test parquet file of image URLs and captions, then runs inference, index, back and front over it. The README suggests clearing CUDA_VISIBLE_DEVICES first if your GPU does not have enough VRAM, which forces the CPU path.

```bash
export CUDA_VISIBLE_DEVICES=
wget https://github.com/rom1504/img2dataset/raw/main/tests/test_files/test_1000.parquet
clip-retrieval end2end test_1000.parquet /tmp/my_output
```

When it finishes, the README says to open http://localhost:1234 and search among your pictures. If you only want the index and not the web service, the same command accepts --run_back False. That flag is the one documented way to split the pipeline, and it is worth knowing before you script anything around this tool.

## Querying a backend from Python with ClipClient

The client is the part most likely to end up inside another program. Initialization takes a backend URL and an index name, both required, and a set of optional parameters with defaults: aesthetic_score 9, use_mclip False, aesthetic_weight 0.5, modality Multimodal.IMAGE, num_images 40, deduplicate True, use_safety_model True, use_violence_detector True.

The README's example points at the hosted Laion5B backend. That endpoint is a public demonstration, and the README does not describe its availability guarantees, rate limits or privacy properties, so treat it as a way to try the client rather than as infrastructure.

```python
from clip_retrieval.clip_client import ClipClient, Modality

client = ClipClient(url="https://knn.laion.ai/knn-service", indice_name="laion5B-L-14")
results = client.query(text="an image of a cat")
```

The README shows the shape of a result: a dictionary with url, caption, id and similarity keys. Querying by image accepts either a local path or a URL, and querying by embedding accepts a precomputed vector. Because deduplicate, use_safety_model and use_violence_detector default to True, a first query against a backend you control may return fewer rows than num_images, and the README does not explain how to tell which filter removed what.

## Where the pipeline breaks down

The clearest limitation is that this is a batch system wearing a search interface. Embeddings are computed for a dataset, an index is built over the result, and the backend serves that fixed index. Nothing in the README describes how to append to an index in place, how to remove an image, or how to keep the embedding store and the index consistent if one step fails midway.

The second limitation is the CPU default. requirements.txt pins faiss-cpu, so the index lives in memory on the machine running clip back. For the Laion5B scale the README points at, the project routes you to a separate document, docs/laion5B_back.md, rather than claiming the default install handles it. The README also does not state memory requirements for the backend at any dataset size, which is the number you most need before provisioning a host.

Finally, the safety and aesthetic filters are opinionated defaults. use_safety_model and use_violence_detector are on by default, and aesthetic_score defaults to 9. If you are building a search tool for a curated archive, those defaults may silently drop exactly the images your users are looking for.

## clip-retrieval compared with embedding search inside a vector database

The obvious alternative is to skip this project and put CLIP embeddings into a general-purpose vector store, then write the HTTP layer yourself. That path gives you incremental inserts, deletes, and a query API you already operate. The difference in approach is real: clip-retrieval owns the whole pipeline including the CLIP model and the FAISS index, so you get a working search UI quickly but you inherit its batch assumptions. A vector database assumes you bring your own embeddings and manage the lifecycle, so you write more code and get a system that tolerates change.

A second alternative is to use all_clip or open_clip directly for embedding computation and stop there. The README lists all_clip as the way to load any CLIP model and open_clip as the way to train them, and clip-retrieval depends on both. If your only need is embeddings for a downstream task, those libraries are the smaller dependency. clip-retrieval earns its place only when you also want the index, the backend and the front end.

## Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-03-28. The most recent release listed is 2.45.0 from 2025-08-15, following 2.44.0 and 2.43.0 in January 2024. The gap between the 2024 releases and 2.45.0 is wide enough that you should read HISTORY.md before assuming a smooth upgrade path, since the README does not describe migration between versions.

The dependency constraints matter more than the version number. requirements.txt pins upper bounds across the stack, including torch below 3, numpy below 2, pyarrow below 16 and faiss-cpu below 2. Those ceilings mean clip-retrieval will hold back packages in a shared environment, and an upgrade to the project can force a coordinated upgrade of torch, torchvision and numpy together.

The licence is MIT, declared in both LICENSE and setup.py. MIT is permissive and places few obligations on how you redistribute the code, but it says nothing about the CLIP checkpoints or datasets you point the pipeline at. The README links to Laion5B and to img2dataset without discussing the terms attached to the images or the model weights, and those terms are your problem to check separately. Nothing here is legal advice.

## Conclusion

Adopt clip-retrieval if you already have a dataset of image URLs and captions and want a self-hosted semantic search endpoint without writing the embedding, indexing and serving layers yourself. Skip it if you need a managed service, or if you want a single library call inside a larger application rather than five CLI stages and a Flask process. Before committing, run the end2end example against the 1000-row parquet file and confirm that the index type autofaiss picks fits your dataset size, since the README does not document what happens when the index and the query volume outgrow a single machine.

## FAQ

### What does clip-retrieval do?

It computes CLIP embeddings for images and text, builds a FAISS index over them, and serves that index through a Flask backend with a browser front end. The README describes the result as a simple semantic search system.

### What is embedding-based retrieval in clip-retrieval?

Queries are converted to embeddings and matched against an index of precomputed embeddings rather than against keywords, so a text query returns images whose embeddings are close to the query embedding. clip-retrieval builds that index with autofaiss and serves it with clip back.

### What does image retrieval mean for clip-retrieval?

In this project, retrieval means sending a text, image or embedding query to a backend and getting back captioned images ranked by similarity. The README's example result carries url, caption, id and similarity fields.

### What is a CLIP model and how does clip-retrieval use one?

A CLIP model maps images and text into a shared embedding space. clip-retrieval uses one to compute embeddings during clip inference, and the README notes that all_clip can load alternative CLIP models while open_clip is used for training.

## Sources

- [License: MIT](https://github.com/rom1504/clip-retrieval/blob/main/LICENSE)
- [Project website](https://rom1504.github.io/clip-retrieval/)
- [README](https://github.com/rom1504/clip-retrieval/blob/main/README.md)
- [Releases](https://github.com/rom1504/clip-retrieval/releases)
- [rom1504/clip-retrieval on GitHub](https://github.com/rom1504/clip-retrieval)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/rom1504-clip-retrieval
