# colpali-cookbooks: notebook recipes for ColPali and ColQwen2 multimodal retrieval

> The repository is a set of six Jupyter notebooks covering inference, fine-tuning, interpretability and an end-to-end RAG pipeline with adapter hot-swapping. It is teaching material pinned to colpali-engine 0.3.x, not a library you import.

**tonywu71/colpali-cookbooks** — Recipes for learning, fine-tuning, and adapting ColPali to your multimodal RAG use cases. 👨🏻‍🍳

- Repository: https://github.com/tonywu71/colpali-cookbooks
- Website: https://huggingface.co/vidore
- Stars: 357 · Forks: 29
- Language: Unknown
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/tonywu71-colpali-cookbooks

## What colpali-cookbooks is for, and who should open it

ColPali retrieves documents by looking at page images rather than extracted text. The README describes it as treating each page as an image and using Paligemma-3B to capture layout, tables and charts, producing multi-vector embeddings compared with pairwise late interaction. That changes what you have to build: no OCR stage, no chunking policy, but a vision-language model in the retrieval path.

The repository is the on-ramp to that. It holds six notebooks under examples, listed from most recent to oldest in the README table, covering inference with the transformers-native ColPali and ColQwen2 implementations, similarity-map generation for both models, LoRA fine-tuning of ColPali, and an end-to-end RAG pipeline for ColQwen2 with adapter hot-swapping. The audience is an engineer or researcher who wants to see the mechanics on a real model before writing production code. It is not aimed at someone who wants a pip-installable retriever, and the packaging confirms that: the wheel target includes only the examples directory.

## How the recipes are organised, and what the pin to colpali-engine 0.3.x means

There is no runtime architecture here. The unit of distribution is the notebook, and pyproject.toml wires the dependencies that the notebooks assume: colpali-engine[interpretability,train]>=0.3.2,<0.4.0, datasets>=3.0.1,<4.0.0, huggingface_hub and ipykernel. The extras tell you which notebooks need what. The interpretability extra backs the similarity-map notebooks, the train extra backs the fine-tuning one, and the Skypilot extra (skypilot==0.6.1,<1.0.0) is for launching runs on a cluster rather than a single machine.

The upper bound on colpali-engine is the detail worth noticing. Anything below 0.4.0 is in range, so a fresh install today can pull a version newer than the one a notebook was written against. That is normal for teaching material that tracks a fast-moving research library, but it means a notebook cell can fail on an API change that has nothing to do with your code. The release history reflects the same moving target: v0.4.0 is titled "transformers-native ColPali cookbook" and v0.5.0 "transformers-native ColQwen2 cookbook", so the transformers-native path is a recent addition rather than the original approach.

Python 3.9 is the declared floor in requires-python. The dev extra adds ruff with a 120-character line length and pytest, but the lint and test configuration applies to a repository whose main artifacts are .ipynb files, so do not read it as a test suite for the recipes themselves.

## Running a notebook locally or in Colab

The README gives two routes. The easiest is to open a notebook from the examples directory and use the Colab button, which runs the code in Google Colab. The local route is to clone the repository and open the notebooks in Jupyter Notebook or an IDE. There is no install command in the README, so the dependency set in pyproject.toml is the closest thing to one. Installing the package itself gets you nothing but the notebooks, because the wheel target includes only examples.

If you want the dependencies without cloning, install the pinned engine and the two supporting packages directly:

```bash
pip install "colpali-engine[interpretability,train]>=0.3.2,<0.4.0" "datasets>=3.0.1,<4.0.0" huggingface_hub ipykernel
```

For a local clone, the README's instruction is to open the notebooks in Jupyter. The project also declares a Skypilot extra for cluster runs, which is installed as a separate group rather than by default:

```bash
pip install "colpali-cookbooks[skypilot]"
```

For a first real use, open examples/use_transformers_native_colpali.ipynb. The README describes it as covering inference, scoring and interpretability with the transformers-native ColPali implementation, so it is the shortest path to seeing a similarity score computed on a page image. Expect a model download from Hugging Face on first run; the README does not state the download size or the VRAM requirement, so treat the first execution as the measurement. If your machine cannot hold the model, the adapter hot-swapping notebook is the one the README says works on Colab's free T4 GPU.

## The adapter hot-swapping recipe and its VRAM trade-off

The most opinionated notebook is the RAG one, titled "ColQwen2: One model for your whole RAG pipeline with adapter hot-swapping". The README says its purpose is to save VRAM by using a unique VLM for the entire RAG pipeline, and that it works on Colab's free T4 GPU. The mechanism implied by the title is that one base model stays resident while task-specific LoRA adapters are swapped in and out, so retrieval and generation share weights instead of loading two models.

The trade-off is real. Adapter swapping buys memory at the cost of serialising work: you cannot retrieve and generate concurrently on the same base model, and every swap has a cost the README does not quantify. For a demo or a single-user pipeline on a small GPU that is a good bargain. For a service with concurrent requests, it is a bottleneck you would have to design around, and the notebook does not claim otherwise. The README also does not document rollback, adapter versioning, or what happens when an adapter fails to load mid-request.

## Fine-tuning ColPali with LoRA and quantized weights

examples/finetune_colpali.ipynb is described as fine-tuning ColPali with LoRA and optional 4bit or 8bit quantization. The quantization options are what make this runnable outside a datacenter: 4bit cuts weight memory substantially, and LoRA keeps the number of trainable parameters small so the optimizer state does not dominate. The train extra in pyproject.toml exists for this notebook.

What the README does not give you is a recipe for the data. examples/data/ exists in the repository layout, but the README does not describe its format, its size, or how to substitute your own pages. It also does not state expected training time, a target metric, or how to evaluate the fine-tuned model against the base one. Fine-tuning a vision-language retriever without a held-out evaluation set is a good way to produce a model that scores well on the pages you trained on and worse on everything else. The ViDoRe benchmark is linked from the README as a separate repository, and that is where evaluation lives, not here.

## Similarity maps as a debugging tool, not a metric

Two notebooks generate similarity maps, one for ColQwen2 and one for ColPali. They visualise which regions of a page image contributed to a retrieval score. This is the part of the repository that is hardest to get elsewhere, because late interaction over multi-vector embeddings is opaque by default: you get a number and no explanation.

Use it as a diagnostic. When a query returns the wrong page, a similarity map tells you whether the model fixated on a chart, a header, or a table that happens to share vocabulary with the query. That is a concrete failure mode for image-based retrieval, and it is the kind of thing that text-extraction pipelines never surface because they never see the layout. The limitation is that the maps are qualitative. The README does not present them as a scoring method, and there is no threshold or aggregate metric attached to them, so they will not substitute for evaluating retrieval quality on a labelled set.

## When to reach for something else

If you need an installable retrieval library rather than notebooks, the ColPali Engine, linked from the README, is the actual package: colpali-engine is what this repository depends on, and this repository only demonstrates it. The difference in approach is that the engine gives you model classes and indexing utilities to embed in an application, while the cookbooks give you executed cells and prose explaining why each step happens. Start here to understand the model, then move to the engine to ship it.

If your documents already extract cleanly to text, a text-based retriever is the wrong comparison and the right choice. ColPali's advantage is layout, tables and charts that OCR degrades; if your corpus is plain prose, you are paying vision-model inference costs for information a text encoder already captures. And if you need a stable API surface across upgrades, the colpali-engine>=0.3.2,<0.4.0 range plus the notebook format means you should expect to re-run cells after an upgrade rather than assume they still work.

## Licence and the cost of keeping notebooks current

The repository is MIT licensed, and pyproject.toml carries the matching classifier. MIT permits commercial use and modification with attribution, but it covers this repository only. The models the notebooks download from Hugging Face have their own licences, and the README does not restate them, so check the model card before you build on a checkpoint. This is a description of what the licence text says, not legal advice.

Maintenance cost is the more practical concern. The last push was on 2026-09-03, and the most recent release is v0.5.0 from 2025-06-02, so the repository is not archived and the notebook set has been extended over time. The upgrade cost sits in the pin: when colpali-engine crosses 0.4.0, the declared range excludes it and the notebooks will need edits before they run against it. Budget for re-executing each notebook you depend on after any engine upgrade, and for the model download on every fresh environment.

## Conclusion

Adopt it if you want to run ColPali or ColQwen2 by hand before committing engineering time: the notebooks cover inference, fine-tuning with LoRA and 4bit or 8bit quantization, similarity maps and an end-to-end RAG example. Do not adopt it if you need an installable retrieval library, because the distribution contains only the examples directory. Verify the colpali-engine version pin, the requires-python floor of 3.9, and whether your GPU or Colab's T4 can hold the model before you plan anything around it.

## FAQ

### What is ColPali in the context of colpali-cookbooks?

The README describes ColPali as a model that retrieves documents by analysing their visual features, treating each page as an image and using Paligemma-3B to capture text, layout, tables and charts as multi-vector embeddings compared with pairwise late interaction. The repository provides notebooks for learning, fine-tuning and adapting it.

### What is the maximum size of a ColPali model?

The README states that ColPali uses Paligemma-3B, which is the only size figure it gives. It does not document a maximum model size or a range of checkpoints, so any other number would be guesswork.

### Do I need to install colpali-cookbooks to use the notebooks?

No. The README's two routes are opening a notebook from the examples directory in Colab, or cloning the repository and opening the notebooks in Jupyter Notebook or an IDE. The wheel target in pyproject.toml includes only the examples directory, so installing the package does not add a runtime library.

### Which notebook should I run first?

The examples/use_transformers_native_colpali.ipynb notebook is listed as covering inference, scoring and interpretability with the transformers-native ColPali implementation, which makes it the shortest path to a computed similarity score. The adapter hot-swapping RAG notebook is the one the README says works on Colab's free T4 GPU.

### Does colpali-cookbooks cover fine-tuning?

Yes. examples/finetune_colpali.ipynb is described as fine-tuning ColPali using LoRA with optional 4bit or 8bit quantization, and the train extra in pyproject.toml backs it. The README does not document the data format under examples/data/ or how to evaluate the result.

## Sources

- [License: MIT](https://github.com/tonywu71/colpali-cookbooks/blob/main/LICENSE)
- [Project website](https://huggingface.co/vidore)
- [README](https://github.com/tonywu71/colpali-cookbooks/blob/main/README.md)
- [Releases](https://github.com/tonywu71/colpali-cookbooks/releases)
- [tonywu71/colpali-cookbooks on GitHub](https://github.com/tonywu71/colpali-cookbooks)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/tonywu71-colpali-cookbooks
