Model or dataset
illuin-tech/colpali avatar
illuin-tech/colpali

ColPali: the idea that outlived its own package

The code used to train and run inference with the ColVision models, e.g. ColPali, ColQwen2, and ColSmol.

2,817 stars265 forksPythonMIT

At a glance

What is it?
The visual document retriever that replaced OCR pipelines with vision language model patches, and the repository that now tells you to migrate off it and into Sentence Transformers.
Who is it for?
ColPali matters less as a package you install than as an idea that turned out to be right. Encoding a page as a set of visual token vectors and scoring a query against all of them removes the OCR and layout pipeline that most document retrieval systems still depend on, and that observation has been adopted well beyond this repository.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 16 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The core idea: a page as a bag of visual vectors

Traditional document retrieval converts a page to text and searches text. That pipeline is old, well understood, and fragile in a specific way: a table loses its structure, a chart becomes an empty string, and a layout carries meaning that the text extraction discards. ColPali's proposal is to skip that pipeline and let a vision language model represent the page directly.

The mechanism is described precisely in the README. Instead of a single vector per page, the model constructs efficient multi-vector embeddings in visual space. The ViT output patches from PaliGemma-3B are fed to a linear projection, producing a multi-vector representation of the document. Training then maximises similarity between those document embeddings and the query embeddings, following the ColBERT method rather than pooling everything into one vector.

Keeping multiple vectors per page is what makes it work. A single pooled vector has to compress an entire page into a fixed number of dimensions, and any specific detail on that page competes with every other detail for the same space. With one vector per patch, a query about a y-axis label can match the vector for that patch directly, and a chart's visual shape stays available as a signal rather than being reduced to whatever words an extractor found near it.

The stated benefit is the removal of what the README calls potentially complex and brittle layout recognition and OCR pipelines, replaced by a single model that accounts for both the textual and the visual content of a document, layout and charts included. The paper behind it is arXiv 2407.01449.

The repository now leads with a deprecation notice

This is the single most important thing to know before reading anything else here. The README puts a deprecation banner above the original documentation, and the position is unambiguous: `colpali-engine` is deprecated, and Sentence Transformers is recommended for new projects and production use. The repository and package remain available for research, reproducibility and existing projects.

The reason is consolidation rather than failure. Sentence Transformers v6 added first class support for ColPali style models through `MultiVectorEncoder`, giving one ecosystem for embedding and retrieval models across dense bi-encoders and multi-vector late-interaction models, and across text, images, audio and video depending on the model. The README credits the ColPali and ViDoRe teams with helping bring the model configurations to the Hugging Face Hub as part of that integration.

So the project did not lose its argument, it won it. The mechanism that was novel in 2024 is now a supported model class inside a library that most retrieval teams already depend on. That is a good outcome for the field and an awkward one for anyone who wants to keep the original API.

The README closes the original documentation with a thank you and says the journey is far from over and continues at Sentence Transformers. It is a polite handover rather than an obituary, and the preserved documentation below it is explicitly there for reproducibility.

The migration table is the most useful page in the file

Anyone with existing `colpali-engine` code has one question, and the README answers it in a single table with a left column of old calls and a right column of replacements.

Model construction moves from `ColQwen2.from_pretrained(...)` with a separate `ColQwen2Processor` to a single `MultiVectorEncoder` constructed from a model id. Query processing moves from `processor.process_queries(...)` followed by a model call to `model.encode_query(queries)`, and image processing moves the same way to `model.encode_document(images)`. Scoring moves from `processor.score_multi_vector(...)` to `model.similarity(...)`. Pooling moves from `HierarchicalTokenPooler` to `HierarchicalTokenPooling`, and the interpretability helpers move from `colpali_engine.interpretability` to `sentence_transformers.multi_vector_encoder.interpretability`.

The one entry with a real behavioural subtlety is `mask_non_image_embeddings=True`, which becomes an explicit `MultiVectorMask` configured with `keep_only_token_ids`. Masking used to be a flag and is now an object you construct, which is the kind of change that compiles fine and then produces different results if you skip it.

The README also explains checkpoint recognition: Sentence Transformers automatically recognises Transformers native `*ForRetrieval` checkpoints, including the `-hf` variants of the Vidore models. An example given is `vidore/colqwen2-v1.0-hf`, which keeps its projection and normalization inside the model while its processor formats text queries and image documents. For most inference migrations the whole thing reduces to four lines.

python
from sentence_transformers import MultiVectorEncoder

model = MultiVectorEncoder("vidore/colqwen2-v1.0")
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(images)
scores = model.similarity(query_embeddings, document_embeddings)

Installation is one command, with image support requested explicitly: `pip install -U "sentence-transformers[image]>=6.0.0"`. For exact parity notes, training guidance and advanced configuration, the README defers to the full Sentence Transformers migration guide.

The package itself, for people who stay

If you are keeping `colpali-engine`, the packaging tells you what environment you are committing to. The project metadata names three authors, Manuel Faysse, Hugues Sibille and Tony Wu, with Faysse and Wu also listed as maintainers, and requires Python 3.10 or later with an upper bound of 3.15.

The core dependencies are `numpy`, `peft` with a range up to but not including 0.21.0, `pillow`, `requests`, `scipy`, `torch` with a range up to 2.14.0, `torchvision`, and `transformers` from 5.3.0 up to but not including 6.0.0. The Transformers floor is worth noting: this project now requires a Transformers major version that did not exist when ColPali was published, which is what the v0.3.18 fix is about.

The optional extras are organised by task rather than lumped together. `train` pulls in accelerate, bitsandbytes for quantisation, configue for config management, datasets, mteb for evaluation, pillow with a tighter upper bound, and typer for a command line interface. `interpretability` adds einops, matplotlib and seaborn for visualising the attention maps that make this model family explainable. `lik` installs late-interaction-kernels, which arrived in v0.3.17 as a faster path for the MaxSim operation. `plaid`, added in v0.3.18, installs fast-plaid and fastkmeans for faster matching on larger corpora. `dev` is pytest and ruff, and `all` pulls the lot.

The build backend is hatchling with hatch-vcs, meaning the version is derived from version control rather than hardcoded, and the wheel includes only the `colpali_engine` package. The repository tree also carries `docs/`, `examples/`, `scripts/`, `tests/`, a `CHANGELOG.md` and a `CITATION.cff` for academic citation, plus an `assets/` directory holding the architecture diagram.

Release history tells you what actually breaks

Three recent releases are visible, and each is a lesson in what goes wrong when a research repository meets a changing dependency stack.

v0.3.18, published in August 2026, is the one that matters most if you load a ColQwen2 checkpoint. It fixes checkpoint conversion so that `model.embed_tokens` and `model.norm` are remapped into the `language_model.*` layout used by Transformers v5. The stated consequence of not having this fix is that those weights were silently dropped and randomly re-initialised when loading from Qwen2-VL checkpoints. A model that loads without error and produces plausible garbage is the worst failure mode in this list, and if you are debugging unexpectedly poor retrieval quality from a ColQwen2 checkpoint, that is where to look. The same release widened the dependency ranges to allow `torch<2.14.0`, `peft<0.21.0` and `pillow<12.4.0`.

v0.3.17, from June 2026, added support for late-interaction-kernels. The practical significance is that MaxSim, the operation that compares every query vector against every document vector, is quadratic in the number of vectors, and it is the bottleneck once you index more than a few thousand pages. A kernel implementation moves that work to something that can use a GPU properly.

v0.3.16, from May 2026, was purely a dependency range extension.

So the pattern is clear. Model architecture work happened years ago and has stopped. What remains is compatibility maintenance against a moving PyTorch and Transformers stack, plus performance work on the scoring path. That is exactly what you would expect from a project whose main library has moved elsewhere.

Model family and where to evaluate

The repository describes itself as the code used to train and run inference with the ColVision models, naming ColPali, ColQwen2 and ColSmol as examples. The original ColPali model is based on the ColBERT architecture and the PaliGemma model, and the README notes it includes later ColVision and bi-encoder retriever variants.

That distinction matters. The multi-vector late interaction models are what this repository is about. The bi-encoder variants mentioned alongside them are a different retrieval shape that produces a single vector per page, which is the faster and less precise option. A team choosing between them should decide based on whether queries tend to target specific details inside a page, which is where multi-vector scoring pays off, or whether coarse page-level relevance is enough, which is where bi-encoders are hard to beat.

The README closes with the beginning of a model table listing each model's score on the ViDoRe leaderboard, its licence, comments and whether it is currently supported, though the collected text stops before the rows. ViDoRe is the evaluation to look at rather than any single number, since it is built specifically for this retrieval setting.

The evaluation ecosystem lives in sibling repositories rather than here: the ViDoRe benchmark and its Hugging Face Space leaderboard, a Hugging Face Space demo, a blog post explaining the approach, and a cookbooks repository under a different maintainer. The model card, the paper and the leaderboard are linked from the README header, which is where to go for scores rather than for code.

Editorial conclusion

ColPali matters less as a package you install than as an idea that turned out to be right. Encoding a page as a set of visual token vectors and scoring a query against all of them removes the OCR and layout pipeline that most document retrieval systems still depend on, and that observation has been adopted well beyond this repository. What you find on `main` today is a project in the middle of a handoff. The README leads with a deprecation notice, points new work at Sentence Transformers v6 and `MultiVectorEncoder`, and provides a side by side table mapping every engine call to its replacement. Existing code still works and the package still installs; treat the maintenance line as frozen and plan against the migration guide instead. If you are training rather than serving, read the training extra and the ColQwen2 checkpoint fix in v0.3.18 before assuming an old checkpoint loads cleanly under Transformers v5.

Frequently asked questions

What does ColPali do?

It builds document embeddings in visual space for retrieval, skipping OCR. Instead of one vector per page, the ViT output patches from PaliGemma-3B are fed to a linear projection to produce a multi-vector representation, and the model is trained to maximise similarity between document and query embeddings following the ColBERT method. The stated benefit is removing layout recognition and OCR pipelines, since one model accounts for both the text and the visual content of a page, including layout and charts.

Is ColPali a VLM?

It is built on one. The original model is based on the ColBERT architecture and the PaliGemma vision language model, and the repository covers later ColVision and bi-encoder retriever variants. The distinction that matters in practice is that the multi-vector late interaction models keep one vector per patch, which is what lets a query match a specific detail on a page, while the bi-encoder variants produce a single vector per page and are faster but coarser.

What are the key differences between ColQwen2 and ColPali?

They share the multi-vector late interaction design and differ in the underlying vision language model. ColPali is built on PaliGemma, while ColQwen2 is built on the Qwen2-VL family, which the package metadata reflects: ColQwen2 needs `ColQwen2.from_pretrained(...)` plus a `ColQwen2Processor` under `colpali-engine`. One practical difference shows up in checkpoint loading. The v0.3.18 release fixed ColQwen2 conversion so `model.embed_tokens` and `model.norm` map into the `language_model.*` layout used by Transformers v5, because without that fix those weights were silently dropped and randomly re-initialised.

Is ColPali open source?

Yes, under the MIT licence, and the package on PyPI is named `colpali_engine`. The catch is that it is deprecated. The README recommends Sentence Transformers for new projects and production use, since version 6 added `MultiVectorEncoder` for ColPali style models, and ships a migration table mapping `ColQwen2.from_pretrained(...)` to `MultiVectorEncoder(...)`, `process_queries` and `process_images` to `encode_query` and `encode_document`, `score_multi_vector` to `similarity`, and `HierarchicalTokenPooler` to `HierarchicalTokenPooling`. The package remains available for research, reproducibility and existing projects.

Official sources

  1. illuin-tech/colpali on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/illuin-tech-colpali.svg)](https://hysenlabs.com/projects/illuin-tech-colpali)