PyLate: training and retrieval for ColBERT late interaction models
Late Interaction Models Training & Retrieval
At a glance
- What is it?
- PyLate is a Python library built on Sentence Transformers for fine-tuning, inference and retrieval with ColBERT models, up to multi-GPU training and knowledge distillation. Here is how it installs, what its losses actually do, and where it stops being the right tool.
- Who is it for?
- Adopt PyLate if you already have a Sentence Transformers pipeline and want ColBERT scoring inside it, or if you need knowledge distillation from a teacher model and multi-GPU training without writing the plumbing yourself. Do not adopt it if you need a retrieval server with a documented rollback story, or if you are not prepared to pin transformers, sentence-transformers and fast-plaid together.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 69 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What PyLate adds on top of Sentence Transformers
Sentence Transformers gives you a trainer, a loss interface and a model abstraction for single-vector embeddings. ColBERT does not fit that shape. It keeps one vector per token, scores a query against a document with a MaxSim operation, and needs its own collator, its own evaluator and its own index format. PyLate is the layer that fills those gaps without asking you to leave the Sentence Transformers training loop.
The README describes it as "a library built on top of Sentence Transformers, designed to simplify and optimize fine-tuning, inference, and retrieval with state-of-the-art ColBERT models." The audience is narrow and specific: people who already have a dataset of query, positive and negative triples, already know what a SentenceTransformerTrainer is, and want late interaction scoring rather than a single pooled vector. If you have never trained a retrieval model, the library assumes a fair amount of context.
One detail worth noticing early: models.ColBERT will accept a base encoder that was never trained for retrieval. The README says that if the checkpoint is not a ColBERT model, a linear layer is added to the base encoder. That means you can start from bert-base-uncased and build a ColBERT model, which is unusual flexibility and also a warning that the initial output is not meaningful until you train it.
The training loop: contrastive loss, GradCache and distillation
Training in PyLate is the Sentence Transformers loop with PyLate objects substituted in. You build a models.ColBERT, wrap it in torch.compile, load a triplet dataset, pick a loss, attach an evaluator and hand everything to SentenceTransformerTrainer.
The contrastive path has two constraints the README states plainly. First, temperature matters a lot in contrastive learning, and the README points to a CVPR 2021 paper on the behaviour of contrastive loss while noting that a temperature around 0.02 is often used in the literature. Second, contrastive learning is not compatible with gradient accumulation. The workaround the library offers is GradCache: losses.CachedContrastive takes a mini_batch_size while you raise per_device_train_batch_size, emulating a larger batch without the memory cost. In multi-GPU setups, both Contrastive and CachedContrastive accept gather_across_devices=True, which pools elements across devices for an even larger effective batch.
The distillation path is the one the README recommends for best performance. Rather than learning only from positive and negative pairs, you train against the scores of a strong teacher model. The example loads lightonai/ms-marco-en-bge in three configurations (train, queries, documents) and wires them together with utils.KDProcessing, which resolves query and document ids to text on the fly. That transform-based join is the part worth understanding: the training set stores ids, not strings, so the dataset stays small on disk and the text is fetched during collation.
Installing PyLate and running a first training job
The README gives one installation command for the base library and a separate extra for evaluation dependencies. The base install pulls in sentence-transformers 5.3.0, datasets, accelerate, transformers, ujson, ninja, fastkmeans and fast-plaid, so expect a heavy environment.
pip install pylateIf you plan to run the evaluation helpers, the README specifies an extra:
pip install "pylate[eval]"The pyproject.toml declares several more extras, including api (fastapi, uvicorn, batched), voyager, scann, warp, tachiom, flash-maxsim and lik. Those are optional index and kernel backends rather than part of a default install.
A minimal first run, adapted from the README's contrastive example, is a ColBERT model over a triplet dataset with a ColBERT-specific collator. The collator is the piece people forget; without utils.ColBERTCollator(model.tokenize) the trainer will not produce the token-level batches the loss expects.
from datasets import load_dataset
from sentence_transformers import SentenceTransformerTrainer, SentenceTransformerTrainingArguments
from pylate import evaluation, losses, models, utils
model = models.ColBERT(model_name_or_path="bert-base-uncased")
dataset = load_dataset("sentence-transformers/msmarco-bm25", "triplet", split="train")
splits = dataset.train_test_split(test_size=0.01)
train_loss = losses.Contrastive(model=model, temperature=0.02)
args = SentenceTransformerTrainingArguments(
output_dir="output/contrastive-bert-base-uncased",
num_train_epochs=1,
per_device_train_batch_size=32,
fp16=True,
)
trainer = SentenceTransformerTrainer(
model=model, args=args, train_dataset=splits["train"],
eval_dataset=splits["test"], loss=train_loss,
data_collator=utils.ColBERTCollator(model.tokenize),
)
trainer.train()The README notes that fp16 should be set to False if your GPU cannot run it, and that bf16 is the alternative on hardware that supports it. After training, the README shows the output directory being passed straight back into models.ColBERT as model_name_or_path, so the saved checkpoint is reloadable without conversion.
Where PyLate is the wrong tool
The dependency pins are the first real constraint. pyproject.toml requires sentence-transformers exactly at 5.3.0, transformers in the range >=4.41.0,<=5.3.0, ujson exactly at 5.12.0, ninja exactly at 1.11.1.4, fastkmeans exactly at 0.5.0, and fast-plaid in a narrow band. If your project already pins a different sentence-transformers version, installing PyLate will either fail to resolve or force an upgrade of something else you depend on. This is not a library you drop into an existing stack casually.
Training data shape is the second constraint. The contrastive example expects triplets with query, positive and negative columns, and the distillation example expects a dataset keyed by ids that a transform resolves later. If your corpus is unlabelled, PyLate does not generate the supervision for you; you need a teacher model to score pairs, or existing relevance judgements.
Indexing is the third. The library depends on fast-plaid for retrieval, and the optional extras list voyager, scann and tachiom as alternative backends, each with its own Python version restrictions (voyager and scann are both marked python_version < '3.14'). The README does not document what happens when an index is corrupted or how to migrate an index between fast-plaid versions. If you need an operational story around index rebuilds, that story is not in the README, and you will be writing it yourself.
Finally, PyLate is a training and retrieval library, not a search service. The api extra pulls in fastapi, uvicorn and batched, which suggests a serving path exists, but the README shown here does not walk through deploying it. Treat the API extra as a starting point, not a supported product.
PyLate versus a single-vector Sentence Transformers pipeline
The closest alternative is the plain Sentence Transformers stack you already have, using a bi-encoder with a pooled embedding and cosine similarity. The difference is architectural, not cosmetic. A bi-encoder compresses a document into one vector, so retrieval is a single nearest-neighbour lookup and the index is small. ColBERT keeps a vector per token and scores with MaxSim, which is more expensive at query time but preserves token-level matching that pooling destroys. PyLate exists because that second approach needs a different trainer, a different collator and a different index, and Sentence Transformers alone does not supply them.
Within the ColBERT world, the practical choice is between PyLate and the original ColBERT repository. PyLate's bet is integration: it reuses SentenceTransformerTrainer, SentenceTransformerTrainingArguments and the datasets library, so a team already fluent in that API pays a small learning cost. The original ColBERT codebase is a research codebase with its own training entry points. If you need the Sentence Transformers ecosystem, PyLate is the shorter path. If you need to reproduce a specific published checkpoint exactly, the original implementation is the reference.
There is also a middle option worth naming: a reranker. If your corpus is small enough that a cross-encoder can rescore the top hundred candidates, you get much of the quality benefit of late interaction without maintaining a token-level index. PyLate is for when the corpus is too large for that.
Maintenance, releases and licence
The repository is not archived. The last push was on 2026-07-23, which is recent. Releases are frequent enough to suggest ongoing work: v1.4.0 on 2026-02-25, v1.5.0 on 2026-05-05, and v1.6.0 on 2026-06-11. The version is dynamic, read from pylate/__version__.py via setuptools_scm, and the Makefile shows the release procedure: make release VERSION=X.Y.Z opens a branch and a pull request, and make tag VERSION=X.Y.Z tags main after merge and triggers the PyPI publish workflow. That is a documented, reproducible release path rather than ad hoc uploads.
Upgrade cost is dominated by the dependency pins rather than by PyLate's own API. Because sentence-transformers is pinned exactly and transformers has both a floor and a ceiling, a PyLate upgrade can force a coordinated bump across your environment. The optional extras add another axis: the voyager and scann extras are both restricted to Python below 3.14, so a Python upgrade can silently remove an index backend from your options.
The licence is MIT, declared both in pyproject.toml and in the README badge. MIT is permissive, which means you can use PyLate in commercial work and modify it, subject to the usual attribution conditions. That is a statement about the licence text, not legal advice; if you are redistributing a modified version, read the actual LICENSE file in the repository root.
For test and lint practice, the Makefile runs pytest across both pylate and tests with -n auto for parallel execution, and lint goes through pre-commit run --all-files. If you fork the project, those two targets are the entry points.
Editorial conclusion
Adopt PyLate if you already have a Sentence Transformers pipeline and want ColBERT scoring inside it, or if you need knowledge distillation from a teacher model and multi-GPU training without writing the plumbing yourself. Do not adopt it if you need a retrieval server with a documented rollback story, or if you are not prepared to pin transformers, sentence-transformers and fast-plaid together. Before committing, verify that the pinned dependency set in pyproject.toml resolves against your existing environment, and check that the fast-plaid version range (>=1.4.6.270,<=1.4.6.2110) matches the index backend you plan to run.
Frequently asked questions
What is PyLate and what problem does it solve?
PyLate is a Python library built on top of Sentence Transformers for fine-tuning, inference and retrieval with ColBERT late interaction models. It supplies the ColBERT-specific pieces that Sentence Transformers does not: a model class that can add a linear layer to a base encoder, a collator, losses, and evaluators.
How do I install PyLate?
The README gives pip install pylate for the base library, and pip install "pylate[eval]" when you also need the evaluation dependencies. The pyproject.toml lists further extras such as api, voyager, scann, warp, tachiom, flash-maxsim and lik for optional backends.
Can PyLate train a ColBERT model from a plain BERT checkpoint?
Yes. The README states that if the model you load is not a ColBERT model, a linear layer is added to the base encoder, and its example starts from bert-base-uncased. The resulting model is not useful until it has been trained.
Why does PyLate recommend knowledge distillation over contrastive training?
The README says that to get the best performance when training a ColBERT model you should use knowledge distillation, training the model against the scores of a strong teacher model. Its distillation example uses the lightonai/ms-marco-en-bge dataset with a KDProcessing transform that resolves query and document ids to text on the fly.
Does PyLate support multi-GPU training?
The README describes fine-tuning on both single and multiple GPUs, and notes that in a multi-GPU setting you can gather elements from different GPUs to build larger batches by setting gather_across_devices to True on the Contrastive and CachedContrastive losses.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lightonai-pylate)