llm2vec: the transformer pin stops at 4.44.2, and the updates list has four date formats
Code for 'LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders'
At a glance
- What is it?
- A three-step recipe that turns decoder-only language models into text encoders, delivered as LoRA adapters on masked-next-token checkpoints. The package metadata caps transformers at a narrow window, flash-attn is a manual second install, and the newest work in this line moved to a different repository with a different task.
- Who is it for?
- This is worth using if you want a text encoder that starts from a decoder checkpoint you already hold, and you are prepared to install a compiled attention kernel yourself. Two things decide whether it works in your environment.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The metadata caps transformers at 4.44.2 while the changelog claims newer support
The declared dependency is a closed window rather than a floor: `transformers>=4.43.1,<=4.44.2`. The other runtime requirements are all open, with numpy, tqdm, torch, peft, datasets, evaluate and scikit-learn carrying no version at all, and mteb is held at >=1.14.12 behind an extra named evaluation. That ceiling is hard to reconcile with an entry in the updates list, which says support was added for the latest transformer versions, the ones that handle Llama 3.1 and 3.2 and other recent models, and points at a custom evaluation script for evaluating any LLM2Vec model. The intent and the constraint point in opposite directions, so a fresh install into a current environment resolves to the source checkout rather than to the published package.
flash-attn is a manual second install, and it must build without isolation
The library has two installation routes and both of them end the same way:
pip install llm2vec
pip install flash-attn --no-build-isolationCloning the repository and running `pip install -e .` is followed by the identical second line. Two details are worth pausing on. flash-attn does not appear anywhere in install_requires, so it is an instruction rather than a dependency, and nothing in the package will fail loudly if you skip it. The flag matters as much as the package: `--no-build-isolation` tells pip to build against the environment that already exists, which is what makes the compiled kernel match the installed torch instead of a throwaway build environment. That also means the build depends on a compiler toolchain being present before you start, rather than being fetched for you.
The version number is executed out of a file during setup
setup.py does not carry a version literal. It opens llm2vec/version.py, reads the text, and runs it through exec into a dictionary, then takes __version__ from the result. The rest of the file is conventional: packages are found with a pattern limited to the library name, the long description is the README itself, include_package_data is on, and the package is marked as not zip safe. The declared floor is python_requires >=3.8, which sits well below the versions most current decoder models assume, and the classifiers list only that the package works on Python 3 and is operating system independent, without naming a single minor version. So the one number the installer checks is the loosest in the file, while the number it reads for its own identity is the one computed at build time.
Twelve published checkpoints, and only two of them carry a claim
The model list is a grid of four base models, Meta-Llama-3-8B, Mistral-7B, Llama-2-7B and Sheared-Llama-1.3B, against three training stages: bidirectional attention with masked next token prediction, the same plus unsupervised SimCSE, and the same plus supervised contrastive training on public E5 data. That is twelve combinations, and two of them carry footnote markers. One is marked as state of the art on MTEB among models trained on public data, and that marker sits on the supervised Meta-Llama-3 cell. The other is marked as unsupervised state of the art on MTEB, on the Mistral unsup-simcse cell. Read carefully, the leaderboard claim covers two of twelve cells, and the first of those is explicitly limited to the public-data category.
The input is a list of pairs, and the instruction convention changes with the task
The encode call does not take plain strings. It takes a list whose entries are either a two-item list of instruction and text, or a bare string, so a batch can mix instructions and plain text. The stated convention is that instructions are provided for both sentences in symmetric tasks, and only for the query in asymmetric tasks, and the example that follows carries a web search instruction on the query side while noting in a comment that instructions are not required for documents. That convention is part of the interface rather than a convenience: a retrieval query encoded with the instruction and a document encoded without it are the intended pair, and encoding both sides the same way is a different input to the same model. The example text that follows is a CDC protein passage used with the query about how much protein a female should eat.
Bidirectionality is on by default, pooling is mean, and the window is 512
Three defaults define the behaviour, and all three are arguments to from_pretrained rather than configuration files. Models load with bidirectional connections enabled, and the change is a single argument, enable_bidirectional=False, which turns the defining step of the whole recipe back off. Pooling defaults to the mean strategy and is changed with pooling_mode. The maximum sequence length defaults to 512 and is changed with max_length. The loading call also takes any HuggingFace model loading argument, and the examples pass a base model identifier, an optional PEFT adapter path, a device map that selects cuda when available and cpu otherwise, and bfloat16 as the dtype. So the LoRA weights are the variable part, and everything else about the encoder is set at construction time.
The updates list uses four date formats and is not in order
The news section opens with an entry reading 04/04/26, the only one carrying a year, and the three entries below it read 03/10, 04/07 and 30/04 with no year at all. Which century and which half of the year those three belong to is left to the reader, and 04/07 in particular reads as either April or July depending on the convention. The order does not resolve it either, since the dated entry sits above three undated ones. The content of the entries is the more useful signal: a new generative embedding recipe, support for the newest transformer versions, support for Gemma and Qwen-2 contributed from outside the group, and Meta-Llama-3 checkpoints in supervised and unsupervised variants. The newest release is 0.2.3 from 2025, while the last recorded push to main is dated 2026-04-04.
The newest work moved out of this repository and changed the task
The top entry announces LLM2Vec-Gen, described as a recipe for interpretable, generative embeddings that encode the potential answer of an LLM to a query rather than the query itself. Three links accompany it, and all three leave this repository: a paper, a separate GitHub repository, and a Hugging Face model built on Qwen3-8B. The direction of travel is the opposite of the rest of the project. LLM2Vec takes a decoder model and turns it into something that embeds text for retrieval, keeping the query as the thing encoded. LLM2Vec-Gen is described as encoding the answer instead, which is a different representation problem and a different evaluation, and it is maintained outside this tree. What remains here is the retrieval recipe, its twelve checkpoints and the training scripts under experiments/, with configuration directories for training and testing beside them.
Editorial conclusion
This is worth using if you want a text encoder that starts from a decoder checkpoint you already hold, and you are prepared to install a compiled attention kernel yourself. Two things decide whether it works in your environment. Check the transformers ceiling first, because the package metadata stops at 4.44.2 while the changelog claims support for the versions that arrived with Llama 3.1 and 3.2, so a modern stack needs a source install or a patched constraint. Second, budget for the build: flash-attn is not a declared dependency and must be compiled with build isolation turned off, which ties it to whatever torch you already have. For evaluation, look at the two starred cells rather than the whole table, since only those two carry a state-of-the-art claim and one of them is limited to models trained on public data. And if you want generative rather than retrieval embeddings, that work now lives in a separate repository.
Frequently asked questions
What are the three steps in the llm2vec recipe?
Enabling bidirectional attention on a decoder-only model, training with masked next token prediction, and then unsupervised contrastive learning. The model can be fine-tuned further, and the published checkpoints deliver the result as a base MNTP model plus a LoRA adapter, in supervised or unsupervised contrastive variants.
How do I install llm2vec, and why is flash-attn separate?
Run pip install llm2vec and then pip install flash-attn --no-build-isolation, or clone the repository and run pip install -e . followed by the same second command. flash-attn is not in install_requires, and the no-build-isolation flag makes the kernel build against the torch already installed in your environment.
Which transformers version does llm2vec support?
The package metadata declares a closed window, transformers>=4.43.1,<=4.44.2. The updates list separately claims support for the newest transformer versions, the ones covering Llama 3.1 and 3.2, which is why a current environment is better served by installing from the repository.
What pooling and sequence length does llm2vec use by default?
Mean pooling, a maximum sequence length of 512, and bidirectional connections enabled. Each of the three is an argument to the from_pretrained method, changed with pooling_mode, max_length and enable_bidirectional respectively.
Where did the generative embedding work from llm2vec go?
Into a separate project announced from this repository. LLM2Vec-Gen trains interpretable generative embeddings that encode the potential answer to a query rather than the query, and it lives in its own repository with its own paper and a Qwen3-8B model on Hugging Face, so it is not part of this package.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mcgill-nlp-llm2vec)