Salesforce LAVIS: A Unified Library for Language-Vision Models
LAVIS - A One-stop Library for Language-Vision Intelligence
At a glance
- What is it?
- LAVIS bundles BLIP, BLIP-2, InstructBLIP and their datasets behind one configuration-driven interface. It suits researchers who need reproducible vision-language baselines, not teams that want a small dependency footprint.
- Who is it for?
- Adopt LAVIS if you are reproducing or extending BLIP, BLIP-2, InstructBLIP or the other models listed in the README, and you can work inside its pinned dependency set. Do not adopt it if you need a current transformers release in the same environment, or if you only want an inference endpoint for one captioning model.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- No. The owners have archived the repository on GitHub, so it is read-only and no longer receives changes.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What LAVIS Actually Solves for Vision-Language Work
Vision-language research has a packaging problem. A captioning model, a retrieval model and a visual question answering model usually arrive as three repositories with three dataset loaders and three training loops. LAVIS is Salesforce's attempt to collapse that into one library. The technical report describes it as "a one-stop comprehensive library" with "a unified interface to easily access state-of-the-art image-language, video-language models and common datasets".
The intended user is a researcher or an engineer who needs to train, evaluate or benchmark on multimodal classification, retrieval, captioning, visual question answering, dialogue or pretraining. The README lists those tasks explicitly. If your job is to compare BLIP against BLIP-2 on the same split with the same preprocessing, LAVIS is built for exactly that, because evaluate.py and train.py sit at the repository root and take configuration files rather than ad hoc scripts.
The library is not aimed at someone who wants one model behind an HTTP endpoint. There is an app/ directory and the requirements include streamlit, so demos exist, but the center of gravity is the training and evaluation pipeline. That distinction matters when you are choosing between LAVIS and a narrower inference wrapper.
How the Configuration and Registry Layer Fits Together
The repository layout tells most of the story. lavis/ holds the library code, projects/ holds per-model implementations such as blip2, instructblip, blip-diffusion, xinstructblip, pnp-vqa and img2llm-vqa, run_scripts/ holds launch scripts, and dataset_card/ documents the datasets. The top level exposes train.py and evaluate.py as the two entry points.
Configuration is handled through omegaconf, which appears in requirements.txt. That is why the examples are notebooks rather than single Python files: a task is described by a configuration tree, and the model, dataset and preprocessing are selected by name inside it. The README calls the library "highly extensible and configurable, facilitating future development and customization", which in practice means you add a model by registering it and pointing a config at it, not by editing the training loop.
Dependencies are pinned hard. requirements.txt fixes transformers==4.33.2, timm==0.4.12, fairscale==0.4.4 and opencv-python-headless==4.5.5.64, and caps diffusers at 0.16.0. Those pins are the mechanism that keeps the released checkpoints loading correctly. They are also the main source of friction, because any other package in your environment that wants a newer transformers will conflict.
Installing salesforce-lavis and Running a First Captioning Task
The README states that LAVIS is available on PyPI, and setup.py declares the distribution name as salesforce-lavis. The package itself imports as lavis. A plain pip install is the documented route.
pip install salesforce-lavisBecause the pinned requirements are strict, install into a fresh virtual environment rather than into an existing project environment. setup.py declares python_requires=">=3.7.0" and requirements.txt asks for torch>=1.10.0, so confirm both before you start; on Windows, setup.py adds the PyTorch stable wheel index as a dependency link, which the README does not otherwise explain.
The fastest way to see the library work is the notebook set under examples/. For captioning, the repository ships examples/blip_image_captioning.ipynb and examples/blip2_instructed_generation.ipynb. Open one and run the cells in order; the notebooks import from lavis.models and load a pretrained checkpoint by name. The README does not document a command-line captioning entry point, so treat the notebooks as the first real use rather than looking for a CLI flag that is not there.
For a batch job rather than a notebook, the repository provides run_scripts/ and the root-level evaluate.py. The README does not spell out the exact argument names for evaluate.py, so read the script before wiring it into a pipeline. Do not guess flags; the file is the source of truth.
Where LAVIS Gets in Your Way
The dependency pins are the first real limitation. transformers==4.33.2 is fixed in requirements.txt, and timm is held at 0.4.12. If your application already depends on a newer transformers, installing salesforce-lavis into that same environment will either fail or force a downgrade that breaks the other package. There is no documented extra or optional-dependency group that relaxes these pins.
The second limitation is release cadence. The most recent tagged release in the repository is v1.0.2 from 2023-03-06, while setup.py still declares version="1.0.1". Model implementations have continued to land on the main branch, including X-InstructBLIP in November 2023, but the version metadata and the tags do not track those additions. If you need a reproducible version number for a paper or a lockfile, the tag and the code you are actually running may not correspond.
The third is scope. LAVIS covers the models its authors implemented. The README does not claim compatibility with arbitrary Hugging Face vision-language checkpoints, and the configuration layer is built around LAVIS model registries. If your model is not one of the projects listed, you are writing an integration, not using one.
Finally, the repository is not archived, but the last push was on 2026-06-02. That is the fact to weigh, not a general impression of activity.
LAVIS Against a General-Purpose Transformers Pipeline
The obvious alternative is to skip LAVIS and use transformers directly with a pipeline object for image-to-text or visual question answering. The difference is architectural, not cosmetic.
A transformers pipeline gives you one model, one processor and a callable. You get a current library version, you can upgrade independently, and you can mix the model with anything else in the ecosystem. What you do not get is a shared evaluation path. If you want to score BLIP and BLIP-2 on the same retrieval split with identical preprocessing, you write that harness yourself.
LAVIS inverts the trade. The evaluation and training harness is the product. evaluate.py, train.py, the run_scripts/ directory and the dataset_card/ documentation are there so that a comparison between two registered models is a configuration change. You pay for it with the pins and with the registry abstraction, which is more indirection than a single inference call needs.
A reasonable split: use transformers for serving one model in production, and use LAVIS when the question you are answering is comparative or when you need the training loop that produced the checkpoint.
Maintenance, Licensing and What Upgrades Cost
The licence is BSD-3-Clause. LICENSE.txt is at the repository root, setup.py declares license="3-Clause BSD", and the README links to the Open Source Initiative page for the same identifier. That is a permissive licence, which generally means you can use the code in commercial products provided the copyright notice and disclaimer are retained. The model checkpoints are separate artifacts with their own terms, and the repository does not consolidate those terms in one place, so check the individual model pages before shipping anything. This is a description of what the repository states, not legal advice.
Upgrade cost is dominated by the pins. Moving to a newer transformers means either waiting for the project to update requirements.txt or maintaining a fork, and the README does not document a supported upgrade path or a compatibility matrix. The gap between the v1.0.2 tag from 2023-03-06 and the model work merged afterward suggests version tags are not the unit of upgrade here; commits on main are.
The last push to the repository was on 2026-06-02. Treat that as the signal for how much you should expect the pins to move on their own.
Verifying LAVIS Fits Before You Commit
Three checks resolve most of the uncertainty. First, open the examples/ directory and confirm a notebook exists for your task. The repository ships notebooks for ALBEF, BLIP, BLIP-2 and CLIP covering feature extraction, VQA, captioning, image-text matching, text localization and zero-shot classification. If your task is not in that list, LAVIS is not offering you a starting point.
Second, run pip install salesforce-lavis in a clean environment and then check the resolved versions of transformers, timm and fairscale against requirements.txt. If pip silently upgrades them, the released checkpoints may not load the way the notebooks expect.
Third, read evaluate.py and the relevant config under projects/ before writing any pipeline code. The README does not document the argument surface of these scripts, so the files are the documentation. Skipping that step is how people end up guessing at flags that do not exist.
Editorial conclusion
Adopt LAVIS if you are reproducing or extending BLIP, BLIP-2, InstructBLIP or the other models listed in the README, and you can work inside its pinned dependency set. Do not adopt it if you need a current transformers release in the same environment, or if you only want an inference endpoint for one captioning model. Before committing, verify that the task you need appears in the examples directory, that the checkpoint you want is reachable from the registry the library uses, and that your Python and PyTorch versions satisfy python_requires >=3.7.0 and torch>=1.10.0 in requirements.txt.
Frequently asked questions
What is LAVIS?
LAVIS is an open-source deep learning library from Salesforce for language-vision research and applications. The README describes it as a one-stop library with a unified interface for image-language and video-language models and common datasets, covering tasks such as classification, retrieval, captioning, visual question answering, dialogue and pretraining.
How do you install LAVIS?
The README states that LAVIS is available on PyPI, and setup.py declares the distribution as salesforce-lavis, so it installs with pip install salesforce-lavis. Because requirements.txt pins transformers==4.33.2 and timm==0.4.12, use a fresh environment rather than an existing project environment.
Which vision-language models does LAVIS include?
The README lists BLIP-2, InstructBLIP, BLIP-Diffusion and X-InstructBLIP as model releases, along with Img2LLM-VQA and PNP-VQA. Each has a corresponding directory under projects/ with its own project page and, for several of them, a Colab notebook.
What licence does LAVIS use?
LAVIS is released under BSD-3-Clause. LICENSE.txt is at the repository root, setup.py declares license="3-Clause BSD", and the README links to the Open Source Initiative page for that identifier.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/salesforce-lavis)