Model or dataset
NVlabs/VILA avatar
NVlabs/VILA

NVlabs/VILA: an Apache-2.0 vision language model family with a separate model licence

VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and cloud.

3,863 stars334 forksPythonApache-2.0

At a glance

What is it?
VILA is a family of open vision language models from NVIDIA Labs, tuned for multi-image and video understanding. The code is Apache-2.0, but the weights carry a non-commercial licence, and the last push to the repository was on 2026-03-12.
Who is it for?
Adopt VILA if you need an open VLM that handles several images or video frames in one context and you can live with the CC BY-NC 4.0 weights licence, which rules out commercial deployment of the released checkpoints. Skip it if you want a permissively licensed model, a small dependency set, or a repository with recent commits; the last push was on 2026-03-12.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What VILA is for, and who should care

VILA is a family of open vision language models built by NVIDIA Labs, described in the README as designed to optimize both efficiency and accuracy for efficient video understanding and multi-image understanding. That phrasing points at the two tasks that separate it from a generic image captioner: reasoning over several images at once, and reasoning over sampled video frames. The README lists VILA examples for video captioning, in-context learning, multi-image reasoning, and a Jetson Orin deployment, which is a fair summary of the intended audience: engineers who need to run multimodal inference on hardware ranging from a laptop or an Orin module up to an A100.

The project ships more than weights. The repository contains training code, evaluation code and finetuning directories, plus serving and demo_trt_llm directories for deployment. That makes it interesting to two groups: people evaluating an open VLM as a drop-in component, and people who want to finetune or reproduce a multimodal model rather than only call an API. The version in pyproject.toml is 2.0.0, matching the NVILA (VILA2.0) release announced in December 2024.

The licence split is the first thing to check. The code badge in the README is Apache-2.0, while the model badge is CC By NC 4.0. Those are two different grants covering two different artifacts.

How the VILA stack is put together

The layout of the repository is the clearest description of the architecture. The llava directory holds the model and CLI code, finetuning holds training entry points, longvila holds the long-video work, vila_hd holds the high-resolution variant, and serving plus demo_trt_llm cover deployment. The package exposes four console scripts through pyproject.toml: vila-run, vila-eval, vila-infer and vila-upload. Those names tell you the intended workflow: run a demo, evaluate a checkpoint, run inference, and push a model to Hugging Face.

VILA's stated origin is interleaved image-text pretraining, which the README credits with enabling multi-image input and in-context learning. Later releases layered on video support (VILA-1.5), a full-stack efficiency rework (NVILA), and long-context video (LongVILA, which the README says supports more than 1M context length with a multi-modal sequence parallel system). VILA-HD swaps in PS3, a vision encoder that the README says scales vision pre-training to 4K resolution.

The dependency list in pyproject.toml is where the practical weight sits. It pins torch==2.3.0, torchvision==0.18.0, transformers==4.46.0, numpy==1.26.4, timm==0.9.12, and pulls two packages straight from Git: s2wrapper from bfshi/scaling_on_scales and ps3-torch from NVlabs/PS3. Training extras add deepspeed==0.9.5. That is a dense, tightly pinned environment, and the pins are the main source of install friction.

Installing VILA and running a first inference

The repository provides environment_setup.sh at the top level, and the Dockerfile shows how the maintainers use it: the image starts from nvcr.io/nvidia/pytorch:24.06-py3, copies pyproject.toml and llava, then runs the setup script with the argument vila, which is also the conda environment name used by the container command. If you follow the container path, the setup step looks like this:

bash
bash environment_setup.sh vila

After that the container runs server.py inside the vila environment, which is the path to the Gradio demo rather than to a library import. The README badge states Python 3.10+, while pyproject.toml declares requires-python >=3.8; the two disagree, and given the pinned torch and transformers versions, treating 3.10 or newer as the safe floor is the reasonable reading.

For a scripted first use, the package installs as vila and registers the CLI entry points. A minimal install followed by the inference entry point looks like this:

bash
pip install -e .
vila-infer --help

The README does not spell out the argument set for vila-infer, so expect to read llava/cli/infer.py before you can pass a model path and an image. That is a real gap: the project documents its models and its benchmarks far better than it documents its own command line.

Weights are not in the repository. The README links a Hugging Face collection under Efficient-Large-Model/nvila, and the vila-upload script exists for pushing checkpoints back out, which implies the expected workflow is to download a checkpoint from Hugging Face and point the CLI at it. If you only want the demo, server.py inside the Docker image is the shortest path.

Where VILA gets expensive or awkward

The licence is the hard limit. The README carries a model licence badge pointing at MODEL_LICENSE and reading CC By NC 4.0. Non-commercial means the released checkpoints are not something you can ship in a commercial product without a separate arrangement, even though the surrounding code is Apache-2.0. Teams that skim the green code badge and stop there will get this wrong.

The second constraint is hardware and precision. The README's deployment notes describe AWQ-quantized 4-bit VILA-1.5 models running through TinyChat and TensorRT-LLM backends on A100, 4090, 4070 Laptop, Orin and Orin Nano. The published throughput and time-to-first-token tables are explicitly labelled as measured with the TinyChat backend at batch size 1, with W4A16 for the language model and W8A8 for the vision tower, against an FP16 baseline. Those numbers are a comparison between two configurations of the same model on specific GPUs, not a promise about your workload. Batch size 1 in particular means the tables say nothing about server-style throughput.

The dependency pins are the third problem. torch==2.3.0 and transformers==4.46.0 are exact versions, and the install reaches into two Git repositories. If your environment already runs a different torch build, VILA will fight it. The Dockerfile sidesteps this by pinning the whole base image, which is why the container is the recommended route in practice even though the README leads with the model family.

Finally, VILA is the wrong tool when you need a small, permissively licensed model for a narrow task. Image classification, OCR, or a single-image captioning pipeline do not need multi-image context or a 1M-token video path, and you would be paying the dependency and licence cost for capability you never use.

VILA against a general-purpose multimodal API

The obvious alternative for many teams is a hosted multimodal model from a commercial provider, and pyproject.toml shows the project is aware of that world: openai==1.8.0 appears in the dependency list, and the evaluation code uses it. The difference in approach is not accuracy, it is control. A hosted API gives you a model you cannot inspect, cannot finetune beyond whatever the vendor exposes, and cannot run on an Orin module in a disconnected environment. VILA gives you checkpoints, training code in finetuning, and a deployment path through TensorRT-LLM and TinyChat that the README documents down to per-GPU token rates.

What you give up is convenience and legal freedom. There is no managed scaling, no SLA, and the non-commercial weight licence blocks the most common commercial use. A team choosing between the two should treat the licence as the deciding factor before comparing benchmark tables, because a model you cannot ship is not a cheaper option regardless of its scores.

Within open models, the meaningful fork is VILA versus its own descendants. The README notes that OmniVinci, an audio-visual omni-modal model, is built on the VILA codebase, and that Long-RL supports RL training on VILA, LongVILA and NVILA models. If your task involves audio or reinforcement learning, those repositories are the closer starting point, and VILA is the substrate rather than the product.

Maintenance, upgrades and what the licence covers

The repository is not archived, and the last push was on 2026-03-12, roughly six months before this writing. That is neither abandoned nor brisk. The README news list is more revealing: the most recent entries are from July 2025, covering OmniVinci and Long-RL, both of which live in separate repositories. Within this repository, the newest functional line is the June 2025 PS3 and VILA-HD release, and the January 2025 note that VILA became part of the Cosmos Nemotron vision language models. The pattern is that new work lands in sibling projects and the VILA repository carries the base models and tooling.

Upgrade cost is dominated by the pins. Moving to a newer torch or transformers is not a configuration change; it is a compatibility exercise across llava, the two Git-sourced packages, and whatever backend you deploy on. The Dockerfile's base image, nvcr.io/nvidia/pytorch:24.06-py3, is the version the maintainers tested against, and drifting from it is where most upgrade pain will come from.

On licensing, the split is explicit in the README badges: Apache-2.0 for code, CC By NC 4.0 for models. The practical consequence is that you can read, modify and redistribute the training and inference code, but the checkpoints themselves carry a non-commercial restriction. This is a description of what the repository states, not legal advice; if your use is commercial, the MODEL_LICENSE file and your own counsel are the places to resolve it.

Editorial conclusion

Adopt VILA if you need an open VLM that handles several images or video frames in one context and you can live with the CC BY-NC 4.0 weights licence, which rules out commercial deployment of the released checkpoints. Skip it if you want a permissively licensed model, a small dependency set, or a repository with recent commits; the last push was on 2026-03-12. Before committing, check the Hugging Face collection for the exact checkpoint you need, confirm the Python and CUDA versions your hardware supports, and read MODEL_LICENSE rather than relying on the Apache-2.0 code badge.

Frequently asked questions

What is the NVIDIA Vision Language Model that NVlabs/VILA refers to?

VILA is a family of open vision language models from NVIDIA Labs, designed for efficient video understanding and multi-image understanding, with model sizes that the README lists as 3B, 8B, 13B and 40B for the 1.5 generation. The README also notes that as of January 6, 2025 VILA became part of the Cosmos Nemotron vision language models.

How do vision language models like VILA work?

VILA uses interleaved image-text pretraining, which the README credits with enabling multi-image input and in-context learning, and later releases added video understanding and long-context support. Deployment runs the model through backends such as TinyChat and TensorRT-LLM, with quantized 4-bit checkpoints for GPUs including A100, 4090 and Orin.

Which vision language models does VILA include?

The README lists VILA-1.5 in 3B, 8B, 13B and 40B sizes, NVILA (also called VILA2.0), LongVILA for long video with more than 1M context length, VILA-HD using the PS3 vision encoder, and VILA-U as a separate unified model. Model checkpoints are published in a Hugging Face collection under Efficient-Large-Model/nvila.

Is VLM better than LLM for VILA's use cases?

The README does not make a general comparison between VLMs and LLMs; it presents VILA as a vision language model family for video and multi-image understanding, alongside text-only models such as those in the Long-RL repository. The relevant distinction in the repository's own framing is task fit, not a ranking.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. NVlabs/VILA on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nvlabs-vila.svg)](https://hysenlabs.com/projects/nvlabs-vila)