# LLaVA-OneVision-2: a fully open 8B multimodal stack, from codec encoders to training logs

> LLaVA-OneVision-2 is an Apache-2.0 release from EvolvingLMMs-Lab that unifies image, long video and spatial reasoning in one 8B model, and ships the encoders, data, configs and logs alongside the weights. The interesting part is the codec-aligned vision encoder; the awkward part is that the repository is a training framework, not a drop-in inference package.

**EvolvingLMMs-Lab/LLaVA-OneVision-2** — Fully Open Framework for Democratized Multimodal Training

- Repository: https://github.com/EvolvingLMMs-Lab/LLaVA-OneVision-2
- Website: https://evolvinglmms-lab.github.io/LLaVA-OneVision-2/projects/index.html
- Stars: 1,215 · Forks: 76
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/evolvinglmms-lab-llava-onevision-2

## What LLaVA-OneVision-2 actually ships, and who it is for

Most multimodal releases hand you weights and a model card. LLaVA-OneVision-2 hands you the pipeline. The README describes it as a fully open 8B multimodal model that unifies image, long-form video and spatial understanding under one architecture, and states that data, encoders, training, checkpoints and logs are released end to end. Four datasets are named: LLaVA-OneVision-2-VideoCaption for dense video captions, LLaVA-OneVision-2-Spatial for 3D-aware spatial reasoning, plus LLaVA-OneVision-1.5-Mid-Training-85M and LLaVA-OneVision-1.5-Instruct carried forward from the previous generation.

That framing tells you the audience. This is for teams that intend to fine-tune, ablate or reproduce, not teams looking for a hosted endpoint. If you only need image question answering, an 8B model with a training harness attached is more machinery than you want. If you need to know what data a model saw and how it was scheduled, the release is unusually complete.

The repository layout supports that reading. Alongside configs/, dockerfile, tests/ and tools/ sit aiak_megatron/, aiak_training_llm/, offline_packing/ and transformers_impl/, which is the shape of a training codebase that also carries a vendored model implementation. Examples are split by generation under examples/llava_onevision2/ and examples/llava_onevision1_5/, with a qwen2_5_vl example alongside them.

## Codec-style patch selection and why the token budget matters

The technical claim that separates this release from LLaVA-OneVision-1.5 is the encoder. OneVision-Encoder and OneVision-Encoder-Lang are described as HEVC-style vision transformers that add a codec-stream input mode next to the existing image and uniform-frame video modes. The mechanism, per the README, is patch selection: rather than sampling sparse frames densely, the encoder selects motion- and residual-rich patches and samples dense frames sparsely.

The diagram caption in the method section states the payoff in concrete terms: the same 54-token budget covers roughly three times the temporal range compared with uniform sampling. That is the whole argument. Long video is a context problem, and if most patches in a clip are static background, uniform frame sampling spends tokens on pixels that carry no information. A codec-style selector spends them on the frames where something moved.

The trade-off is real and the README does not hide it so much as skip past it. A codec-aligned encoder needs a codec stream at inference time, which means your serving path has to produce one. That is why the evaluation reproduction notes mention both a frames backend and a codec backend: the model runs either way, but only one of them gets the temporal coverage the method is built around. If your pipeline decodes video to raw frames before it reaches the model, you are on the frames backend and the headline advantage does not apply to you.

## Installing LLaVA-OneVision-2 and running a first example

The README does not give a pip install line for the package itself. What it does give is a requirements.txt at the repository root and a Quick Start section scoped to 4B on a single node, with example directories per model generation. The dependency list is the practical starting point, and it is opinionated: transformers==5.7.0, accelerate==1.9.0, datasets==2.19.2, timm==1.0.3, qwen_vl_utils, megatron-energon==5.0.0, hydra-core==1.3.2 and omegaconf==2.3.0.

Install those with the requirements file before anything else, because the pins are strict and several of them are not what a current environment would resolve to on its own.

```bash
pip install -r requirements.txt
```

After that, the examples directory is where a first run lives. The repository splits examples by generation, so a LLaVA-OneVision-2 run belongs under examples/llava_onevision2/ rather than the 1.5 directory. The README does not spell out the flags for the scripts in that directory, so read the files there rather than assuming a command line; the Quick Start heading promises a single-node 4B path, and that is the configuration the examples are written against.

For evaluation rather than training, the README points elsewhere entirely. The reported numbers are reproduced from the llava-onevision2 branch of lmms-eval, which the project says contains the evaluation model wrapper, benchmark and task configs, a Docker environment and launcher scripts for both the frames and codec backends. That branch, not this repository, is where you check whether a published score matches your setup.

## Where the release is thin, and when it is the wrong tool

The documentation gap is inference. There is no packaged runtime here. The README links to a vLLM model page for llava_onevision2 and to an online playground, which is a clear signal that serving is expected to happen in someone else's stack. If you want to load the model, call generate and move on, this repository is the wrong entry point; the Hugging Face model page for LLaVA-OneVision-2-8B-Instruct is the one you want, and the training code here is overhead.

The second constraint is hardware. The Quick Start is explicitly a 4B single-node recipe, and the released instruct model is 8B. Nothing in the README promises a smaller or quantized variant, and the dependency set includes megatron-energon and hydra, which are training-side tools rather than inference conveniences. Budget accordingly.

The third is version churn. requirements.txt pins transformers==5.7.0 while the repository also carries a transformers_impl/ directory, which suggests the model code is vendored rather than upstreamed at the pinned version. The README does not document what happens when the two diverge, and it does not document a rollback path if a training run fails partway. Treat the pinned environment as the supported one and do not assume a newer transformers will work.

Finally, the spatial claim is harder to verify than the video claim. LLaVA-OneVision-2-Spatial is named as a dataset and 3D-aware spatial reasoning is listed as a capability, but the README's mechanism section is about codec patch selection, which is a temporal technique. Depth and layout reasoning are asserted rather than explained, and the evaluation reproduction branch is the place to look for the tasks that back them up.

## LLaVA-OneVision-2 against Qwen2.5-VL and the 1.5 generation

The nearest comparison is the previous generation, and the repository keeps it visible: examples/llava_onevision1_5/ sits next to examples/llava_onevision2/, the 1.5 datasets are still shipped, and a 1.5 branch is linked from the table of contents. The difference in approach is the encoder. LLaVA-OneVision-1.5 used a conventional vision transformer with uniform frame sampling; version 2 adds the codec-stream mode and the OneVision-Encoder weights, which the news section dates to a separate February 2026 release with its own technical report. If you are already running 1.5, the upgrade is not a checkpoint swap: it is a change to how frames reach the encoder, and it only pays off if you can supply a codec stream.

The other comparison the repository invites is Qwen2.5-VL, which is present as examples/qwen2_5_vl/ and as a dependency in the form of qwen_vl_utils. That is worth noticing: the project uses Qwen's video utility package inside its own requirements. The architectural difference is that Qwen2.5-VL is a model you consume, while LLaVA-OneVision-2 is a training framework that produces a model. If your goal is to fine-tune on your own video corpus and own the resulting weights under Apache-2.0, this repository is built for that. If your goal is to call a strong multimodal model today, the training harness is a detour.

## Licence, maintenance and what an upgrade costs

The repository is Apache-2.0, which is permissive and includes an explicit patent grant. That matters more here than for a typical library, because the release includes encoder weights, datasets and training code. Apache-2.0 covers the code in this repository; the README does not state the licence of the released model weights or of the four datasets, which live on Hugging Face under separate namespaces. Check those pages individually before you build a product on them. None of this is legal advice.

The last push to the default branch was on 2026-09-10, and the repository is not archived. Release 2.0 is dated 2026-08-06, with 1.5 before it on 2025-12-26. That is a young release line with a recent commit, and the news section shows the project has been shipping steadily across encoders, datasets and an RL recipe for the previous generation.

Upgrade cost is the part to weigh carefully. The pinned transformers==5.7.0 and the vendored transformers_impl/ mean that moving this stack forward is not a matter of bumping a version number. A generation upgrade also changes the encoder input mode, so any preprocessing you wrote against uniform frame sampling needs revisiting. The README does not describe a migration path from 1.5 to 2.0, and it does not document rollback. Plan for the upgrade as a re-run, not a patch.

## Conclusion

Adopt LLaVA-OneVision-2 if you need a permissively licensed multimodal stack whose data mixture, encoders and training logs are all inspectable, and if you have the GPU budget to fine-tune an 8B model or the patience to wire the transformers_impl path into your serving stack. Do not adopt it if you want a one-line inference dependency: the README points at vLLM and lmms-eval rather than at a packaged runtime, and requirements.txt pins transformers==5.7.0, which will collide with most existing environments. Before committing, check the lmms-eval llava-onevision2 branch for the exact task configs behind the reported numbers, and confirm that your serving path supports the codec backend and not only the frames backend.

## FAQ

### What is LLaVA-OneVision-2?

It is the next-generation release of the LLaVA-OneVision family: a fully open 8B multimodal model that unifies image, long-form video and spatial understanding under a single architecture, released with data, encoders, training code, checkpoints and logs. The repository is the training framework around that model, licensed Apache-2.0.

### How do I install LLaVA-OneVision-2?

The README does not give a package install command. It ships a requirements.txt at the repository root, so the documented path is to install those pinned dependencies and then work from the examples/llava_onevision2/ directory, which the README's Quick Start section frames as a 4B single-node setup.

### What is the OneVision-Encoder and how is it different from a normal vision transformer?

OneVision-Encoder and OneVision-Encoder-Lang are described as HEVC-style vision transformers that add a codec-stream input mode alongside image and uniform-frame video. Instead of sampling sparse frames densely, they select motion- and residual-rich patches and sample dense frames sparsely, which the README says gives about three times the temporal range under the same 54-token budget.

### Can I run LLaVA-OneVision-2 without training it myself?

Yes, but not from this repository. The README links to a vLLM model page for llava_onevision2 and to the LLaVA-OneVision-2-8B-Instruct weights on Hugging Face, which is where inference is expected to happen. The code here is the training and evaluation pipeline.

## Sources

- [EvolvingLMMs-Lab/LLaVA-OneVision-2 on GitHub](https://github.com/EvolvingLMMs-Lab/LLaVA-OneVision-2)
- [License: Apache-2.0](https://github.com/EvolvingLMMs-Lab/LLaVA-OneVision-2/blob/main/LICENSE)
- [Project website](https://evolvinglmms-lab.github.io/LLaVA-OneVision-2/projects/index.html)
- [README](https://github.com/EvolvingLMMs-Lab/LLaVA-OneVision-2/blob/main/README.md)
- [Releases](https://github.com/EvolvingLMMs-Lab/LLaVA-OneVision-2/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/evolvinglmms-lab-llava-onevision-2
