Model or dataset
NVlabs/OmniVinci avatar
NVlabs/OmniVinci

OmniVinci: NVIDIA's Open-Source Omni-Modal LLM for Video, Audio, and Text

OmniVinci is an omni-modal LLM for joint understanding of vision, audio, and language.

682 stars56 forksPythonApache-2.0

At a glance

What is it?
OmniVinci is an open-source 9-billion-parameter language model from NVIDIA that processes video, audio, and text jointly. Published at ICLR 2026, it introduces three architectural components for cross-modal alignment and outperforms Qwen2.5-Omni on cross-modal and video benchmarks using a training budget six times smaller.
Who is it for?
OmniVinci is the right starting point for research or production work that requires a single model to handle video, audio, and text together without assembling separate specialist models. It is not the right choice for tasks that need only a single modality, where smaller, specialized models will perform better for the same inference cost.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 30 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What OmniVinci Is and the Problem It Addresses

Most publicly available language models handle one or two modalities: text, or text plus images. Handling video with synchronized audio as a single input stream requires understanding temporal alignment across modalities, matching what is heard to what is seen at a precise time. Separate models for audio and video do not share representations and cannot reason across those boundaries.

OmniVinci is NVIDIA's open-source contribution to this problem. The model, called OmniVinci-9B, was released on October 19, 2025 and accepts video, audio, and text as joint input. The project was presented as a paper at ICLR 2026 under the title "Enhancing Architecture and Data for Omni-Modal Understanding LLM." The README describes three downstream application domains: robotics, medical AI, and smart factory environments, all of which generate sensor streams that include vision, audio, and other signals together.

The model weights are available on HuggingFace as nvidia/omnivinci under the Apache-2.0 license. The last push to the repository was on August 31, 2026.

Three Architectural Innovations for Cross-Modal Alignment

The README describes three specific architectural contributions. First, OmniAlignNet strengthens the alignment between vision and audio embeddings in a shared omni-modal latent space. The goal is to produce a representation where the model can draw on both modalities simultaneously rather than routing them through separate encoding paths that only merge at a late stage.

Second, Temporal Embedding Grouping captures relative temporal alignment between vision and audio signals. When a video contains speech, a door slam, and background music, these sounds occur at different frames. Temporal Embedding Grouping provides a mechanism for the model to track which audio segments correspond to which video frames.

Third, Constrained Rotary Time Embedding encodes absolute temporal information in omni-modal embeddings. Rotary position embeddings are a standard technique in transformer models for encoding sequence position; the "constrained" variant is adapted for multi-modal temporal inputs where the position index must reflect actual time rather than token count.

These three components work together so that the model's internal representations carry both which modality is being processed and when in time that content occurs.

Training Data: 24 Million Cross-Modal Conversations

The README describes a curation and synthesis pipeline that generated 24 million conversations for training OmniVinci. These include both single-modal conversations (text-only, image-only, audio-only) and omni-modal conversations that combine modalities. The README notes a key finding: modalities reinforce one another in both perception and reasoning, meaning that training on combined modality data improves performance on individual modalities compared to training on each modality independently.

The total training token count for OmniVinci is 0.2 trillion. The README compares this directly to Qwen2.5-Omni, which used 1.2 trillion training tokens. This six-to-one reduction in training compute is the efficiency claim the paper makes for the architecture and data approach combined. The codebase builds on the NVILA codebase, as indicated by the environment_setup.sh script and the pyproject.toml naming the package vila.

Downloading and Running OmniVinci-9B

OmniVinci-9B is distributed through HuggingFace. Download the model weights and switch into the downloaded directory:

bash
huggingface-cli download nvidia/omnivinci --local-dir ./omnivinci --local-dir-use-symlinks False
cd ./omnivinci

The repository includes a setup script that installs the environment based on the NVILA codebase:

bash
bash ./environment_setup.sh omnivinci

After setup, inference examples are in the repository root. The README references example_mini_audio.py for audio and image inference examples, and example_mini_video.py for video. The top-level entry point for general inference is example_infer.py. Loading the model requires trust_remote_code=True because the model architecture uses custom code not present in the standard transformers package. The pyproject.toml pins torch==2.3.0, which means the environment must match that version to avoid compatibility issues with the model's dependencies.

The inference code uses AutoModel and AutoConfig from the transformers library. The model loads with torch_dtype set to float16 for reduced memory use.

Benchmark Results Against Qwen2.5-Omni

The README provides a benchmark comparison table against Qwen2.5-Omni. On the DailyOmni benchmark, which measures cross-modal understanding, OmniVinci scores 66.5 against Qwen2.5-Omni's 47.5, a difference of +19.05 points. On Worldsense (a second cross-modal benchmark), OmniVinci scores 48.2 against 45.4. On the MMAR audio benchmark, OmniVinci scores 58.4 against 56.7, a difference of +1.7. On Video-MME (video understanding without subtitles), OmniVinci scores 68.2 against Qwen2.5-Omni's 64.3, a difference of +3.9 points.

On MVBench (a video understanding benchmark), OmniVinci scores 70.6 against Qwen2.5-Omni's 70.3, a marginal difference. The README does not provide results on other commonly used benchmarks such as MMMU or VQAv2, so the comparison is limited to the benchmarks included in the paper.

Limitations and License Considerations

The installation path is non-standard. The environment_setup.sh script builds on the NVILA codebase, which is a research framework rather than a stable library. The pyproject.toml lists numerous pinned dependencies (accelerate==0.34.2, transformers==4.46.0, peft>=0.9.0, bitsandbytes==0.43.2) that may conflict with other packages in an existing environment. Setting up a dedicated virtual environment or conda environment is practically necessary.

The repository has no GitHub releases. The model weights were released through HuggingFace rather than a versioned GitHub tag, which means there is no formal release process tied to the code repository. The README documents only OmniVinci-9B; the README does not describe other model sizes or fine-tuning procedures.

The Apache-2.0 license permits commercial use, modification, and distribution. The requirements include preserving the license and NOTICE files in any distribution and stating changes made to the original code. The license does not grant trademark rights.

Editorial conclusion

OmniVinci is the right starting point for research or production work that requires a single model to handle video, audio, and text together without assembling separate specialist models. It is not the right choice for tasks that need only a single modality, where smaller, specialized models will perform better for the same inference cost. Before deploying it, verify that your environment can run the NVILA-based setup script (environment_setup.sh) and that torch==2.3.0 is compatible with the rest of your stack, since the pyproject.toml pins that version explicitly.

Frequently asked questions

What modalities does OmniVinci support?

OmniVinci-9B supports video with audio, standalone audio, and images as inputs, combined with text prompts. The README references example scripts for video, audio, and image inference separately. The model processes these modalities jointly rather than routing them through independent pipelines.

What is the license for OmniVinci?

OmniVinci is licensed under Apache-2.0. This permits commercial use, redistribution, and modification. Distributions must include the original Apache-2.0 license text and the NOTICE file, and must state any changes made to the original code.

How does OmniVinci compare to Qwen2.5-Omni?

According to the benchmarks in the README, OmniVinci scores 66.5 versus Qwen2.5-Omni's 47.5 on DailyOmni (cross-modal understanding), 68.2 versus 64.3 on Video-MME, and 58.4 versus 56.7 on MMAR (audio). OmniVinci achieved these results with 0.2T training tokens versus 1.2T for Qwen2.5-Omni.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. NVlabs/OmniVinci on GitHub
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nvlabs-omnivinci.svg)](https://hysenlabs.com/projects/nvlabs-omnivinci)