Model or dataset
OpenGVLab/InternVideo avatar
OpenGVLab/InternVideo

InternVideo: a monorepo of video foundation models, from InternVideo1 to InternVideo3

[ECCV2024] Video Foundation Models & Data for Multimodal Understanding

2,398 stars161 forksPythonApache-2.0

At a glance

What is it?
OpenGVLab's InternVideo repository collects five generations of video foundation models plus the InternVid dataset. It is a research monorepo, not a single installable package, and the README points you at subdirectories rather than a pip install.
Who is it for?
InternVideo is for research groups and engineers who need video foundation model weights, evaluation scripts and pretraining code, and who are comfortable working per subdirectory. It is not for teams that want one pip install and a stable Python API, because the README does not document a packaged release or a versioning policy.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 91 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What problem InternVideo solves, and for whom

Video understanding research has a distribution problem: the weights, the pretraining recipe, the evaluation scripts and the dataset annotations usually live in different places. InternVideo bundles them. The repository describes itself as containing "InternVideo series and related works in video foundation models", and the top level lists five model lines plus a dataset: InternVideo1, InternVideo2, InternVideo2.5, InternVideo3, InternVideo-Next and Data/InternVid.

The audience is narrow and specific. Someone training or fine-tuning a video model, someone who needs a pretrained video encoder for retrieval or action recognition, or someone who wants the InternVid video-text annotations. The topics list on the repository names the tasks directly: action recognition, temporal action localization, video question answering, video retrieval, zero-shot classification. If your work sits in one of those, the repository is aimed at you. If you want a hosted video API, it is not.

How the InternVideo repository is organised

There is no single package. Each generation is a subdirectory with its own model zoo, scripts and documentation. The README's update log is effectively the map: InternVideo2 checkpoints and scripts arrived in 2024.04, InternVideo2.5 in 2025.01, InternVideo-Next in 2025.12, and InternVideo3 in 2026.06 with a technical report, an 8B instruct model, a long-video SFT dataset, evaluation scripts and an initial video-agent implementation in a separate repository called Vidify.

The split matters because the generations are not interchangeable. InternVideo2 is described as scaling video foundation models for multimodal video understanding, with smaller S/B/L variants distilled from InternVideo2-1B and VideoCLIP variants built with MobileCLIP. InternVideo2.5 is about long and rich context modeling. InternVideo-Next is described as general video foundation models for genuine world understanding, and one of the related search phrases frames it as being without video-text supervision. InternVideo3 is about multimodal contextual reasoning via efficient long-horizon agents.

So the data flow is: pick a generation, download its checkpoint from Hugging Face, and run that subdirectory's scripts. The repository itself is the index and the shared home, not the runtime.

Getting the code and finding the right checkpoint

The README does not give a top-level installation command, and there is no pip package name in it. What it does give is a set of subdirectories, each with its own model zoo file. For InternVideo2 the README points to ./InternVideo2/single_modality/MODEL_ZOO.md for the smaller S/B/L models and ./InternVideo2/multi_modality/MODEL_ZOO.md for the VideoCLIP variants. Start there rather than at the repository root, because the per-generation instructions live in those files.

The repository URL appears in the README's own links, and the top level contains a .gitmodules entry, so some content is pulled in separately. Clone it and read the model zoo for the generation you want before writing any code:

bash
git clone https://github.com/OpenGVLab/InternVideo.git
cd InternVideo
cat InternVideo2/single_modality/MODEL_ZOO.md

For the chat-style models, the README links InternVideo2-Stage3-8B at https://huggingface.co/OpenGVLab/InternVideo2-Chat-8B and InternVideo2-Stage3-8B-HD at https://huggingface.co/OpenGVLab/InternVideo2_chat_8B_HD, where 8B means InternVideo2-1B paired with a 7B LLM. The InternVideo2.5 model is published at https://huggingface.co/OpenGVLab/InternVL_2_5_HiCo_R16, InternVideo-Next weights sit in the collection https://huggingface.co/collections/OpenGVLab/internvideo-next, and the InternVideo3 8B instruct model is at https://huggingface.co/yanziang/InternVideo3-8B-Instruct.

Those are web links, not commands. The README does not show a download invocation, so how you fetch a checkpoint depends on the Hugging Face tooling you already use. What you should end up with is a local copy of the checkpoint files. What the README does not give you is a single inference entry point that works across all five generations, so the next step is generation-specific and documented in the subdirectory, not at the root.

The InternVid dataset and where annotations actually live

Data/InternVid is a separate concern from the models, and it is the part with the longest release history. The README records that video instruction data were released in 2023.05 for tuning end-to-end video-centric dialogue systems, that the full InternVid video annotation of 230M video-text pairs was released in 2024.06 on OpenDataLab and Hugging Face, and that InternVid2 annotations arrived in 2024.07. InternVid itself was accepted for spotlight presentation at ICLR 2024.

This is the most concrete deliverable in the repository for anyone whose bottleneck is data rather than architecture. The annotations are hosted externally, which means the repository is a pointer rather than a mirror. That is worth knowing before you plan a download: the storage and bandwidth cost sits with OpenDataLab or Hugging Face, and the repository's role is documentation and loader code.

One caveat visible in the README: the full 230M-pair annotation and the InternVid2 annotation are distinct releases with distinct hosts. If you are reproducing a paper that used one, check which one before assuming the other is equivalent.

Where InternVideo is the wrong tool

The repository is a research monorepo, and that shows in ways that matter operationally. There are no releases listed, so there is no versioned artifact to pin. If you build a product on InternVideo2 code and InternVideo3 later changes shared conventions, nothing in the README describes a deprecation or migration path. The README does not document rollback either.

Second, the generations overlap in scope. InternVideo2, InternVideo2.5 and InternVideo3 all address video understanding at different scales, and the README presents them as parallel entries rather than a single recommended path. Choosing between them is a research decision you make from the technical reports, not something the README resolves for you.

Third, the support channel is a WeChat group, which the README invites people to join for questions about trial, running or deployment. That is a real constraint for teams that need searchable, asynchronous issue history. There is a GitHub repository, but the README does not describe an issue triage process or a support commitment. If you need an SLA, this is not the project for you.

InternVideo compared with InternVL

The search data shows people asking about internvideo vs internvl, and the repository gives a partial answer. InternVL appears in the README only as a Hugging Face model identifier: the InternVideo2.5 model is published at OpenGVLab/InternVL_2_5_HiCo_R16. That naming suggests the 2.5 line is built on the InternVL family rather than competing with it.

The practical difference in approach: InternVL is the vision-language model line, and InternVideo is the video-specific line, with InternVideo2.5 explicitly tying the two together for long and rich context modeling. If your task is single-image or general multimodal dialogue, the InternVideo README gives you no reason to start here. If your task is temporal, spanning many frames and often long video, the InternVideo subdirectories and the InternVid annotations are the parts that do not exist in the InternVL naming.

The README does not provide a benchmark comparison between the two, so treat any claim about relative accuracy as something you would have to establish yourself from the technical reports.

Licence and the cost of keeping up

The repository is Apache-2.0. That covers the code in this repository. It does not automatically cover the model weights, which are hosted on Hugging Face and OpenDataLab under their own terms, and the README does not restate those terms. If you plan to redistribute a checkpoint or use it commercially, read the licence on the specific model page. This is a description of what the repository states, not legal advice.

The upgrade cost is the more interesting number. The README's update log shows a release roughly every few months across 2025 and 2026: InternVideo2.5 in 2025.01, InternVideo-Next in 2025.12, InternVideo3 in 2026.06. Each brings its own report, weights and in some cases its own dataset. For a team that adopts one generation and stays there, that cadence is irrelevant. For a team that wants the newest capability, it means re-reading a technical report and re-validating a checkpoint on a regular schedule, and the repository offers no compatibility shim between generations. The last push to the repository was on 2026-07-02.

Editorial conclusion

InternVideo is for research groups and engineers who need video foundation model weights, evaluation scripts and pretraining code, and who are comfortable working per subdirectory. It is not for teams that want one pip install and a stable Python API, because the README does not document a packaged release or a versioning policy. Before committing, verify which subdirectory matches your task, confirm the checkpoint's licence on its Hugging Face page, and check whether the code you need lives in InternVideo2, InternVideo2.5 or InternVideo3, since the README treats them as separate entries with separate reports.

Frequently asked questions

How do I install InternVideo?

There is no top-level install command in the README. You clone the repository and then follow the model zoo file inside the generation you want, such as InternVideo2/single_modality/MODEL_ZOO.md, and fetch weights from Hugging Face.

Where are the InternVideo model weights hosted?

The README links checkpoints on Hugging Face, including a collection for InternVideo2, OpenGVLab/InternVL_2_5_HiCo_R16 for InternVideo2.5, a collection for InternVideo-Next, and yanziang/InternVideo3-8B-Instruct for InternVideo3.

What is the difference between InternVideo2 and InternVideo2.5?

The README describes InternVideo2 as scaling video foundation models for multimodal video understanding, and InternVideo2.5 as empowering video mllms with long and rich context modeling. They are separate subdirectories with separate technical reports.

Is InternVideo the same as InternVL?

No. InternVL appears in the InternVideo README only as the Hugging Face identifier OpenGVLab/InternVL_2_5_HiCo_R16, which is where the InternVideo2.5 model is published. The README does not present them as the same project.

What is the InternVid dataset?

The README describes InternVid as a large-scale video-text dataset for multimodal understanding and generation, accepted for spotlight presentation at ICLR 2024. The full 230M video-text pair annotation was released in 2024.06 on OpenDataLab and Hugging Face.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. OpenGVLab/InternVideo on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/opengvlab-internvideo.svg)](https://hysenlabs.com/projects/opengvlab-internvideo)