# Video-ChatGPT: the transformers pin is a commit hash and the weights are not on GitHub

> An ACL 2024 video conversation model from MBZUAI whose requirements.txt installs transformers from one commit, whose weights sit on a personal file share, and whose benchmark questions live in a sibling repository.

**mbzuai-oryx/Video-ChatGPT** — [ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.

- Repository: https://github.com/mbzuai-oryx/Video-ChatGPT
- Website: https://mbzuai-oryx.github.io/Video-ChatGPT
- Stars: 1,512 · Forks: 129
- Language: Python
- License: CC-BY-4.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/mbzuai-oryx-video-chatgpt

## requirements.txt names a commit hash where a version number belongs

The pin file has fifteen lines, and the one that decides how the language half of the model behaves carries no version at all. Line three reads `transformers@git+https://github.com/huggingface/transformers.git@cae78c46`, so pip clones the transformers repository and checks out that one commit before building it. Which release cae78c46 sits on is written nowhere in the tree, and with no GitHub release there is nothing to compare it against. Everything else is frozen at 2023 values: torch==2.0.1, numpy==1.24.3, Pillow==9.5.0, decord==0.6.0, pydantic==1.10.7, gradio==3.23.0, accelerate==0.20.3, sentencepiece==0.1.99. One line stays open, tokenizers>=0.13.3, so a fresh install resolves that package to whatever is newest while the framework it has to work with stays on one unpublished commit.

## PYTHONPATH is the install step, because nothing in the tree is packaged

The top-level entries are LICENSE, README.md, data/, docs/, quantitative_evaluation/, requirements.txt, scripts/ and video_chatgpt/. There is no setup.py and no pyproject.toml among them, so the importable model code in video_chatgpt/ and the separate evaluation harness in quantitative_evaluation/ are made importable by hand. The documented sequence sets up conda, installs the pins, then exports the repository root onto the path:

```shell
conda create --name=video_chatgpt python=3.10
conda activate video_chatgpt

git clone https://github.com/mbzuai-oryx/Video-ChatGPT.git
cd Video-ChatGPT
pip install -r requirements.txt

export PYTHONPATH="./:$PYTHONPATH"
```

The interpreter is pinned by the conda command rather than by the requirements file, which never mentions Python at all. Anything that imports video_chatgpt from another checkout needs the same export, and moving to a different checkout means editing the path rather than reinstalling.

## FlashAttention is scoped to training and frozen at tag v1.0.7

FlashAttention appears once, described as an addition for training, and it is built from a checkout rather than taken from a wheel:

```shell
pip install ninja

git clone https://github.com/HazyResearch/flash-attention.git
cd flash-attention
git checkout v1.0.7
python setup.py install
```

Those five lines decide the attention kernels. The tag is fixed at v1.0.7, a 2023 point in that project's history, so the extension compiles against whatever CUDA toolchain is already on the machine and against the torch 2.0.1 the pin file installed. `python setup.py install` is the legacy invocation this line of flash-attention expects, not the pip build path later versions use. None of it is needed for inference. Offline demo instructions and training instructions are delegated to docs/offline_demo.md and docs/train_video_chatgpt.md, two links that go nowhere else in the top-level page.

## Weights and extracted features sit on a personal file share, with no release

The repository publishes no GitHub releases, so nothing here can be fetched by tag. The Jun-08-23 update says models, datasets and extracted features are all available from one link, and that link opens a personal folder on the mbzuaiac-my file share domain rather than a package registry or an object store with durable URLs. Distribution is split across three hosts with three different lifetimes: code in this repository, datasets on HuggingFace under the MBZUAI organisation including VideoInstruct-100K, and trained weights plus pre-extracted visual features on that personal folder. The browser demo at ival-mbzuai.com/video-chatgpt is a fourth location, and the demo clips sit in yet another personal folder. Nothing in the tree records a checksum for the weights or a tag saying which ones were used.

Next to that split sits the licensing asymmetry: the declared license is CC-BY-4.0, a Creative Commons attribution license written for content, and one root LICENSE covers the Python packages, the data directory and the weights without separating them.

## The evaluation harness is here, the benchmark questions are in VideoGPT-plus

quantitative_evaluation/ is in this tree, but neither of the two evaluation assets the project promotes is. The Jun-14-24 entry announces VCGBench-Diverse, given as 4,354 human annotated QA pairs spread across 18 video categories, together with a semi-automatic video annotation pipeline. Both are linked to the mbzuai-oryx/VideoGPT-plus repository and to HuggingFace datasets under MBZUAI, not here. The same date released VideoGPT+ itself from that sibling repository. So the harness that scores a run and the questions it scores against sit in two different projects, and the code you can read locally is not the code that produced the published comparisons. VideoInstruct-100K, the 100,000 video instruction pairs used for training, is a third location again, announced Sep-30-23 on HuggingFace.

## The navigation table links to two sections the page never reaches

The table near the top of the README links to Offline Demo, Training, Video Instruction Data, Quantitative Evaluation and Qualitative Analysis. The body that follows covers the first three and then stops. The Video Instruction Dataset section ends inside its own link label, one word short of a finished anchor: `VideoInstructionDatas`. The human-assisted and semi-automatic annotation framework that sentence points at is named in the link and never described, and the two closing sections the table promises, the benchmark tables and the sample answers, do not appear in the page at all. Read from the top, the installation sequence is complete and the evaluation story is only a list of section names, so anyone planning to reproduce a published number has nothing here to reproduce it from.

## The state of the art claim rests on six badges and no numbers

Six leaderboard badges sit under the title: VCGBench-Diverse, video-based generative performance, and zero-shot question answering on MSVD-QA, MSRVTT-QA, TGIF-QA and ActivityNet. The Jun-28-23 update says the page was changed to feature benchmark comparisons against Video Chat, Video LLaMA and LLaMA Adapter, which names three conversational baselines from that year. No score, no table and no evaluation protocol appear in the text. The claimed coverage is video reasoning, creativity, spatial and temporal understanding, and action recognition, and the one place a number would sit is the missing Quantitative Evaluation section. The paper is arXiv 2306.05424 and the acceptance at ACL 2024 is dated May-16-24, so those leaderboard links resolve to a 2023 conference artifact rather than to the current generation of video language models.

## The newest update announces a different owner's repository

The Updates section runs newest first, and its top line, Mar-28-25, announces Mobile-VideoGPT. That project lives at github.com/Amshaker/Mobile-VideoGPT under a different owner and is credited with results on multiple benchmarks at 2x higher throughput. Every entry about Video-ChatGPT itself stops at Jun-14-24. That is a news log, not an activity record: the last recorded push to main is 2026-09-05 and no release has ever been tagged, so commits have landed for over a year with nothing to pin them to. For anyone planning to build on this, the gap says more than the push date alone. A push tells you something is still being committed; it does not tell you which version of the model those commits produce, or whether the pinned transformers commit still imports against them.

## Conclusion

Read this as a 2023 research code drop rather than a product. Video-ChatGPT is worth using if you want the ACL 2024 architecture, the 100K instruction pairs and the evaluation harness in one place, and you are willing to assemble the environment yourself: work out what transformers commit cae78c46 corresponds to, fetch the weights from the authors' file share, and reconcile a local harness with questions that live in VideoGPT-plus. Before you quote any published number, check whether the benchmark table you are citing exists in the project's own page, because the leaderboard badges are the only comparison it shows. For a later line of work from the same organisation, VideoGPT-plus is where the benchmark assets and the annotation pipeline ended up.

## FAQ

### What does the mbzuai-oryx/Video-ChatGPT repository actually contain?

The tree holds the model package in video_chatgpt/, an evaluation harness in quantitative_evaluation/, plus data/, docs/, scripts/, requirements.txt and a single root LICENSE. It publishes no GitHub releases, and the trained weights with pre-extracted features are reached through a personal file share link rather than through the repository.

### How do you set up Video-ChatGPT, and why does it need PYTHONPATH?

Conda with Python 3.10, then pip install -r requirements.txt, then an export that puts the repository root on PYTHONPATH. No setup.py or pyproject.toml exists in the top-level entries, so the path export is the only documented way to make video_chatgpt importable. FlashAttention is a separate source-built extra the project scopes to training.

### Which version of transformers does Video-ChatGPT need?

The pin file gives no version number. It installs from a commit, transformers@git+https://github.com/huggingface/transformers.git@cae78c46. The surrounding pins are exact, from torch 2.0.1 and pydantic 1.10.7 down to gradio 3.23.0, and tokenizers>=0.13.3 is the single open range in the file.

### Where are the VideoInstruct 100K training pairs and the VCGBench-Diverse questions?

VideoInstruct-100K, the 100,000 video instruction pairs used for training, is on HuggingFace under MBZUAI, announced Sep-30-23. VCGBench-Diverse, with 4,354 human annotated QA pairs across 18 video categories, and the semi-automatic annotation pipeline are linked to the VideoGPT-plus repository instead of this one.

### Does Video-ChatGPT generate video files, or discuss video?

It discusses video. The overview calls it a video conversation model that combines LLM capabilities with a pretrained visual encoder adapted for spatiotemporal video representation, so the output is conversation about a video rather than frames. The related datasets and leaderboard links cover QA and generative performance, not video synthesis.

## Sources

- [Issues](https://github.com/mbzuai-oryx/Video-ChatGPT/issues)
- [License: CC-BY-4.0](https://github.com/mbzuai-oryx/Video-ChatGPT/blob/main/LICENSE)
- [mbzuai-oryx/Video-ChatGPT on GitHub](https://github.com/mbzuai-oryx/Video-ChatGPT)
- [Project website](https://mbzuai-oryx.github.io/Video-ChatGPT)
- [README](https://github.com/mbzuai-oryx/Video-ChatGPT/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/mbzuai-oryx-video-chatgpt
