# Ovis: a multimodal model whose base language changed three times, and whose pins are exact

> An Apache-2.0 multimodal architecture from a research group, released across nine milestones on three different base language models, with the setup script making every line of the requirements file an install-time pin, performance claims delivered as images, and an install command that clones from a different organisation than the repository.

**ATH-MaaS/Ovis** — A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.

- Repository: https://github.com/ATH-MaaS/Ovis
- Website: https://huggingface.co/AIDC-AI/Ovis2.5-9B
- Stars: 1,522 · Forks: 89
- Language: Python
- License: Apache-2.0
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/ath-maas-ovis

## The install clones from a different organisation than the repository you are reading

There are three names in this project and they are not the same name. The repository sits under one organisation. Every model identifier in the readme sits under a second one, including the two released checkpoints and the model family collection. The install command clones from the second: 

```bash
git clone git@github.com:AIDC-AI/Ovis.git
conda create -n ovis python=3.10 -y
conda activate ovis
cd Ovis
pip install -r requirements.txt
pip install -e .
```

And the contact address in the hiring notice is at a third organisation entirely, a large cloud provider rather than either repository host. The declared homepage is a model card rather than a project site, which fits a project whose primary artefact is weights rather than a tool. None of this is concealed and none of it is unusual for a research lab that publishes models under one account and code under another, but it produces two practical frictions. The clone command is an SSH URL, so it needs a key with access to that account, which excludes the anonymous tarball path and the read-only clone that most evaluators would take. And a search for the model family and a search for the code will not obviously lead to each other. If you are evaluating this, resolve all three names before you start, because the readme alone will not get you to a working checkout without a guess.

## The setup script makes every line of the requirements file an install requirement

The packaging here is a single setup script, and it does something that will affect your environment more than anything in the model card. It opens the requirements file, splits it into lines, and hands the whole list to the install requires argument. So the file you can read is the file pip enforces, and the style of that file is the style of the package's constraints. Read it and the pattern is clear. The machine-learning stack is pinned exactly, with a fixed version for the deep learning framework, the transformers library, the tokeniser, a sentence piece library, an arrow library, an accelerate library, a pydantic core library, a markdown renderer with its optional extras, a numeric library, a machine-learning toolkit, a web framework, an einops pair, an image models library, a tokenizer library, a streaming generator, a video library, a deep learning training framework, a subtitle library, a movie library and an imaging library. Meanwhile the utility libraries, the two HTTP clients, the ASGI server, the tensor library, the tabular library, the audio library, the attention library, the tokenizer library and the interface library, are all unpinned. The consequence is a package that fixes the parts everyone else also depends on and leaves free the parts nobody fights over. In a dedicated container that is sensible. In a shared environment it is a collision waiting to happen, and the deep learning training framework in the list is the heaviest of them for a project whose documented workflows are inference and serving.

## One vision tower, three base languages, nine releases

The release list is the most informative document in this repository and it is a list of nine entries running from June 2024 to August 2025. Read the model names in order and the architecture becomes visible. The first generation was built on one model's family. The next generation added a second family at a much larger size, and a third small family, and then published quantised versions of both. The second generation of the project came as a spread of six sizes from one to thirty-four billion parameters, and the current generation is two sizes. So the base language model has changed three times in about fourteen months, and the repository's own topic list carries two of those families as separate tags, which is a small confirmation. What stayed constant is the vision side, and the current model table shows the same vision tower for both released sizes while the language model changes size underneath it. That is the design claim in one table: the visual encoder is fixed and the language model is a pluggable component, and the readme says so directly, that the architecture can be instantiated with popular language models. It also explains the release cadence. A new base model is a new checkpoint rather than a new codebase, which is why this project can ship nine milestones without nine forks.

## The performance claims are images

The performance section contains one sentence of text and three images. The sentence says the current generation shows strong results on general multimodal benchmarks, complex chart analysis and reasoning tasks, and achieves leading performance among open-source models under a given parameter ceiling. Then the evidence is a figure for overall performance, a figure for something labelled with an initialism, and a figure for reasoning. So the actual measurements are not in the readme, in any form you can copy into a spreadsheet or check against a leaderboard you already trust. That is a common pattern in this field and it is not a criticism of honesty; benchmark numbers go stale and images are easier to regenerate. It is a practical problem for an adopter, though, because the readme is where you look first and the numbers are not there. The linked technical report and the model card are where the tables are, and the model card is the one that will be versioned alongside the weights. If you are making a selection decision, go to the model card for the size you intend to run and read the numbers for the tasks you care about, and treat the phrase about leading performance among models under the parameter ceiling as a pointer rather than a finding.

## The serving path is OpenAI-compatible, and the thinking mode is a template argument

Two inference paths are documented and both are worth reading for what they reveal. The first is a Gradio interface launched as a script, taking a model path and a port. The second is a serving engine, and the two details that matter are the trust flag and the client. 

```bash
vllm serve AIDC-AI/Ovis2.5-9B \
     --trust-remote-code \
     --port 8000
```

The trust flag means the server downloads and executes Python from the model repository to build the model class, which is the same trust decision you make with any model that is not natively implemented in the transformers library. Read that code before you run it in a service. The second detail is the client: the example talks to the local server with an OpenAI-compatible client, using an empty key and a local base URL, and passes a multimodal message with an image part and a text part. So the model is exposed through a standard chat completions interface rather than a bespoke one, which is the single most useful fact for integration. It means an existing OpenAI-shaped client works unchanged, and a gateway or proxy that speaks that protocol can sit in front of it. The extension arguments are where the model-specific features appear. The thinking mode is enabled through a chat template argument and its length is bounded by a separate budget argument, so you can cap the reasoning phase. And there is a minimum pixel count for images, set in the example to a value annotated as a megapixel, which is a floor rather than a target: a small image gets scaled up, and that directly affects how many visual tokens the request costs.

## A committed .DS_Store, an undocumented plugin directory, and no numbers

Three small observations that together describe the repository rather than the project. The top-level listing begins with a macOS metadata file, which is an artefact that should never be committed and is present in a repository that also has an ignore file, so either it predates the ignore file or it was added afterwards. It is harmless and it is the kind of thing that tells you the contributor workflow is not enforcing anything automatically. Next, there is a plugin directory that the visible documentation never mentions. Given the serving story, it is plausibly related to the serving engine, and the readme excerpt does not say, so if you are extending the serving path that is the first directory to open and the first question to ask. And then there is the version story. The repository publishes no GitHub releases at all, so the release history is the list in the readme, which is a list of announcements rather than of artefacts, and the artefacts live on a model host with their own versioning. The last push was in July 2026, and the newest announcement in the readme is from August 2025, so there is a gap between the last announced milestone and the last commit with nothing in the readme explaining it. The licence is Apache-2.0 for the code, and the weights are a separate question the readme does not address.

## Conclusion

Adopt Ovis if you need a multimodal model small enough to serve on one card and you are prepared to treat the repository as the specification, because the serving path is an OpenAI-compatible endpoint with a documented thinking mode and a fixed vision tower, which makes it easy to drop into an existing client. Do not adopt it if your environment is a shared Python environment, because the setup script turns every line of the requirements file into an install requirement and the machine-learning stack is pinned to the patch version, so torch, transformers, numpy and deepspeed will be fixed by this package rather than by you. Two things to check before you install. Compare the pins against what is already in your environment, because a serving-only use case still pulls a distributed training framework at a fixed version. And read the performance material rather than the readme, because the claims in the readme are images with no numbers in the text, and the model card and the linked technical report are where the measurements are.

## FAQ

### How do I install and run Ovis?

Clone the repository with the command given, create a Python 3.10 environment, install the requirements file and then the package in editable mode. The tested versions are stated as Python 3.10, Torch 2.4.0, Transformers 4.51.3 and DeepSpeed 0.15.4, and a matching serving engine version is pinned separately with a versioned wheel index.

### What base models has Ovis been built on?

Three families across the release history: an earlier family, then a second generation built on two different families at several sizes, and the current generation on a third. The readme states the architecture can be instantiated with popular language models, and the current model table shows the same vision tower for both released sizes while the language model differs.

### Can I serve Ovis through an OpenAI-compatible client?

Yes. The serving example launches the model with a trust-remote-code flag and a port, and the client example talks to it with the OpenAI Python SDK using an empty key and a local base URL, passing a message with image and text parts. Model-specific features such as the thinking mode are passed as extra arguments.

### What are the exact dependencies of the Ovis package?

The setup script reads the requirements file and passes its lines to the install requires argument, so the file is the contract. The machine-learning stack is pinned to exact versions including torch, transformers, numpy, accelerate, deepspeed, pillow and timm, while utility libraries such as the HTTP clients, the tabular library and the interface library are unpinned.

### How is the Ovis thinking mode controlled?

Through a chat template argument that enables it, and a separate budget argument that controls the length of the reasoning phase. Both are passed as extra arguments in the request body, and the demo file for it is named separately from the basic inference demo.

## Sources

- [ATH-MaaS/Ovis on GitHub](https://github.com/ATH-MaaS/Ovis)
- [Issues](https://github.com/ATH-MaaS/Ovis/issues)
- [License: Apache-2.0](https://github.com/ATH-MaaS/Ovis/blob/main/LICENSE)
- [Project website](https://huggingface.co/AIDC-AI/Ovis2.5-9B)
- [README](https://github.com/ATH-MaaS/Ovis/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ath-maas-ovis
