Model or dataset
OpenBMB/MiniCPM-V avatar
OpenBMB/MiniCPM-V

MiniCPM-V: A 1.3B Vision-Language Model Built for Phones, Not Servers

A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone

26,470 stars2,076 forksPythonApache-2.0

At a glance

What is it?
MiniCPM-V 4.6 is a 1.3B-parameter multimodal model from OpenBMB aimed at on-device image and video understanding, with an Ollama library entry and open-sourced iOS and Android adaptation code. Here is what the repository documents, what it leaves out, and when a different model is the better call.
Who is it for?
Adopt MiniCPM-V if you need image or video understanding inside a phone app, a laptop, or a single modest GPU, and you accept that the 1.3B model will lose to 8B-class vision models on dense OCR and fine-grained chart reading.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 22 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem MiniCPM-V 4.6 solves: vision understanding where there is no GPU rack

Most capable vision-language models assume a server. They are large, they need batching to be economical, and they are useless in an app that has to work on a train. MiniCPM-V takes the opposite constraint as its starting point. The README describes the series as "multimodal LLM series designed for strong performance and efficient deployment on devices," and the flagship model, MiniCPM-V 4.6, is 1.3B parameters in total.

The intended user is not a research lab benchmarking OCR accuracy. It is an engineer shipping an app that has to look at a photo or a video frame and answer a question about it, on hardware the user already owns. The README names iOS, Android and HarmonyOS as supported deployment targets and says the edge adaptation code is open-sourced. That is a narrower and more specific promise than "runs locally," and it is the reason to care about this project at all.

The second audience is anyone who wants a small vision model in a local inference stack. MiniCPM-V 4.6 was merged into Ollama's official model library on 2026-06-25, which means the shortest path to running it does not involve this repository's Python at all.

How MiniCPM-V 4.6 gets small: intra-ViT early compression and mixed 4x/16x tokens

The mechanism the README highlights is visual token compression. A vision-language model turns an image into a sequence of visual tokens that the language model then attends over, and that sequence is usually the dominant cost. MiniCPM-V 4.6 uses what the README calls "intra-ViT early compression technique" from LLaVA-UHD v4, applied inside the vision transformer rather than after it. The README states this reduces visual encoding computation cost by more than 50%.

The second piece is a mixed 4x/16x visual token compression rate. Rather than fixing one compression ratio for every input, the model supports two, which the README frames as "a more flexible performance-efficiency trade-off in different tasks." The practical reading: a task that needs fine spatial detail can spend more tokens, and a task that only needs a coarse summary can spend far fewer. The README does not document how the ratio is selected at inference time, and that is a real gap for anyone planning to tune it.

The architecture is split across the repository in a way that reflects the model family. The top-level entries include chat.py for inference, omnilmm/ for the model code, finetune/ for training, eval_mm/ for evaluation, and web_demos/ for the interactive demos. A separate requirements_o2.6.txt exists alongside requirements.txt, which tells you the omnimodal branch has its own dependency set rather than sharing one environment.

Installing MiniCPM-V 4.6 and running your first image query

The README does not print a step-by-step pip install block. It points to the model on Hugging Face (openbmb/MiniCPM-V-4.6), to an Ollama library entry, and to a Cookbook repository for scenario guides. The requirements.txt in the repository is the concrete dependency list, and it is pinned tightly: torch==2.1.2, torchvision==0.16.2, transformers==4.40.0, accelerate==0.30.1, and gradio==4.41.0 among others. Two heavy optional packages, xformers and flash_attn, are commented out in that file.

If you already have a Python environment, the dependency install is the first step. Note that requirements.txt also pulls a wheel from an Aliyun OSS URL for modelscope_studio, which will fail on a network that cannot reach that host.

bash
pip install -r requirements.txt

For the fastest first result, the Ollama route avoids the Python environment entirely. The README states MiniCPM-V 4.6 was merged into Ollama's official model library on 2026-06-25.

bash
ollama run minicpm-v4.6

Once the model is available, the repository's own entry point is chat.py. The README does not reproduce its flags, so read the file before assuming an interface. The repository also ships an API service documented in docs/api.md, released on 2026-05-17 for both MiniCPM-V 4.6 and MiniCPM-o 4.5. The README does not state the port the service listens on, so check docs/api.md rather than guessing.

For the interactive route, web_demos/ holds the realtime web demo that the 2026-02-06 news entry says is deployable on your own devices such as a Mac or a GPU. The README notes that the hosted web demo may experience latency issues due to network conditions, which is exactly why the local demo path exists.

Where MiniCPM-V 4.6 is the wrong choice

The 1.3B parameter count is the whole point and also the ceiling. The README claims MiniCPM-V 4.6 "surpasses larger models like Gemma4-E2B-it in performance," but the comparison set in that sentence is small models, not frontier systems. If your task is dense OCR over multi-column scanned documents, or reading precise values off a chart, a larger vision model will generally do better, and the README offers no evidence to the contrary.

The second limitation is the dependency pinning. torch==2.1.2 and transformers==4.40.0 are exact pins. If your project already runs a newer transformers, installing requirements.txt will downgrade it. There is no documented optional-dependency split for the vision-only path versus the omnimodal path beyond the separate requirements_o2.6.txt file, and the README does not explain which file applies to which model.

The third gap is mobile packaging. The README says the model "can be deployed across common mobile platforms, including iOS, Android and HarmonyOS, with edge adaptation code open-sourced," but the download link for the app points to a separate repository, OpenBMB/MiniCPM-V-Apps. Nothing in this repository's top-level layout is an Xcode project or a Gradle build. Treat the mobile story as documented and externally hosted, not as something you get by cloning this repo.

Finally, the README notes the hosted web demo may suffer latency from network conditions. A real-time streaming demo over a network is a different reliability profile from local inference, and the README itself says a Docker image for local deployment of the realtime demo is still being worked on.

MiniCPM-V versus MiniCPM-o, and versus a general small vision model

The most consequential alternative is inside the same family. MiniCPM-o 4.5 is 9B parameters and, per the README, "approaches Gemini 2.5 Flash in vision, speech, and full-duplex multimodal live streaming." The difference in approach is not just size. MiniCPM-o supports full-duplex streaming, meaning its speech and text outputs do not block its real-time video and audio inputs, which the README says lets it "see, listen, and speak simultaneously" and perform proactive interactions such as proactive reminding. MiniCPM-V does not do that. If you need a camera that talks back in real time, MiniCPM-V is the wrong branch of the family, and the 9B model is the one to evaluate.

Against a general small vision model, the difference is the compression strategy. Most small VLMs reduce cost by shrinking the language model and accepting whatever visual token count the vision encoder produces. MiniCPM-V 4.6 attacks the visual encoding cost directly, claiming more than 50% reduction, and exposes a 4x/16x compression choice. That is a more targeted bet: it assumes the bottleneck is visual tokens, which is true for high-resolution images and video frames, and less true for short text-heavy prompts.

The repository also ships an API service, documented in docs/api.md, which positions it as a drop-in HTTP endpoint rather than a library. That matters if you want to compare it against a hosted vision API without rewriting your client.

Maintenance, licensing and the cost of keeping up

The repository is not archived, and the last push was on 2026-09-08, which is recent enough that the project is still moving. The most recent tagged release is 20250527, covering MiniCPM-V and MiniCPM-o, dated 2025-05-27. That gap between the release tag and the latest push is worth noting: the repository is being updated, but the release naming does not track the model versions described in the README, where MiniCPM-V 4.6 and MiniCPM-o 4.5 are the current models. If you pin to a release tag, verify which model weights that tag corresponds to.

The upgrade cost is real because of the pinned dependencies. Moving to a newer transformers or torch means editing requirements.txt and re-validating, and the repository gives no compatibility matrix. The separate requirements_o2.6.txt file suggests the omnimodal path already diverged once.

Licensing is Apache-2.0, which is permissive and generally compatible with commercial use, but this article is not legal advice. Two things to check yourself: whether the model weights on Hugging Face carry the same license as the repository, and whether the separate MiniCPM-V-Apps repository, which holds the mobile adaptation code, uses the same terms. The repository's LICENSE file covers this repository only.

Editorial conclusion

Adopt MiniCPM-V if you need image or video understanding inside a phone app, a laptop, or a single modest GPU, and you accept that the 1.3B model will lose to 8B-class vision models on dense OCR and fine-grained chart reading. Do not adopt it as a drop-in replacement for a hosted frontier model on hard document tasks, and do not expect the repository to hand you a packaged mobile build: the iOS and Android adaptation code is open-sourced but the README points at a separate MiniCPM-V-Apps download page. Before committing, verify three things: that your target platform's adaptation code is in the apps repository, that the transformers==4.40.0 and torch==2.1.2 pins in requirements.txt coexist with your existing stack, and that the Hugging Face model card for the exact checkpoint you plan to serve matches the capability you need.

Frequently asked questions

What is MiniCPM-V?

It is a multimodal LLM series from OpenBMB designed for efficient deployment on devices, focused on vision-language understanding across image, video and text inputs. The current model described in the README is MiniCPM-V 4.6, a 1.3B-parameter model.

What is MiniCPM-o, and how does it differ from MiniCPM-V?

MiniCPM-o extends the family toward real-time end-to-end omnimodal interaction with streaming video and audio inputs plus text and speech outputs. MiniCPM-o 4.5 is 9B parameters and supports full-duplex multimodal live streaming, so its outputs do not block its real-time inputs; MiniCPM-V focuses on vision-language understanding.

How do I use MiniCPM-V?

The README points to three routes: the Hugging Face model page, an Ollama library entry for MiniCPM-V 4.6, and the repository's own chat.py entry point plus an API service documented in docs/api.md. The README does not print a step-by-step install block, so the concrete dependency list is requirements.txt.

How does MiniCPM-V compare to other LLMs?

The README claims MiniCPM-V 4.6 surpasses larger models like Gemma4-E2B-it in performance while showing superior efficiency than smaller models like Qwen3.5-0.8B, at roughly 1.5x token throughput. It does not publish a comparison against 8B-class or frontier vision models in the README text.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. OpenBMB/MiniCPM-V on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/openbmb-minicpm-v.svg)](https://hysenlabs.com/projects/openbmb-minicpm-v)