Open-source project
OpenBMB/MiniCPM-Robot avatar
OpenBMB/MiniCPM-Robot

MiniCPM-Robot trades parameters for a policy that runs on the robot

A Smarter and Faster On-Device AI Brain for Robots

340 stars28 forksPythonApache-2.0

At a glance

What is it?
MiniCPM-Robot pairs a 1.5 billion parameter manipulation policy with a 0.9 billion parameter on-device tracker, claiming a minute of visual memory at close to single-frame inference cost. The efficiency numbers are specific, and the latency figure excludes action decoding.
Who is it for?
MiniCPM-Robot suits a mobile robot where putting a network in the control loop is unacceptable and a policy has to run on hardware the machine carries, particularly where a task needs memory of something no longer in view. It is the wrong choice for a fixed installation with reliable networking, where a policy several times the size can run off-board and be updated without touching the fleet.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 36 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Two small models for two different robot problems

MiniCPM-Robot is an embodied intelligence model family aimed at running on the robot rather than on a server behind it. Two models were released together, and they address different halves of what a robot needs.

The first is a 1.5 billion parameter vision-language-action model for manipulation, described as a generalist policy covering downstream tasks with one set of weights, in simulation and on real hardware. The second is a 0.9 billion parameter target tracker, which the project describes as the first fully on-device embodied tracker, covering static, dynamic and adversarial targets through natural language and vision alone.

The size is the entire argument. Models of this kind are usually large enough that inference happens somewhere with a datacentre GPU, which puts a network between perception and action. For a robot that has to close a control loop, that latency is not an inconvenience but a limit on what the robot can do. Compressing the policy until it fits on the machine changes which behaviours are possible.

The audience is robotics teams evaluating whether current vision-language-action models can run on hardware they can actually mount, and researchers comparing generalist policies at the small end of the size range.

The streaming context result is the striking number

The most interesting claim concerns how history is handled, and it is specific enough to reason about.

A policy that benefits from remembering what it has already seen faces a cost problem. Recomputing over sixty frames of history at every decision step is stated as requiring 125 trillion floating point operations per step. The project's streaming approach reduces that to 3.3, which is roughly a fortieth of the work, and supports up to a minute of visual context memory while keeping the online cost comparable to conventional single-frame reactive inference.

If that holds, it dissolves a real trade-off. Memory and reactivity normally pull against each other, because the cheap approach forgets and the remembering approach is too slow to act. Getting a minute of visual memory at close to the cost of no memory is the kind of result that changes what a policy can be asked to do, since tasks involving an object that has gone out of view stop requiring a separate memory system.

A second efficiency claim compounds it. The model inherits visual token compression from its vision-language ancestor, reducing each frame from 256 tokens to 64, a fourfold reduction in what the language model must process per frame.

Read the latency comparison with its own caveat attached

The headline latency figure is 120 milliseconds of model-forward time per decision step on a datacentre GPU in half precision with single-frame input, against 234 milliseconds for a named larger competitor.

The project then states, in the same breath, that this measurement excludes task autoregressive decoding. That disclosure deserves credit and it also changes what the number means. Model-forward time is the cost of pushing observations through the network; generating the action tokens is separate and sequential, and for an autoregressive policy that decoding is frequently the larger share of wall clock time. So the comparison is like for like against the competitor, and it is not the end-to-end latency a robot experiences.

Anyone evaluating this should measure the full step on their own hardware rather than planning a control loop around the published figure.

The throughput claim comes from elsewhere and is worth separating. A third-party inference project added support on the day of release and reports raising throughput on a different datacentre GPU from 10 to 37 hertz, attributed to graph capture and custom fused kernels. That is a substantial gain and it belongs to that project's optimisation work rather than to the model, which is a useful distinction when deciding what you would need to reproduce it.

Getting the models, and what the repository holds

There is no package to install and no command in the documentation available here. The repository is organised as two directories, one per model, each with its own quick start, inference and evaluation documentation, alongside an assets directory and a Chinese translation of the README.

The weights themselves are distributed through two model hubs rather than the repository, with a collection published on each, so obtaining the models means fetching them from one of those rather than cloning anything. The release note points at both.

For a first real use the practical sequence is to decide which of the two models matches your problem, since they are separate models solving separate tasks and nothing here suggests they compose, then follow the quick start inside that model's directory. The manipulation model targets sim and real deployment; the tracking model is demonstrated on a specific quadruped platform.

That platform detail is worth reading carefully. The tracking figures, quoted as more than five frames per second at around 180 milliseconds, are measured on one named robot. Those numbers describe that hardware, and the point of the project is that the model runs on what the robot carries, so the hardware is part of the claim rather than incidental to it.

What the benchmark claims can and cannot establish

The manipulation model is said to beat a 3 billion parameter policy and a policy of more than 5 billion parameters on representative evaluations. Read the qualifier: the evaluations are representative in the authors' judgement, and they were selected by the same group that trained the model.

That is normal practice and it is why the comparison should be treated as a reason to test rather than as a result. A small generalist policy beating larger ones is a surprising claim, and surprising claims in this area usually depend heavily on which tasks were chosen and how success was defined.

The tracking model's claim is framed differently and is stronger for it. Leading among open-source systems on a named public benchmark is a checkable statement, because the benchmark exists independently and others can run against it.

Other limits are structural. There are no tagged releases, so anyone depending on this should pin a commit. The hardware named across the various figures spans two datacentre GPU classes and a commercial quadruped, none of which is a modest requirement. And the repository documents evaluation and inference rather than training, so reproducing the models rather than running them is not something the published material supports.

Running a larger policy off-board is the alternative

The alternative is the arrangement this project is arguing against: run a larger vision-language-action model on a server and send actions to the robot over a network.

The difference in approach is where the trade-off lands. A larger policy has more capability per decision, is easier to update because it lives on infrastructure you control, and can serve several robots from one deployment. It also puts a network in the control loop, which adds latency you cannot bound and a failure mode where the robot is blind the moment connectivity drops. For anything mobile, that second point is decisive.

On-board inference removes both problems and pays for it in capability, since 1.5 billion parameters is a real constraint against models several times the size, and in operational awkwardness, because updating a model means updating the fleet.

The choice follows from the robot rather than from the model. A fixed manipulator in a facility with reliable networking can reasonably run a larger policy off-board. A mobile robot operating where connectivity is unreliable cannot, and for that case a smaller policy that runs locally is not a compromise but the only option, which is the ground this family is built for.

Apache terms and what to verify first

The repository is Apache-2.0 licensed with the file present, which includes an express patent grant and permits commercial use, and is the permissive end of the range for robotics research code. The model weights are distributed separately through the hubs and carry their own terms, which is the licence that governs deploying them in a product and is a separate question from this repository. This is not legal advice.

The repository is compact: two model directories, a contributors file, assets, and documentation in both English and Chinese. The last push was on 2026-08-25, with the models released on 2026-07-19.

The verification order is short. Establish which of the two models addresses your task, since they do not overlap. Then measure a complete decision step on your own hardware, including action decoding, because the published latency figure explicitly excludes it. Finally, if the streaming context behaviour is the reason you are interested, test it on a task that actually requires remembering something out of view, since that capability is where the fortyfold compute reduction would earn its place.

Editorial conclusion

MiniCPM-Robot suits a mobile robot where putting a network in the control loop is unacceptable and a policy has to run on hardware the machine carries, particularly where a task needs memory of something no longer in view. It is the wrong choice for a fixed installation with reliable networking, where a policy several times the size can run off-board and be updated without touching the fleet. Measure a full decision step including action decoding on your own hardware before planning around the published 120 millisecond figure, since the project states that measurement excludes autoregressive decoding, and pin a commit, because the repository has no tagged releases.

Frequently asked questions

What models does MiniCPM-Robot include?

Two. A 1.5 billion parameter vision-language-action model for generalist manipulation in simulation and on real hardware, and a 0.9 billion parameter embodied target tracker covering static, dynamic and adversarial targets using vision and natural language only.

How does MiniCPM-Robot handle visual history cheaply?

Through streaming inference rather than recomputation. The project states that sixty frames of history would require 125 trillion floating point operations per decision step by recomputation, while streaming needs 3.3, supporting up to a minute of visual context at a cost comparable to single-frame reactive inference.

Is the quoted MiniCPM-Robot latency the full decision time?

No. The 120 millisecond figure is model-forward latency per decision step on a datacentre GPU in half precision with single-frame input, and the project states the measurement excludes task autoregressive decoding, so end-to-end time will be higher.

Where do I get the MiniCPM-Robot weights?

From the model hub collections published alongside the release rather than from the repository itself. The repository holds a directory per model, each with its own quick start, inference and evaluation documentation.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. OpenBMB/MiniCPM-Robot on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/openbmb-minicpm-robot.svg)](https://hysenlabs.com/projects/openbmb-minicpm-robot)