Open-source project
mudler/locate-anything.cpp avatar
mudler/locate-anything.cpp

locate-anything.cpp: open-vocabulary detection in C++ without a Python runtime

Port of Nvidia LocateAnything-3B on ggml

598 stars79 forksC++NOASSERTION

At a glance

What is it?
A ggml port of NVIDIA's LocateAnything-3B that returns the same boxes as the official implementation from a text prompt, and runs on CPU or through ggml's GPU backends. The trade-off is a 3B model, a large GGUF, and a build that expects you to know CMake.
Who is it for?
Adopt locate-anything.cpp if you need open-vocabulary boxes inside a C++ or LocalAI deployment and can accept a ~6.3 GB model file, a CMake build with a vendored ggml submodule, and a one-shot prompt per image. Do not adopt it if you need video tracking, a hosted API, or a model small enough for a constrained edge device; a 3B VLM plus a vision tower is not that.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 69 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What locate-anything.cpp does that a fixed-label detector cannot

A conventional detector ships with a closed label set. If the class you need was not in the training list, you retrain or you accept the nearest label. locate-anything.cpp takes the other route: you pass an image and a free-text description, and the model returns boxes with those labels attached. The README's own example asks for "person</c>car" and gets back a JSON array of detections plus an annotated PNG. The project describes itself as a C++17 inference port of NVIDIA's LocateAnything-3B, an open-vocabulary detection and visual-grounding VLM, built on ggml.

The intended reader is someone who already has images flowing through a C++ or C service and does not want to bolt a Python process onto the side of it. The README states plainly that there is no Python runtime at inference time. The Python side exists only to convert and quantize the checkpoint. That split is the whole point of the port: the awkward dependency lands in a one-time offline step, and the deployed artifact is a binary plus a GGUF file.

Token-space detection: how the Qwen2.5-3B plus MoonViT stack emits boxes

The architecture is three pieces stacked together. A Qwen2.5-3B language model, a MoonViT vision tower, and a two-layer MLP projector that maps vision features into the language model's embedding space. That is a standard VLM arrangement, and the interesting part is what happens at the output.

Detection is performed in token space. The model does not regress continuous coordinates through a detection head. It emits coordinate tokens drawn from the range <0> to <1000>, and those tokens decode to boxes. This is why the port can claim identical detections rather than merely similar ones: the pipeline from pixels to labeled boxes is discrete at the coordinate stage, so small numeric differences in the floating-point math do not move a box unless they flip a token. The README attributes the identical output to porting the full pixel-to-labeled-boxes pipeline and validating it against the official implementation.

Three decode modes are exposed. The default is called hybrid. There is a pure autoregressive mode labelled slow, and an MTP-only mode labelled fast. The README does not expand what MTP stands for, and the mode names are the only description given, so treat the distinction as a performance knob rather than something you can reason about from the README alone.

Building locate-anything.cpp and running a first detection

The build assumes CMake and a C++17 compiler. ggml is vendored at third_party/ggml, so clone with submodules or the build will fail on a missing directory. The README gives this exact sequence, and it enables both the tests and the CLI:

bash
git clone --recursive https://github.com/mudler/locate-anything.cpp
cd locate-anything.cpp
cmake -B build -DLA_BUILD_TESTS=ON -DLA_BUILD_CLI=ON && cmake --build build -j

After that you should have a locate-anything-cli binary. The CMake options that matter for hardware are LA_GGML_CUDA, LA_GGML_METAL, LA_GGML_VULKAN and LA_GGML_CANN, each defaulting to OFF and each forwarding the corresponding flag to the ggml submodule. LA_SHARED builds liblocate_anything as a shared library with ggml static-linked in, which is the option you want if you are embedding this rather than shelling out. Ascend NPU builds have a separate document at docs/ASCEND_NPU.md and, per the README, an f16-only caveat.

You still need weights. The prebuilt GGUFs live on Hugging Face at mudler/locate-anything.cpp-gguf. The README recommends q8_0 at roughly 6.3 GB, described as near-lossless and box-identical to f32. Downloading that file is the shortest path to a working detection. If you would rather produce the file yourself, the repository includes a converter that needs the Python dependencies in scripts/requirements.txt:

bash
python scripts/download_model.py
python scripts/convert_locateanything_to_gguf.py     # -> models/locate-anything-f32.gguf

That produces a full-precision f32 GGUF of about 15 GB. The README notes the Python gguf writer cannot emit K-quants, so quantization goes through the CLI instead, starting with locate-anything-cli quantize models/locate-anything-f32.gguf plus the target type. The README's truncation cuts off the rest of that command line, so check the CLI help rather than guessing the remaining arguments.

With a model in hand, detection is a single command. This is the README's example:

bash
locate-anything-cli detect --model models/locate-anything-q8_0.gguf \
    --input street.jpg \
    --prompt "Locate all the instances that matches the following description: person</c>car." \
    --annotated out.png

The output is a JSON object with a detections array, each entry carrying a label and a box, and an annotated PNG written to out.png. Note the prompt shape. It is not a bare noun list. The description is embedded in a sentence and terminated with a </c> token before the labels. If you drop that framing and pass "person, car", you are off the format the example demonstrates.

Where the port is slower, weaker, or simply the wrong choice

The GPU comparison is not a clean sweep, and the README says so. Against the official model run exactly as its model card documents, in bf16 and greedily, the port's f16 build is roughly 1.7 to 2.1 times faster and the q8_0 build roughly 1.9 to 3.1 times faster, with boxes identical. Against the official sampled out-of-box run, the result is mixed: faster on sparse scenes, comparable on dense ones, because sampling stops earlier on dense inputs. That is a real caveat buried in the performance section, and it means the headline speedup depends on which baseline you consider fair.

Size is the harder constraint. The smallest published file, q4_k, is about 4.7 GB, and it carries sub-pixel box drift relative to f32. q5_k has the same drift at about 5.1 GB. Only q8_0 (6.3 GB) and q6_k (5.5 GB) are described as box-identical. There is no small variant here. A 3B language model plus a vision tower plus a projector does not shrink to something you can drop into a microcontroller, and anyone hoping for a few-hundred-megabyte artifact is looking at the wrong project.

Quantization is also partial by design. Only the Qwen2 LM matmuls are quantized; the vision tower, the projector and the norms stay f32. That is why q8_0 is only about a third smaller than the 9.2 GB f16 file rather than four times smaller. It is a deliberate choice that protects the boxes, but it caps how much memory you can reclaim.

One more limit worth naming: the README documents image input and a single detection call. There is no mention of video, tracking, batching, or a server mode. If your problem is temporal, this is not the tool.

locate-anything.cpp against the official PyTorch LocateAnything-3B

The obvious alternative is the thing this project ports: NVIDIA's nvidia/LocateAnything-3B running under PyTorch. The difference is not accuracy, since the README claims identical boxes, but the deployment shape.

The official path gives you a Python process, the full PyTorch runtime, and the model card's own sampling defaults. It is the reference. If you are evaluating prompt formats, comparing against published numbers, or fine-tuning, that is where you want to be, because the port is a fixed inference path and does not offer training.

The port gives you a C++17 binary, a vendored ggml, and no Python at inference. That matters when the surrounding system is already C++ and adding a Python interpreter is a packaging problem you would rather not have. It also matters for hardware reach: ggml's backends mean CUDA, Metal, Vulkan and CANN are all reachable through the same CMake options, whereas the PyTorch path ties you to whatever the official stack supports. The cost is that you inherit the port's mode names, its quantization menu, and its release cadence rather than upstream's.

A second alternative worth considering is a closed-label detector with a fixed class list. It will be smaller and faster. It will also fail the moment your query does not map onto its classes, which is the exact problem this project exists to solve.

Release cadence, licence status and what upgrading costs

There is one release, v0.1.0, published on 2026-07-22, and the last push to master was on 2026-07-22. That is a single tagged version with no upgrade history to read, so anyone adopting this should assume the API surface may still move. The CLI flags shown in the README are the only contract documented so far.

Upgrade cost is dominated by the model file, not the code. Changing quantization means re-downloading a multi-gigabyte GGUF and re-running your own accuracy checks, because the README's box-identity claims are tied to specific quantization types rather than to the model in general. Moving from q8_0 to q4_k to save 1.6 GB trades away exact box parity, and that is a decision you have to measure on your own images.

On licensing, the README displays an MIT badge and links to a LICENSE file, but the repository metadata reports the licence as NOASSERTION, meaning GitHub could not classify it automatically. Those two signals disagree. The README also states that the GGUF files are derived from nvidia/LocateAnything-3B, which carries its own terms separate from whatever covers this port. Before shipping anything, read the LICENSE file in the repository and the upstream model's licence rather than relying on the badge. This is a description of what the repository shows, not legal advice.

Editorial conclusion

Adopt locate-anything.cpp if you need open-vocabulary boxes inside a C++ or LocalAI deployment and can accept a ~6.3 GB model file, a CMake build with a vendored ggml submodule, and a one-shot prompt per image. Do not adopt it if you need video tracking, a hosted API, or a model small enough for a constrained edge device; a 3B VLM plus a vision tower is not that. Verify three things before committing: that the prebuilt GGUF you download is box-identical to f32 for your prompt style (the README marks q8_0 and q6_k as box-identical and q5_k and q4_k as having sub-pixel drift), that your prompt follows the "Locate all the instances that matches the following description:" form with the closing </c> token, and that the LICENSE file confirms the MIT badge the README displays, since the repository metadata reports NOASSERTION.

Frequently asked questions

What is NVIDIA LocateAnything-3B, and how does locate-anything.cpp relate to it?

It is an open-vocabulary detection and visual-grounding VLM from NVIDIA, combining a Qwen2.5-3B language model, a MoonViT vision tower and a two-layer MLP projector. locate-anything.cpp is a C++17 inference port of it built on ggml, with the full pixel-to-labeled-boxes pipeline reimplemented and validated against the official implementation.

What is locate anything?

In this repository, it refers to LocateAnything-3B, an open-vocabulary detection and visual-grounding VLM that takes an image and a text prompt and returns labeled boxes. locate-anything.cpp ports that model to C++ on ggml, running on CPU and on GPU through ggml's backends.

How do you use locate-anything.cpp to detect objects from a text prompt?

Build the CLI with CMake, download a GGUF such as locate-anything-q8_0.gguf, then run locate-anything-cli detect with --model, --input and a --prompt. The prompt follows the form shown in the README, embedding the description in a sentence and closing it with a </c> token before the labels; the command returns a JSON detections array and can write an annotated PNG via --annotated.

What is the LocateAnything dataset?

The repository does not describe a dataset of that name. It documents the LocateAnything-3B model, its GGUF conversions published at mudler/locate-anything.cpp-gguf, and a download script at scripts/download_model.py that fetches the Hugging Face checkpoint.

Official sources

  1. Issues
  2. mudler/locate-anything.cpp on GitHub
  3. README
  4. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mudler-locate-anything-cpp.svg)](https://hysenlabs.com/projects/mudler-locate-anything-cpp)