Cactus by cactus-compute: an on-device inference engine for phones, wearables and robots
Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots.
At a glance
- What is it?
- Cactus is a C++ engine that runs LLMs, vision models and speech models locally on Apple and Android hardware, with an optional confidence-based handoff to the cloud. Its headline trade-off is the CQ2 quantization tier, which collapses reasoning scores.
- Who is it for?
- Adopt Cactus if you are shipping an Apple or Android app that needs text, speech or vision inference on the device and you want a single C API with OpenAI-shaped payloads. Do not adopt it if you are targeting desktop servers, NVIDIA GPUs, or if your product needs reliable multi-step reasoning at 2-bit weights: the published CQ2 numbers put GSM8K at 0.40 against 73.67 for F16.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What Cactus actually replaces on a mobile codebase
Running a transformer on a phone normally means stitching together a runtime, a kernel library, a quantization pipeline and a chat template layer, then writing platform glue for Metal on Apple and something else on Android. Cactus presents itself as one stack for that: an engine with an OpenAI-compatible surface, a graph layer, ARM kernels, and its own quantization format. The README describes it as "a hybrid edge-cloud AI engine for mobile devices & wearables," and the repository layout backs that up, with android/, apple/, bindings/ for Swift, Kotlin, Flutter, React Native, Python and Rust, plus separate cactus-engine/, cactus-graph/ and cactus-kernels/ directories.
The intended user is an application developer, not a model researcher. The Quick Demo is two commands on a Mac, and the C API takes a folder of weights and a JSON message array. If your team already has a Python training pipeline and a serving stack, Cactus is not competing with that. It is competing with the option of not running the model locally at all.
The four-layer stack, from quantization up to the API
The README draws the architecture as four stacked boxes. At the bottom, Cactus Quants implements what it calls a "custom rotation-based quantization technique" covering 4-bit down to 1-bit. Above that, Cactus Kernels provides CPU and GPU kernels named for Apple, Samsung and Pixel hardware. Cactus Graph sits above the kernels as a "zero-copy computation graph" with tensor ops, matmul, attention and normalization. Cactus Engine is the top layer, exposing text, speech and vision through OpenAI-compatible APIs.
Data flow through the graph layer is explicit. You declare inputs with a shape and a precision, chain operations, write host buffers into the graph, call execute, and read a void pointer back out. The example mixes precisions in one graph: input a is FP16, input b is INT8, and the matmul that consumes them is declared without a precision argument. That is the mechanism worth noting, because mixed-precision graphs are where a lot of hand-written mobile code gets messy. The graph also exposes hard_reset(), which suggests the intended usage is repeated execution cycles rather than one-shot inference.
At the engine layer the contract is JSON in, JSON out. cactus_init takes a weight folder, an optional path to text files for automatic RAG, and a boolean. cactus_complete takes that handle, a JSON chat array, a response buffer, an options object, a tools JSON, a streaming callback, a PCM audio buffer and its size. The response object carries success, error, cloud_handoff, response text, function_calls, segments, confidence, confidence_threshold, timing fields and ram_usage_mb.
Installing Cactus and running a first completion
The README documents a Homebrew path for macOS. After the tap install, the cactus CLI is on your PATH and the bare run command starts the demo, which per the README downloads or converts a model if it is not already present.
brew install cactus-compute/cactus/cactus
cactus runTo pick a specific model instead of the default demo, pass a Hugging Face name. The README notes that some models are pre-uploaded under the Cactus-Compute organization, and that run will download or convert the model if it is not found locally.
cactus download Cactus-Compute/needle
cactus run Cactus-Compute/needleNeedle is the project's own 26m parameter model for on-device tool calling. It accepts a tools file in OpenAI function-calling format, and the README says a demo toolset is used by default when none is given.
cactus run Cactus-Compute/needle --tools my_tools.jsonOn the C side, the first real call is cactus_init followed by cactus_complete. The README's example passes a weight folder path, a path to text or a directory of texts for auto-RAG, and false for the third argument. The response buffer is a fixed char array, so the caller decides the ceiling before generation starts.
#include "cactus_engine.h"
cactus_model_t model = cactus_init(
"path/to/weight/folder",
"path to txt or dir of txts for auto-rag",
false
);The README does not document what the third boolean controls, and it does not show error handling around cactus_init. Treat both as gaps you will have to close by reading the header.
Quantization tiers are where the real decision lives
Cactus publishes an accuracy table for Gemma-4-E2B-it across bit widths, averaged over three seeds. The pattern is not uniform decay. CQ4 is close to F16 on most tasks, and on HumanEval it actually scores higher, 57.11 against 54.88. CQ3.26, a mixed-precision tier, holds ARC-E slightly above F16 at 74.20 and keeps HumanEval at 53.66, but GSM8K drops from 73.67 to 66.20.
CQ2.54 is where the table turns. GSM8K falls to 22.00 and HumanEval to 15.24. At uniform CQ2, GSM8K is 0.40 and HumanEval is 1.02. Tool-calling degrades on the same curve: BFCL Simple stays at 92.42 under CQ4 and 91.50 under CQ3.26, then falls to 82.25 at CQ2.54 and 18.75 at CQ2. BFCL Parallel-Multi goes from 83.33 at CQ4 to 1.33 at CQ2.
The honest reading is that CQ4 is the safe default and CQ3.26 is defensible for chat and retrieval-style work. Anything at or below CQ2.54 should be treated as a different product: it will answer questions and classify text, but it will not do arithmetic or write code. The README does not offer guidance on choosing a tier, so the table is the only signal you get.
Cloud handoff and the confidence threshold
Cactus Hybrid is the feature that distinguishes this from a pure local runtime. The engine returns a confidence value and a confidence_threshold, described in the README as the "resolved handoff threshold (model-dependent)," plus a cloud_handoff boolean. In the sample response from Gemma4-E2B, confidence is 0.8193 against a threshold of 0.7, so the request stayed local.
The design implication is that you need a cloud endpoint configured for the fallback path to mean anything. The README does not document which providers are supported, how credentials are supplied, or what happens to a request when the cloud call fails. That is a substantial omission for a feature whose entire purpose is routing production traffic. If you plan to rely on handoff, the failure semantics are the first thing to establish from the hybrid documentation rather than from the README.
There is also a latency argument against handoff for interactive use. The sample response shows time_to_first_token_ms of 45.23 and total_time_ms of 163.67 for a 50-token generation. A round trip to a cloud model will not beat that, so handoff only pays off when the local answer would have been wrong.
Benchmark numbers and what they do not cover
The README publishes a device table generated by cactus benchmark, with optional --ios and --android flags. On Mac M5 Max it reports 2964 tps prefill and 154 tps decode for LLM, 0.09s image encode and 168 tps decode for VLM, 0.15s for a 20-second transcription, and 1348MB peak RAM at 1k context. On iPhone 15 Pro the same row reads 517 tps prefill, 26 tps decode, 1.15s image encode, 0.82s transcription, 633MB RAM.
Two caveats are stated in the README itself. The LLM figures use a 1k-context prefill and decode for 100 tokens, and the project notes "No speculative decode or MTP, pure decode." That means these are raw decode rates, not the accelerated numbers you would see from a runtime using speculative decoding. The RAM column is peak during the LLM benchmark only, so it does not tell you the footprint of a vision or speech pipeline.
The table reports Mac and iPhone rows. There is no Android device row in the README, even though cactus benchmark accepts --android and the repository contains an android/ directory. If Android is your target, you have no published reference point and will have to produce your own.
Where Cactus is the wrong tool
The strongest case against Cactus is model coverage. The README states that any Hugging Face model can be converted with cactus convert, but immediately labels that path experimental. The families described as especially tested are Liquid, Gemma, Whisper, Parakeet and Qwen. If your product depends on a model outside those families, you are on the experimental path, and the README gives no success rate, no list of unsupported architectures and no error taxonomy for failed conversions.
Server-side deployment is the second mismatch. Every published benchmark is Apple silicon or iPhone, the kernels are described as ARM NEON SIMD, and the demo installs through Homebrew. Nothing in the README describes a CUDA path or a Linux server deployment. Choosing Cactus for a data-center workload means fighting the project's stated scope.
The third case is reasoning-heavy work at small bit widths. If your application needs arithmetic, code generation or multi-step tool calls, the accuracy table rules out the smallest tiers, and you should budget for CQ4 memory rather than assuming you can shrink to CQ2 later. The README does not document rollback between quantization tiers, so changing your mind after shipping means re-converting and re-testing.
Alternatives and the difference in approach
The topics list on the repository includes llamacpp, which points at the most direct comparison. llama.cpp is a general-purpose C/C++ inference project built around the GGUF format, with a broad backend story spanning CPU, CUDA, Metal and Vulkan, and a large community of converted models. Cactus is narrower by design: ARM kernels for mobile silicon, its own CQ quantization family rather than GGUF, and a graph layer plus an OpenAI-shaped engine API rather than a single inference binary.
The practical difference shows up in two places. First, model availability: llama.cpp benefits from a wide ecosystem of pre-converted GGUF files, while Cactus maintains its own Hugging Face organization and treats arbitrary conversion as experimental. Second, the API shape: llama.cpp gives you a generation loop, while Cactus gives you JSON in and a JSON response object out, with confidence, timing and RAM fields already populated. If you want a chat endpoint inside an app and you do not want to build the surrounding plumbing, Cactus does more for you. If you need a model that only exists as GGUF, or you need a backend other than ARM and Apple, Cactus is the wrong side of that trade.
For speech specifically, the README lists Parakeet-TDT-0.6B-CQ4 as the transcription benchmark model, so the speech path is coupled to the same quantization and kernel stack rather than being a separate integration.
Editorial conclusion
Adopt Cactus if you are shipping an Apple or Android app that needs text, speech or vision inference on the device and you want a single C API with OpenAI-shaped payloads. Do not adopt it if you are targeting desktop servers, NVIDIA GPUs, or if your product needs reliable multi-step reasoning at 2-bit weights: the published CQ2 numbers put GSM8K at 0.40 against 73.67 for F16. Before committing, run cactus benchmark on your own target device, check the LICENSE file directly because GitHub reports NOASSERTION, and confirm whether the model you need is already on the Cactus-Compute Hugging Face account or has to go through cactus convert.
Frequently asked questions
How do I install Cactus on macOS?
The README gives a two-step Quick Demo for Mac: install through the Homebrew tap with brew install cactus-compute/cactus/cactus, then run cactus run. The run command downloads or converts a model if one is not already present.
How do I use Cactus to run a model?
Pass a Hugging Face model name to the CLI. The README states that cactus run [HF-Name] downloads or converts the model if it is not found, and that some models are pre-uploaded under the Cactus-Compute organization for cactus download. The project's own tool-calling model, Needle, runs with cactus run Cactus-Compute/needle.
What hardware does the Cactus engine support?
The published benchmark table covers Mac M5 Max, M4 Pro, M3 Pro, iPad and Vision Pro M5, iPhone 17 Pro and iPhone 15 Pro. The kernels are described as ARM NEON SIMD kernels for Apple, Samsung and Pixel hardware, and the bindings cover Swift, Kotlin, Flutter, React Native, Python and Rust.
What is the difference between CQ4, CQ3.26 and CQ2 in Cactus?
They are quantization tiers of the Cactus Quants format. The README's accuracy table for Gemma-4-E2B-it shows CQ4 close to F16 on most tasks, CQ3.26 still holding ARC-E above F16 but dropping GSM8K to 66.20, and CQ2 collapsing GSM8K to 0.40 and HumanEval to 1.02. CQ3.26 and CQ2.54 are mixed-precision; CQ2, CQ3 and CQ4 are uniformly quantized.
Does Cactus send requests to the cloud?
Only when the hybrid path decides to. The engine response includes a cloud_handoff boolean, a confidence value and a confidence_threshold described as the resolved handoff threshold, which is model-dependent. The README does not document which cloud providers are supported or how credentials are configured.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/cactus-compute-cactus)
Community notes