Runanywhere SDKs: one C++ core for on-device LLMs, speech and vision
Production ready toolkit to run AI locally
At a glance
- What is it?
- RunanywhereAI's runanywhere-sdks packages eight platform SDKs over a single C++ core, with a capability registry that routes each call to the best engine the device can run. The Python package is the quickest way in; the licence terms are the first thing to check.
- Who is it for?
- Adopt Runanywhere SDKs if you need LLM, speech, vision or RAG inference inside a shipped app on iOS, Android, web or desktop, and the pip install runanywhere path or the Swift and React Native bindings match your stack.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 19 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Runanywhere SDKs actually solve
Shipping a model inside an app usually means choosing a runtime per platform and writing the glue yourself. llama.cpp on iOS, ONNX Runtime on Android, something else in the browser. Runanywhere SDKs is an attempt to collapse that work into one semantic API. The README describes eight SDKs sitting over one C++ core, with the same call surface whether you are writing Swift, Kotlin, React Native, Flutter or web TypeScript.
The scope is wider than chat. The README lists LLM chat with token streaming and LoRA adapters, schema-validated structured output, local tool calling, vision and image Q&A, speech-to-text with Whisper and Moonshine, text-to-speech with Piper, Kokoro, Kitten, MeloTTS and Magpie, embeddings, RAG over local documents, and image generation with Stable Diffusion on Core ML plus inpainting on the Hexagon NPU.
That breadth is the pitch and also the first thing to interrogate. A toolkit that covers speech, vision, retrieval and diffusion in one repository has a lot of surface area, and the README is candid about two gaps: wake-word detection is not implemented, and the computer-use action parser only converts Fara1.5-style action strings into viewport-scaled coordinates rather than acting as a full autonomous agent framework. Read the capability list as a map of what the project intends to cover, not as a promise that every item works on every device.
The capability registry and engine routing
The architectural claim is that your code rarely picks hardware. Engines register what they can run, and the highest-priority engine that fits the device wins. The README names the engines: QHexRT on the Snapdragon Hexagon NPU, MLX on Apple silicon, llama.cpp everywhere with Metal on Apple, CUDA on NVIDIA as an opt-in build and WebGPU in the browser, sherpa plus ONNX for speech and embeddings, and Core ML for diffusion.
This is a priority list, not a benchmark table, and the README does not publish per-engine performance numbers. What it does publish is a discovery call. RunAnywhere.capabilities() in the v4 API reports what the current package and device can actually execute, and the README warns explicitly that enum presence alone does not mean an engine is installed. That warning matters: a capability enum can exist in the bindings while the runtime behind it is absent from your build, which is the kind of mismatch that turns into a runtime failure on a user's phone rather than on your desk.
The README also states that LiteRT and ExecuTorch are reserved framework values only and are not integrated runtimes yet. If your platform strategy depends on either of those, the toolkit does not currently cover it, regardless of what the enum suggests.
Installing Runanywhere and running a first generation
The README gives a Python path as the fastest way to see the toolkit work. Install the package from PyPI:
pip install runanywhereThen initialize the runtime and generate from a small model. The README notes the model downloads on first use, so the first call needs network access even though inference itself is local:
import runanywhere as ra
from runanywhere import LlmOptions
ra.initialize()
# downloads on first use
print(ra.llm.generate("Explain on-device AI in one sentence.",
LlmOptions(model="qwen2.5-0.5b")).text)If you prefer a terminal, the README points to a separate repository, RunanywhereAI/RCLI, which consumes the C++ desktop kit this repository publishes. It gives two install routes, Homebrew and a shell installer, and then a single run command:
brew install runanywhereai/tap/rcli
# or
curl -fsSL https://raw.githubusercontent.com/RunanywhereAI/RCLI/main/install.sh | sh
rcli run qwen3 "Explain on-device AI in one sentence."For mobile and desktop targets the README shows a Swift flow built from three steps: register a backend with LlamaCPP.register(), initialize with try RunAnywhere.initialize(), then load a model through RAModelLoadRequest with modelID, category and framework set, and generate through RALLMGenerateRequest. The README adds that importing RunAnywhereMLX and calling MLX.register() enables the Apple-native backend for LLM, VLM, STT, TTS and embeddings on Apple silicon. Installation for Swift is via Swift Package Manager; the README's URL is truncated in the excerpt, so take the package URL from the repository's Package.swift rather than from the README text.
Voice agents, RAG and the parts that are gated
The voice pipeline is VAD, STT, LLM and TTS in one chain, with playback scoped through what the README calls SpeechHandle. If you are building a hands-free assistant, note the boundary again: wake-word detection is not implemented, so the trigger has to come from your application code, not from the SDK.
RAG is described as local document ingestion and retrieval-augmented answers with streaming, and embeddings are L2-normalized vectors intended for search and retrieval. The README does not describe chunking strategy, index persistence or how ingestion scales with document count, so those are questions for the documentation site rather than the README.
Image generation is the most conditional capability. Stable Diffusion is listed on Core ML, and inpainting is listed on the Hexagon NPU with an explicit note that it is platform and backend gated. Gated means exactly what it sounds like: whether you get it depends on the device and the engine that registered. Treat diffusion support as a per-device question you answer with capabilities(), not as a feature you assume from the README feature list.
Where the toolkit is the wrong choice
The licence is the first real constraint. The README badge reads RunAnywhere License, and the GitHub licence field for the repository reports NOASSERTION. That combination means the terms are not a standard OSI licence that a tool can identify. If your organisation requires a recognised open source licence, or if you need to know redistribution terms before embedding a runtime in a shipped app, read the LICENSE file itself. Nothing in the README summarises what the licence permits.
Second, the engine coverage is narrower than the topic list suggests. CUDA on NVIDIA is described as an opt-in build, not a default. LiteRT and ExecuTorch are reserved values with no integrated runtime. If your deployment target is a server with a specific accelerator, check whether an engine actually registers there before designing around the toolkit.
Third, the repository layout signals a monorepo with several independent toolchains. The root package.json declares a Yarn 3.6.1 workspace covering the TypeScript proto bindings, the React Native packages and the React Native example, and its own note explains that three separate workspace roots resolve the same proto-ts package by different relative paths, with a stated future consolidation once the npm-versus-yarn lockfile policy is decided. That is an honest note, and it also tells you the build story has seams. If you only need one platform, you are pulling in a repository organised around eight.
How it compares with picking one runtime
The obvious alternative is to pick a single inference runtime and wire it yourself. llama.cpp is the clearest comparison because the README lists it as one of the engines this toolkit routes to. Using llama.cpp directly gives you one build, one set of bindings, and full control over model loading, threading and memory, with a much smaller dependency surface. What you give up is the routing layer: you decide the backend, you handle Metal or CUDA configuration, and you build the speech, embedding and diffusion pieces separately or not at all.
Runanywhere's difference is the abstraction over engines plus the breadth of capabilities behind one API. That trade is worth it when you are shipping across several platforms and want STT, TTS and embeddings from the same call surface as chat. It is worth less when you need exactly one capability on exactly one platform, because then the registry, the proto bindings and the multi-workspace layout are overhead rather than leverage.
A second comparison point is the model source. The README links a Hugging Face organisation for models, which suggests the intended path is pulling pre-converted models rather than converting your own. If you have a fine-tuned checkpoint in a format none of the named engines consumes, the toolkit does not shorten that conversion step.
Maintenance, releases and what upgrading costs
The repository is not archived, and the last push was on 2026-09-08. Recent releases are frequent and versioned: cpp-desktop-v0.20.37 on 2026-09-07, v0.20.36 on 2026-09-02, and v0.20.35 earlier the same day. The 0.20.x line is pre-1.0, so treat minor version bumps as potentially breaking and read the release notes rather than assuming compatibility.
The release naming is worth noticing. cpp-desktop-v0.20.37 is labelled as cpp-desktop kits, which implies the desktop kit is versioned separately from the umbrella v0.20.36 tag. If you depend on the CLI or the C++ desktop kit, track that lane's tags rather than the repository-wide ones.
The upgrade cost is concentrated in the bindings. With eight SDKs over one core, a change in the core's interface can propagate to Swift, Kotlin, React Native, Flutter, web and Python at once, and the proto-ts workspace is shared across the React Native and web lanes. Pin versions, and check the capability surface after upgrading rather than assuming an engine you relied on still registers. On licensing, the repository ships a LICENSE file and a SECURITY.md, but the README does not state what the RunAnywhere License permits for redistribution or commercial embedding, so that is a question for the licence text and, if it matters commercially, for your own counsel.
Editorial conclusion
Adopt Runanywhere SDKs if you need LLM, speech, vision or RAG inference inside a shipped app on iOS, Android, web or desktop, and the pip install runanywhere path or the Swift and React Native bindings match your stack. Do not adopt it if you need wake-word detection, a full autonomous computer-use agent, LiteRT or ExecuTorch runtimes, or licence terms you can read without a lawyer, since the repository badge says RunAnywhere License and the GitHub licence field reports NOASSERTION. Verify three things before you commit: the exact terms in the LICENSE file, that RunAnywhere.capabilities() reports an engine for the capability you need on your target device, and that the model you intend to ship is available from the Hugging Face model collection the README links.
Frequently asked questions
What does the Python SDK in Runanywhere SDKs do?
The README shows the Python package as the quickest way to run a model locally: install it with pip install runanywhere, call ra.initialize(), then call ra.llm.generate with an LlmOptions model such as qwen2.5-0.5b. The README notes the model downloads on first use.
What are SDK examples in Runanywhere SDKs?
The README gives three worked examples: a Python snippet that initializes the runtime and generates text, a Swift flow using LlamaCPP.register, RAModelLoadRequest and RALLMGenerateRequest, and a terminal path through the separate RunanywhereAI/RCLI repository with rcli run qwen3. The repository also contains a React Native example workspace under bindings/react-native/example.
What does SDK mean in Android for Runanywhere SDKs?
In this project an Android SDK is one of the eight platform bindings over the single C++ core described in the README, alongside Swift, Kotlin, React Native, Flutter, web and Python. The README states that the same semantic API is used across those platforms, and that RunAnywhere.capabilities() reports which engines the current package and device can execute.
How do SDKs work in Runanywhere SDKs?
The README describes eight SDKs over one C++ core, with a capability registry that routes each call to an engine the device can run, such as QHexRT on the Hexagon NPU, MLX on Apple silicon, llama.cpp, sherpa plus ONNX, or Core ML. RunAnywhere.capabilities() reports what the current package and device can actually execute.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/runanywhereai-runanywhere-sdks)