Model or dataset
mybigday/llama.rn avatar
mybigday/llama.rn

llama.rn: Running llama.cpp LLM Inference Inside React Native Apps

React Native binding of llama.cpp

1,047 stars125 forksC++MIT

At a glance

What is it?
llama.rn is a React Native binding of llama.cpp that runs large language model inference on-device, using Metal on iOS and OpenCL or the Hexagon NPU on Android. It requires the New Architecture and handles model initialization, chat completion, tokenization, and embedding inside the mobile process.
Who is it for?
llama.rn is the practical choice when you need to run an LLM inference inside a React Native app without sending data to an external API. The binding handles Metal on iOS and OpenCL or Hexagon NPU on Android.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What llama.rn Makes Possible on Mobile

Running language model inference on a mobile device means no network call to an inference server and no data leaving the device. llama.rn delivers this by wrapping llama.cpp, the C/C++ LLM inference library, as a native module for React Native. The binding exposes model initialization, chat completion, tokenization, detokenization, embedding, and reranking as JavaScript-callable async functions.

The library targets both iOS and Android. On iOS it uses Metal for GPU acceleration. On Android it supports OpenCL for Qualcomm Adreno 700 series and newer GPUs, and an experimental Hexagon NPU path for Qualcomm SM8450 and newer SoCs. Both Android acceleration paths require specific entries in the app manifest and setting n_gpu_layers to a positive value in the initLlama call.

Beyond text completion, the library supports multimodal input through mmproj projector integration, on-device text-to-speech through codec.cpp for a range of TTS models including OuteTTS and CSM, slot-based parallel decoding for concurrent request handling, tool calling via Jinja templates, and grammar sampling with GBNF and JSON schema for constrained output.

The current release is version 0.13.0-rc.6, published on 2026-09-26.

New Architecture Requirement and Expo Integration

Since version 0.10, llama.rn requires React Native's New Architecture. The README marks this as important. Apps still on the Old Architecture must migrate to the New Architecture or remain on the v0.9 branch, which the README links directly for teams that cannot upgrade.

For Expo projects using CNG builds, the expo-build-properties plugin is needed to enable iOS and OpenCL features. The app.json or app.config.js configuration looks like this:

js
module.exports = {
  expo: {
    plugins: [
      [
        'llama.rn',
        {
          enableEntitlements: true,
          entitlementsProfile: 'production',
          forceCxx20: true,
          enableOpenCL: true,
        },
      ],
    ],
  },
}

The README lists these as the default values, so this block is optional if you want default behavior.

For Bun users, lifecycle scripts do not run automatically unless the package is trusted. The README provides the trustedDependencies entry to add to package.json, or an alternative manual command to run the downloader before building.

Installing and Initializing a Model

Installation from npm is one command:

sh
npm install llama.rn

The postinstall script downloads prebuilt native binaries for iOS (rnllama.xcframework) and Android (jniLibs) from the matching GitHub release. The download is verified with SHA-256 before extraction, and existing downloads are reused.

After installation, iOS requires running pod-install again:

code
npx pod-install

For Android, if ProGuard is enabled, add the keep rule to the ProGuard configuration file:

proguard
-keep class com.rnllama.** { *; }

To initialize a model and run a completion, the main API entry point is initLlama:

js
import { initLlama } from 'llama.rn'

const context = await initLlama({
  model: modelPath,
  use_mlock: true,
  n_ctx: 2048,
  n_gpu_layers: 99,
})

Setting n_gpu_layers to 99 requests that all layers be offloaded to the GPU. The returned context object exposes a gpu boolean, a reasonNoGPU string, and a devices array, which let you confirm at runtime whether GPU acceleration is active. The README notes this is useful for verifying Android OpenCL behavior on actual devices.

Chat Completion, Tokenization, and Embeddings

The context object returned by initLlama mirrors the server.cpp API from llama.cpp. The completion method handles both standard completion and chat completion:

js
const stopWords = ['</s>', '<|end|>', '<|eot_id|>', '<|end_of_text|>',
  '<|im_end|>', '<|EOT|>', '<|END_OF_TURN_TOKEN|>', '<|end_of_turn|>',
  '<|endoftext|>']

const msgResult = await context.completion(
  {
    messages: [
      { role: 'system', content: 'You are a helpful assistant.' },
      { role: 'user', content: 'What is React Native?' },
    ],
    stop: stopWords,
  },
  (partial) => { /* handle streaming token */ }
)

The second argument to completion is a callback that receives partial tokens as they are generated, enabling streaming output in the UI.

Additional context methods include tokenize for converting text to token ids, detokenize for the reverse, embedding for generating vector embeddings, and rerank for reranking documents against a query. The README directs users to the docs/API directory for the full method reference.

MTP Speculative Decoding for Supported Models

For GGUF models that include MTP or NextN draft layers, MoonEP-style speculative decoding can be enabled at context initialization:

js
const context = await initLlama({
  model: modelPath,
  n_ctx: 4096,
  n_batch: 1024,
  n_ubatch: 512,
  n_gpu_layers: 99,
  flash_attn_type: 'auto',
  cache_type_k: 'q8_0',
  cache_type_v: 'q8_0',
  speculative: {
    type: 'draft-mtp',
    n_max: 3,
  },
})

Speculative decoding lets the model draft multiple candidate tokens ahead and verify them in a single forward pass, which improves generation throughput on capable hardware. The n_max parameter sets how many draft tokens to attempt per step. Passing speculative: false on an individual completion call disables MTP for that request without reinitializing the context.

The README notes that for recurrent or hybrid models, MTP must be enabled at initLlama time with a positive spec_draft_n_max because llama.cpp needs to allocate rollback state at initialization.

Where llama.rn Has Real Constraints

The hardware requirements are significant. Metal acceleration on iOS requires a device with a Metal-capable GPU, which covers all modern iPhones and iPads. OpenCL acceleration on Android currently requires Qualcomm Adreno 700 series or newer, and the Hexagon NPU path requires Qualcomm SM8450 or newer (8 Gen 1 or newer). Devices with other GPU vendors or older Snapdragon chips are limited to CPU inference.

The model must be in GGUF format. The README points to HuggingFace for GGUF models and to the llama.cpp quantize documentation for converting models. Managing model files on mobile devices, including the download process, the storage location, and the memory footprint at runtime, is the developer's responsibility. The README does not provide a built-in model manager.

The New Architecture requirement is a hard constraint. React Native's New Architecture requires a migration that touches native module code and may not be straightforward for large existing apps. The v0.9 branch remains available for Old Architecture users, but it will not receive new features.

The binding's design is inspired by server.cpp in llama.cpp, so the parameter names and concepts align with that project. Teams unfamiliar with llama.cpp configuration parameters like n_batch, n_ubatch, cache_type_k, and flash_attn_type will need to consult the llama.cpp documentation.

Maintenance Status and License

The last push to the repository was on 2026-09-27, and three release candidates for version 0.13.0 were published in September 2026: rc.4 on the 17th, rc.5 on the 23rd, and rc.6 on the 26th. The project is under active development.

The library is published under the MIT license. The vendor directory in the repository contains the llama.cpp source, which has its own license. Teams using llama.rn in a shipped product should review the llama.cpp license separately, as it governs the C++ inference code that does the actual work.

The repository includes documentation in docs/API generated by TypeDoc from the TypeScript source, and a complete example application in the example directory. The CONTRIBUTING.md documents how to set up a local development build and how to run the test suite via jest.

Editorial conclusion

llama.rn is the practical choice when you need to run an LLM inference inside a React Native app without sending data to an external API. The binding handles Metal on iOS and OpenCL or Hexagon NPU on Android. The key prerequisite is React Native's New Architecture, which has been required since version 0.10: apps still on the Old Architecture must either migrate or stay on the v0.9 branch. Before shipping a production app, verify the GGUF model size fits within the device memory budget, test on physical hardware rather than the emulator for GPU path verification, and check the n_gpu_layers setting against the result's gpu and reasonNoGPU fields to confirm the GPU is actually being used.

Frequently asked questions

what is llama rn

llama.rn is a React Native package that wraps llama.cpp, the C/C++ LLM inference library, as a native module. It lets React Native apps run large language model inference on-device, with GPU acceleration via Metal on iOS and OpenCL or the Hexagon NPU on Android.

Does llama.rn support streaming output?

The context.completion method accepts a second callback argument that is called with each partial token as it is generated. This enables streaming output in the UI without buffering the full response.

What model format does llama.rn require?

Models must be in GGUF format. The README points to HuggingFace for available GGUF models using the search keyword GGUF, and to the llama.cpp quantize documentation for converting models from other formats.

Official sources

  1. Issues
  2. License: MIT
  3. mybigday/llama.rn on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mybigday-llama-rn.svg)](https://hysenlabs.com/projects/mybigday-llama-rn)