# React Native AI: on-device LLM inference with a Vercel AI SDK front end

> Callstack's react-native-ai packages wrap Apple Foundation Models, llama.rn and MLC LLM behind the AI SDK's generateText and streamText calls. The Apple provider ships with the OS; the other two download GGUF or MLC weights at runtime.

**callstackincubator/ai** — On-device LLM execution in React Native with Vercel AI SDK compatibility

- Repository: https://github.com/callstackincubator/ai
- Website: https://react-native-ai.dev
- Stars: 1,406 · Forks: 59
- Language: TypeScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/callstackincubator-ai

## What react-native-ai solves, and for whom

The pitch is inference that never leaves the handset. The README frames it as "privacy-preserving, low-latency inference without server costs", which is a real architectural difference rather than a marketing one: no API key ships in the bundle, no request leaves the device, and there is no per-token bill. For an app that transcribes a voice note or embeds a document the user just typed, that removes an entire class of compliance questions.

The audience is narrower than "React Native developers". It is developers who already use the Vercel AI SDK's generateText, streamText, embed and transcribe functions and want the model argument to point at the device instead of a hosted endpoint. The README calls this a "drop-in replacement with familiar APIs", and the code samples bear that out: the call sites are unchanged, only the model object differs. If you have never used the AI SDK, you are adopting two things at once, and the AI SDK's own abstractions will be the steeper part of the learning curve.

## Three providers, three very different runtimes

The repository is a monorepo of packages under packages/, each exposing a provider object. They are not interchangeable, and the differences matter more than the shared API surface.

Apple is the only built-in provider. It wraps Apple Foundation Models, NLContextualEmbedding, SpeechAnalyzer and AVSpeechSynthesizer, and it is iOS-only. No weights are downloaded because the models belong to the operating system. Availability is version-gated: text generation and transcription need iOS 26 or later plus an Apple Intelligence device, embeddings need iOS 17 or later, and speech synthesis works from iOS 13 with Personal Voice requiring iOS 17.

Llama wraps llama.rn and runs GGUF files pulled from HuggingFace. It works on iOS and Android, and it is the provider to pick when you need a specific open-weight model rather than whatever Apple ships. MLC wraps the MLC LLM runtime, also on both platforms, and its package has a postinstall step: the root package.json runs bash packages/mlc/scripts/fetch-prebuilt.sh after install, and the README notes MLC requires the "Increased Memory Limit" capability in Xcode.

That last detail is the honest signal about the whole category. On-device inference is bounded by the device, and the build configuration has to acknowledge it.

## Installing @react-native-ai/apple and generating your first completion

Start with the Apple provider if you are on iOS, because it needs no model download and no extra native modules. Install it alongside the AI SDK:

```bash
npm install @react-native-ai/apple
```

The README states no additional linking is needed and the package is autolinked. Then call generateText with apple() as the model. The import below is copied from the README's usage sample:

```typescript
import { apple } from '@react-native-ai/apple'
import { generateText } from 'ai'

const { text } = await generateText({
  model: apple(),
  prompt: 'Explain quantum computing',
})
```

If the model is available on the device, text holds the completion. If it is not, this is where you find out: the README's availability table puts text generation behind iOS 26+ and an Apple Intelligence device, so a simulator or an older handset is the wrong place to test. The same package exposes apple.textEmbeddingModel(), apple.transcriptionModel() and apple.speechModel() for the other three capabilities.

## Running a GGUF model with the Llama provider

The Llama path is more work but gives you a model you chose. Install three packages, not one:

```bash
npm install @react-native-ai/llama llama.rn react-native-blob-util
```

Model IDs follow the format owner/repo/filename.gguf. The README's example loads a small SmolLM3 build, downloads it with a progress callback, then prepares it in memory before any generation happens:

```typescript
import { llama } from '@react-native-ai/llama'
import { generateText } from 'ai'

const model = llama.languageModel(
  'ggml-org/SmolLM3-3B-GGUF/SmolLM3-Q4_K_M.gguf'
)

await model.download((progress) => {
  console.log(`Downloading: ${progress.percentage}%`)
})
await model.prepare()
```

Download and prepare are separate steps, and that separation is the design decision worth noticing. Downloading is a network operation you can show progress for; prepare() loads weights into memory and is the point where a device without enough RAM will fail. The README also shows model.unload() for cleanup, which implies you are expected to manage the model's lifetime rather than leave it resident. Any GGUF file from HuggingFace works, so the memory ceiling is yours to reason about, not the library's.

## The AI SDK version split is the sharpest edge

The compatibility table is short and unforgiving: react-native-ai 0.11 and below target AI SDK v5, and 0.12 and above target v6. There is no overlap. An app pinned to AI SDK v5 that upgrades the provider to 0.12 is mixing two incompatible generations of the same API, and the failure will surface as type errors or runtime mismatches in the model interface rather than as a clear message about version skew.

The latest release listed is v0.12.0 from 2026-01-28, with v0.11.0 in October 2025 and v0.10.0 in September 2025 before it. The repository's last push was on 2026-07-07, so work has continued past the last tagged release. Treat the version pairing as a constraint you check before anything else, and read the release notes for the major you are moving to rather than assuming the provider API is stable across the boundary.

## Profiling on-device calls with the AI SDK Profiler

Debugging inference that happens inside the app is awkward, because there is no server log to read. The @react-native-ai/dev-tools package addresses that: it captures OpenTelemetry spans from Vercel AI SDK requests and shows them in a Rozenite DevTools panel called AI SDK Profiler.

```bash
npm install @react-native-ai/dev-tools
```

Rozenite must already be installed and enabled in the app; the README points at the Rozenite getting started guide for that. The Expo demo under apps/expo-example contains the native-development wiring for the plugin, which is the reference to copy. DevTools are runtime agnostic here, so the same panel covers on-device and remote runtimes, which is useful while you are deciding between them.

The README documents one specific failure: if the AI SDK Profiler panel is visible but stays empty after you send a chat message, close the React Native DevTools window and open a fresh one, because a stale debugger session can keep the panel mounted without receiving the current app's telemetry. That is a workaround, not a fix, and it is the kind of detail that tells you the tooling is younger than the providers.

## Where on-device inference is the wrong tool

The Apple provider's version floor is the first disqualifier. Text generation requires iOS 26 or later on an Apple Intelligence device, so an app with a meaningful population below that line cannot rely on it, and there is no fallback described in the README for those devices. You would need the Llama or MLC provider, which means a download.

Downloads bring the second problem. A GGUF model is a file the user pays for in bandwidth and storage, and prepare() then holds it in memory. The README does not document a size limit, a memory budget, or what happens when prepare() fails on a constrained device, so that analysis is yours. MLC pushes the same concern into the build: the README requires the Increased Memory Limit capability in Xcode, which is an admission that the default ceiling is too low.

Finally, on-device is the wrong shape for work that needs a frontier model. A three-billion-parameter GGUF file running on a phone is not going to match a hosted large model on hard reasoning, and no amount of API compatibility changes that. If your feature's quality depends on the biggest available model, run it on a server and use this project for the parts that are genuinely local: transcription, embeddings, short completions.

## Alternatives and the difference in approach

The obvious alternative is calling a hosted model through the AI SDK's own provider packages. The call sites look almost identical, which is the point of the AI SDK, but the trade is inverted: you get larger models and no device constraints, and you give up the privacy guarantee, the offline behaviour, and the absence of per-token cost. The README's framing of on-device inference exists precisely because that trade is not always acceptable.

Within the on-device space, the real choice is between the three providers here rather than between this project and something else. Apple gives you zero download and zero model choice, and it is iOS-only. Llama gives you any GGUF file on both platforms, at the cost of a download and manual memory management. MLC gives you MLC's optimized runtime and its prebuilt fetch step, with a documented Xcode capability requirement. Picking between them is a decision about which constraint you can live with, not about which is better.

The licence is MIT, which is permissive and imposes no copyleft obligation on your app. That covers the packages in this repository; the models you download carry their own licences from HuggingFace, and the README does not discuss them, so check each model separately before shipping.

## Conclusion

Adopt react-native-ai if your app already speaks the Vercel AI SDK and you want inference on the device rather than behind an API key. It is the wrong choice if you must support Android with no model download, or if your product depends on iOS versions below the Apple provider's floor (text generation needs iOS 26+). Before committing, verify three things on a physical device: that the Apple model actually reports available on your target hardware, that a GGUF or MLC model of the size you need fits in memory after llama.rn is linked, and which AI SDK major your installed version of the package expects, because 0.12 and above target v6 while 0.11 and below target v5.

## FAQ

### Can React Native be used with AI?

Yes. React Native AI provides on-device AI primitives for React Native with Vercel AI SDK support, covering text generation, embeddings, transcription and speech synthesis, so AI calls can run on the user's device rather than on a server.

### How do I install React Native AI?

It depends on the provider. The Apple provider installs with npm install @react-native-ai/apple and needs no linking; the Llama provider needs npm install @react-native-ai/llama llama.rn react-native-blob-util; the MLC provider installs with npm install @react-native-ai/mlc and requires the Increased Memory Limit capability in Xcode.

### Which AI SDK version does React Native AI require?

The README's compatibility table maps react-native-ai 0.11 and below to AI SDK v5, and 0.12 and above to AI SDK v6. The two ranges do not overlap, so the version you install on each side has to match.

### Does React Native AI work on Android?

The Llama and MLC providers list iOS and Android as supported platforms. The Apple provider is iOS only, since it wraps Apple Foundation Models and related system frameworks.

### Why is the AI SDK Profiler panel empty?

The README states that a stale debugger session can keep the Rozenite panel mounted without receiving the current app's telemetry stream. Closing the React Native DevTools window and opening a fresh one is the documented workaround.

## Sources

- [callstackincubator/ai on GitHub](https://github.com/callstackincubator/ai)
- [License: MIT](https://github.com/callstackincubator/ai/blob/main/LICENSE)
- [Project website](https://react-native-ai.dev)
- [README](https://github.com/callstackincubator/ai/blob/main/README.md)
- [Releases](https://github.com/callstackincubator/ai/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/callstackincubator-ai
