# Xybrid pins one version into five package registries and two build systems

> Xybrid is a Rust toolkit for running language models, speech recognition, and speech synthesis inside an app, with llama.cpp and whisper.cpp wrapped as workspace crates and bindings published to pub.dev, Maven Central, npm, Swift Package Manager, and OpenUPM. The interesting part sits underneath the samples: a Cargo workspace and a Bazel graph that maintain overlapping test targets, five patch files against third party build rules, and a dev profile that optimizes every dependency.

**xybrid-ai/xybrid** — Cross-platform on-device AI toolkit

- Repository: https://github.com/xybrid-ai/xybrid
- Website: https://xybrid.ai
- Stars: 466 · Forks: 51
- Language: Rust
- License: Apache-2.0
- Published: 2026-09-15 · Updated: 2026-09-15 · Language: en
- Canonical page: https://hysenlabs.com/projects/xybrid-ai-xybrid

## Fourteen workspace members wrap two inference engines

The Cargo.toml in the repository root is a resolver 2 workspace with fourteen members, and their names describe the layering in order. The two `-sys` crates, `llama-cpp-sys` and `whisper-cpp-sys`, are the raw C bindings. `xybrid-llama` and `xybrid-whisper` sit on top of them and are the two engines you actually call. `xybrid-core` holds the shared types, `xybrid-sdk` the high level surface, `xybrid-ffi-facade` the flat surface the language bindings talk to, `xybrid-bolt` and `xybrid-cli` the tools, `macros` the proc macros, `xtask` the build scripts, and `integration-tests` the cross crate checks. Exactly one binding lives in the workspace: `bindings/flutter/rust`.

The shared dependency list is short for something that ships inference. serde, serde_json, serde_yaml, anyhow, thiserror, and tokio, with tokio asked for precisely three features: `rt`, `rt-multi-thread`, and `sync`. There is no HTTP client and no networking crate in that list, which lines up with the promise in the README that Xybrid runs private and offline with no cloud required. That part is checkable rather than a slogan: a feature that needed to reach a server would have to introduce its own dependency, and the manifest shows none.

One key tells you which crate the maintainers want read. The `documentation` field points at docs.rs for `xybrid-core`, not for the SDK or the CLI, so the core crate is the one intended as the public reference.

## Debug builds optimize dependencies and strip their symbols

One profile setting carries more weight than anything in the sample code:

```toml
[profile.dev.package."*"]
opt-level = 3
debug = false
```

The selector is the wildcard, so it applies to every dependency while workspace crates keep the default. The comment written above it gives the reason with numbers. Without the setting, a `flutter run` debug build ships unoptimized inference kernels: measured on a Pixel 8, one five second Whisper-tiny window took minutes in debug versus roughly 3.5 seconds in release. The price is a slower first cold build, and dependency artifacts are cached after that.

`debug = false` is there for a different reason, and the reason is disk. A single Android cargokit build with optimized dependency artifacts and full debuginfo exceeded 45 GB. Workspace crates keep their debuginfo, so a stack trace through Xybrid's own Rust code still carries line numbers; dependencies keep their names in backtraces through symbols only, so a fault inside the inference kernels gives you a function name and nothing else.

Both sides of that trade are stated in the comment rather than left to be discovered. It is a sensible call for anyone iterating on an app, and the wrong call for anyone whose next task is a line-by-line bisect inside llama.cpp.

## Cargo and Bazel hold overlapping test targets side by side

The justfile defines every task twice. `cargo build --workspace` has a Bazel twin, `cargo test --workspace` has a Bazel twin, and each has an analyze or verbose variant. The Bazel test recipe is annotated as the tests represented in the graph that CI uses, and it enumerates seven target patterns: the CLI, core, the FFI facade, `xybrid-llama`, the SDK, the integration tests, and xtask.

Two things Cargo covers are missing from that list. `xybrid-whisper` has no Bazel target pattern, so the Whisper wrapper is exercised only through the Cargo workspace, and the Flutter binding is not in the graph at all. Calling the Bazel set the same targets as CI is a claim about the graph, not about the whole workspace.

The analyze recipe is the cheap one, running `bazelisk build --nobuild` across crates, bindings, macros, xtask, integration tests, and a single `//:llama` target, which resolves a dependency edge without paying for a compile. One variable sets the host config: on macOS it injects `--config=macos-metal`, and elsewhere the flag is empty. Metal is an explicit per host choice rather than something detected, so a Linux runner and a developer laptop build deliberately different things.

The rest of the recipe list is ordinary, and one entry is worth noting for what it rejects: clippy runs with `-D warnings`, so a lint warning fails the build rather than scrolling past.

## One version number written five different ways

Every platform in the quick start resolves to 0.10.1, and each one spells that constraint in a different dialect. Flutter takes a caret range:

```yaml
dependencies:
  xybrid_flutter: ^0.10.1
```

Kotlin takes an exact coordinate with no range at all:

```gradle
dependencies {
    implementation("ai.xybrid:xybrid-kotlin:0.10.1")
}
```

Swift does not use a package index at all. The Package.swift entry points the dependency at the repository URL itself with `from: "0.10.1"`, so SwiftPM resolves against tags in Git and every consumer compiles from source rather than from a binary artifact. npm takes an exact version of `@xybrid/react-native`. Unity, the one platform with a recommended install, takes no version in the command:

```sh
openupm add ai.xybrid.sdk
```

So a single workspace version arrives as a caret range, an exact pin, an open lower bound over git, another exact pin, and no constraint whatsoever. Auditing which build you shipped means reading five manifests in five ecosystems, and a version bump is five edits rather than one. The documented escape from the unpinned Unity default is adding `https://package.openupm.com` as a scoped registry for the `ai.xybrid` scope, which puts the constraint back under your control.

## Expo Go is the one install path that cannot work

The React Native binding has the strictest install requirements of any platform, and one of them rules out the path of least resistance. It needs a React Native 0.76 or newer app on the New Architecture, or an Expo SDK 52 or newer development build. A bare React Native project needs `cd ios && pod install`. An Expo project needs `npx expo prebuild`, because Expo Go cannot load the native module. There is no configuration of Expo Go that makes this work.

```sh
npm install @xybrid/react-native@0.10.1
```

The second restriction concerns hardware and is easy to read backwards. On the iOS Simulator, llama.cpp uses the CPU for reliable inference, while real iOS devices keep Metal acceleration. That makes the simulator the slower path by design rather than a leftover fallback, and it is stated rather than implied. The practical consequence is that a developer benchmarking on a simulator is measuring unaccelerated CPU inference and will draw a conclusion about the library that does not hold on hardware.

The third thing you will look for is also absent from the inline sample. Streaming and cancellation for React Native are not shown in the quick start at all; they live in a separate guide linked from that section, so the surface you can read in the README itself covers loading a model, one generation call, and a release call.

## The same three lines in Dart, Kotlin, and Swift

Three of the five platform samples are one piece of code with different punctuation. Dart, Kotlin, and Swift each load a model called `kokoro-82m`, each send `Hello world` as a text envelope, and each comments that the result is 24kHz WAV audio. Only the wrapper name and the async spelling move. Dart calls `Xybrid.model` with an `XybridEnvelope`, Kotlin calls `Xybrid.model` with an `Envelope` and a `runAsync`, Swift calls `Xybrid.model` with `Envelope.text` passed as a labelled argument.

The TypeScript sample is the outlier twice over. It is the only one naming a real text model, loading `lfm2.5-230m` through `ModelLoader.fromRegistry` rather than handing an identifier to a model factory, and it is the only one bounding generation, passing `GenerationConfigs.greedy({ maxTokens: 64 })` instead of leaving the defaults in place.

```ts
import { Envelope, GenerationConfigs, ModelLoader } from '@xybrid/react-native';

const model = await ModelLoader.fromRegistry('lfm2.5-230m').load();
const result = await model.run(Envelope.text('Name three rivers.'), {
  generationConfig: GenerationConfigs.greedy({ maxTokens: 64 }),
});
console.log(result.text);
await model.release();
```

The third difference carries the most practical weight. This is the only sample that frees the model, calling `await model.release()` after reading the output. The Dart, Kotlin, and Swift snippets stop at the run call, so a reader who copies the first one they find is trusting a native heap to be reclaimed without being told it happened.

## Vendored sources and five patch files replace a dependency graph

The root directory is where the real cost of this project sits. Next to the ordinary Cargo and Bazel files there is a `vendor/` directory, a `hermetic_llvm.patch`, and four more patches aimed at build rules the project does not own: `rules_android_ndk.patch`, `rules_foreign_cc.patch`, `rules_rs.patch`, and `rules_rust.patch`. There is also `.gitmodules`, so submodules are part of the arrangement.

Five patches against upstream rules is a maintenance obligation no dependency manager can discharge. Every rules bump becomes a merge, and every merge is a chance to change how a native library compiles on a platform the maintainers do not use daily. That is the bill for hermetic native builds, and hermetic native builds are exactly what let the two `-sys` crates behave like ordinary crates in a Cargo workspace.

The rest of the root explains how a project this size gets built rather than what it does. `MODULE.bazel` and `BUILD.bazel` sit beside a pinned `.bazelversion`, a `.bazelrc`, and a `.bazelignore`. `xtask`, `tools/`, and a `spike/` directory cover build automation and experiments. `install.sh`, `install.ps1`, and `clean.sh` live at the top level instead of inside `tools/`, which says something about how often they are reached for.

One gap is worth naming. The quick start links a Python binding under `bindings/python/` and a web SDK page, but neither has a workspace member, and the only binding inside the Cargo workspace is Flutter's. The example directories cover Android, Flutter, iOS, and Unity, with no Rust example among them.

## A precompiled binaries release sitting between two version tags

The release list has three entries and one of them is not a version. `v0.10.0` went out on 2026-09-27 at 07:04 UTC. A release tagged `precompiled_cdf8614faba14dfd701f4e8ea3ca82e9`, named Precompiled binaries cdf8614f, followed at 23:19 the same day. `v0.10.1` landed at 00:11 UTC the next morning. The precompiled artifact is named after a commit, so the project publishes binaries alongside source builds, and the releases page is one of the links in the quick start for anyone who wants them instead of compiling.

The last push to master is dated 2026-10-01, so nothing here describes an abandoned checkout. The repository shows 466 stars, 51 forks, and 40 open issues. The licence is Apache-2.0 and the workspace manifest repeats it at edition 2021.

The governance file list is longer than the sample code. MAINTAINERS, GOVERNANCE, CODE_OF_CONDUCT, SECURITY, and CONTRIBUTING sit next to two separate licence agreements, CLA and CLAUDE, plus AGENTS.md and an `agents/` directory with `.agents/` and `.claude/` variants. Four trust badges point at the OpenSSF Scorecard and OpenSSF Best Practices. The README exists in English, Simplified Chinese, and Japanese, which is an unusual commitment for a tool whose entire interface is English command lines and English identifiers.

One gap is genuine rather than inferred. The Unity section stops mid sentence at the manual install path that follows the OpenUPM command, so OpenUPM is the only Unity install that can be read end to end and the git based alternative the section begins to introduce cannot be checked.

## Conclusion

Xybrid fits teams that need inference to run inside an app or a game with no network path and that are willing to own a native build. The Rust core carries no networking crate in its dependency list, which is the strongest available evidence for the offline claim, and the workspace optimizes dependencies even in debug builds so a first measurement is not an artifact of your own build profile. Three things are worth checking before you commit. The same version reaches you through five registries written in five different constraint dialects, so pin them yourself rather than inheriting them. Only the React Native sample shows how to bound generation and how to release a loaded model. And the Bazel target list omits the Whisper wrapper and the Flutter binding, so the continuous integration signal is narrower than the workspace it represents. Judge Unity on the OpenUPM path alone, because the manual install path in the README stops mid sentence.

## FAQ

### What does Xybrid actually run on the device?

Xybrid wraps two inference engines in Rust: xybrid-llama over llama-cpp-sys and xybrid-whisper over whisper-cpp-sys. The workspace covers language models, speech recognition, and speech synthesis, and the shared dependency list contains no HTTP or networking crate, so nothing in the core requires a cloud endpoint.

### Which platforms does Xybrid publish bindings for, and how is each installed?

The quick start covers Flutter through pub.dev, Kotlin through Maven Central, Swift through a git URL in Package.swift, React Native through npm, and Unity through OpenUPM, plus a crates.io package for Rust. All of them resolve to version 0.10.1, though each states the constraint differently.

### Can I use Xybrid in an Expo Go project?

No. The React Native binding needs React Native 0.76 or newer on the New Architecture, or an Expo SDK 52 or newer development build, and Expo Go cannot load the native module. An Expo project must run npx expo prebuild first.

### Why does Xybrid optimize its dependencies even in debug builds?

The workspace sets opt-level 3 for all dependencies under the dev profile because unoptimized debug kernels are unusably slow: on a Pixel 8, one five second Whisper-tiny window took minutes in debug versus about 3.5 seconds in release. Debug info is disabled for dependencies because one Android cargokit build with full debuginfo exceeded 45 GB.

### Is Metal acceleration available in the Xybrid iOS Simulator build?

No. On the iOS Simulator, llama.cpp uses the CPU for reliable inference, while iOS devices keep Metal acceleration. Performance measured in the Simulator therefore reflects the unaccelerated path rather than the one real devices use.

### How does Xybrid keep native dependencies in CI, and what does it patch?

The repository vendors its third party sources and carries five patch files: hermetic_llvm.patch plus patches against the Android NDK, foreign_cc, rs, and rust build rules. A pinned .bazelversion, MODULE.bazel, and BUILD.bazel sit beside a justfile that offers both cargo and bazelisk recipes for build, test, and analysis.

## Sources

- [License: Apache-2.0](https://github.com/xybrid-ai/xybrid/blob/master/LICENSE)
- [Project website](https://xybrid.ai)
- [README](https://github.com/xybrid-ai/xybrid/blob/master/README.md)
- [Releases](https://github.com/xybrid-ai/xybrid/releases)
- [xybrid-ai/xybrid on GitHub](https://github.com/xybrid-ai/xybrid)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/xybrid-ai-xybrid
