# uzu: a Rust inference engine that ships models inside your app

> uzu from trymirai runs AI models locally through one API in Rust, Python, Swift and TypeScript. The appeal is on-device inference with no server round trip, but the project is Apple-first and the workspace is mid-rewrite.

**trymirai/uzu** — A high-performance inference engine for AI models

- Repository: https://github.com/trymirai/uzu
- Website: https://trymirai.com
- Stars: 1,820 · Forks: 92
- Language: Rust
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/trymirai-uzu

## The problem uzu solves: a model in the app, not behind an endpoint

Most teams that want a language model in a product end up calling a hosted API. That choice buys capability and costs three things: a network round trip on every token, a copy of the user's input leaving the device, and a per-request bill. uzu takes the other path. Its README describes it as "a high-performance inference engine for AI models" and lists three claims that map directly onto those costs: zero latency, full data privacy, and no inference costs. Read those as design goals rather than measurements. The README does not publish latency figures or throughput numbers, and nothing in the repository layout suggests a benchmark suite that would settle the question.

The intended audience is narrow and identifiable. The README's badge row advertises Python, TypeScript, Swift and SPM, and the Swift badge names iOS and macOS as platforms. One bullet says the engine "utilizes unified memory on Apple devices." The topics list includes metal. That is an Apple-shaped project, and the quick-start samples are written for someone embedding a model in a shipped application rather than someone standing up a service. If your deployment target is a Linux box behind a load balancer, the README does not describe your case.

## How the engine is put together: an engine, a model handle, a session

The public API in all four languages follows the same shape, which is the clearest signal about the internal design. You construct an Engine from an EngineConfig. You ask it for a model by identifier. You download that model through an async iterator that yields progress updates. You open a chat session from the engine and the model. You send a list of messages and get replies back. Nothing in that sequence mentions tensors, tokenizers or sampling parameters, so the engine owns those decisions.

Model identifiers are namespaced strings, not file paths. The README example uses alibaba:qwen3.5:0.8b:mirai:mirai-m:4, and the README links to a model page at trymirai.com/local-models for the supported set. The engine.model call returns an optional in Rust and Python, so resolution can fail and the caller handles it. The download step is separate from loading, which means a first run and a warm run behave differently: the first run pulls weights, later runs do not.

The repository layout tells you the project is in transition. Cargo.toml lists two current members, crates/uzu-engine and crates/uzu-engine-macros, then a block of nineteen crates under crates/legacy/ preceded by the comment "to be rewritten." The Python, TypeScript and Swift bindings all live under crates/legacy/uzu/bindings. So the engine core has been reworked while the language surfaces have not been migrated yet. That is worth knowing before you plan an upgrade path, because it means the binding layer you call is the older half of the tree.

## Installing uzu and running the first local chat

The README gives a separate dependency line and a separate code sample per language. Start with Python, because the install is one command and the sample is self-contained. The package is published on PyPI as uzu, and the README pins the version:

```bash
uv add uzu==0.5.26
```

After that resolves, the sample builds an engine, resolves a model, downloads it while printing progress, opens a chat session and sends two messages. The identifier below is the one the README uses; substituting another requires checking the model page first.

```python
import asyncio

from uzu import ChatConfig, ChatMessage, ChatReplyConfig, Engine, EngineConfig


async def main() -> None:
    engine = await Engine.create(EngineConfig.create())
    model = await engine.model("alibaba:qwen3.5:0.8b:mirai:mirai-m:4")
    if model is None:
        return

    async for update in (await engine.download(model)).iterator():
        print(f"\rDownload progress: {update.progress:.2%}", end="", flush=True)
```

What you should see is a progress line that advances as weights arrive, then a prompt returned to you. The model handle is None when the identifier does not resolve, and the sample returns silently in that case, so a typo looks like a no-op rather than an error.

For a Rust project, the README adds the crate from git rather than crates.io:

```toml
[dependencies]
uzu = { git = "https://github.com/trymirai/uzu", branch = "main", package = "uzu" }
```

Note the branch. Tracking main means your build follows the repository, including the legacy-to-current migration described above. If you want the version the README pins elsewhere, pin a tag instead. The TypeScript path is a normal npm install of the scoped package:

```bash
pnpm add @trymirai/uzu@0.5.26
```

And Swift uses Swift Package Manager with a version requirement:

```swift
dependencies: [
    .package(url: "https://github.com/trymirai/uzu.git", from: "0.5.26")
]
```

The README does not document a system-level prerequisite list beyond the platform badges, and it does not show a CLI. Everything runs through the library.

## Where uzu is the wrong tool

The README is silent on several things a production deployment usually needs. There is no documented server mode, no HTTP interface, and no container image. If your architecture assumes a model behind an endpoint that many clients share, uzu inverts that: the model runs in the client process, which means every install carries the weights and every device does its own compute. That is the point of the design, but it also means you cannot scale inference independently of your user count.

Platform coverage is the second boundary. The unified-memory bullet and the metal topic point at Apple hardware, and the Swift badge lists iOS and macOS. The README does not describe a CUDA or ROCm backend, and it does not describe what happens on Windows or Linux. A team standardizing on NVIDIA servers should treat uzu as unproven rather than portable.

The third issue is versioning discipline. The Rust dependency in the README tracks branch = "main". For a library whose workspace still contains a legacy tree marked "to be rewritten," following main means your build can change without a version bump on your side. The README does not document rollback, deprecation policy, or how the legacy crates will be retired. The release history shows several releases within a few days of each other in early September 2026, which is consistent with rapid iteration and inconsistent with a frozen API surface.

## Alternatives: llama.cpp, MLX and ONNX Runtime

The obvious comparison is llama.cpp, which also runs models locally and also targets consumer hardware. The difference is in the interface. llama.cpp exposes a C API plus a CLI and a server binary, so you can run it as a process and talk to it over HTTP, and you can choose your quantization at load time. uzu instead presents a high-level engine-and-session API in four languages and treats model configuration as a unified, namespaced identifier. If you want to tune quantization or run a shared endpoint, llama.cpp gives you the knobs. If you want the same three lines of chat code in a Swift app and a Node script, uzu's API shape is the reason to pick it.

On Apple silicon specifically, MLX is the closer neighbor, since it is also built around unified memory and Metal. MLX is a general array framework with a model zoo on top; uzu is an inference engine with a chat session abstraction. The trade-off is control versus convention. MLX lets you write the forward pass; uzu expects you to call engine.chat and accept its defaults, and the README does not expose sampling or decoding parameters in the quick start.

ONNX Runtime occupies a different position again: it is a cross-platform graph executor with a wide operator set and a long list of language bindings, and it is not chat-shaped. It will run a transformer, but you supply the tokenizer loop and the generation logic. uzu's value is that generation is already there. Its cost is that you inherit its model support list rather than bringing your own graph.

## Licence, maintenance and what an upgrade costs

uzu is MIT licensed, both in the repository metadata and in the workspace manifest, which sets license = "MIT" under [workspace.package]. MIT is permissive: it allows commercial use, modification and redistribution provided the copyright notice and permission notice are preserved. That is a description of the licence text, not legal advice, and it says nothing about the model weights you download, which carry their own terms from their publishers. The README does not discuss weight licensing, and the model page is where that would live.

The repository is not archived, and the last push was on 2026-09-09, so development is current. Releases 0.5.23, 0.5.25 and 0.5.26 landed between 2026-09-03 and 2026-09-06. That cadence is the real upgrade cost: patch releases arriving every day or two means you should pin, and the README's own Rust snippet does the opposite by tracking main. The workspace version is 0.5.26 and the README pins 0.5.26 for Python, TypeScript and Swift, so those three ecosystems have a version to hold. Rust does not, unless you replace the branch with a tag.

The other upgrade cost is structural. The bindings you call live under crates/legacy/, and the comment in Cargo.toml says that tree is to be rewritten. When that rewrite lands, the crate paths and possibly the binding internals change. The README does not say whether the public API in Python, TypeScript and Swift will survive intact. Any integration should be thin enough that a binding-layer change is a dependency bump rather than a rewrite.

## Conclusion

Adopt uzu when you ship an Apple-platform app and want a local model behind one API instead of a hosted endpoint: the Swift package and the unified-memory note in the README point at exactly that case. Skip it if you need a documented server deployment or a stable crate layout, because the workspace still carries a legacy tree marked to be rewritten and the README does not document rollback or a server mode. Before committing, build the quick-start example, run the Python or TypeScript snippet against the alibaba:qwen3.5:0.8b:mirai:mirai-m:4 identifier to confirm model resolution, and check the crates/uzu-engine and crates/legacy split in Cargo.toml to see which crates your build actually depends on.

## FAQ

### What are the two main types of inference engines?

The README does not classify inference engines into types. It describes uzu as a high-performance inference engine for AI models that runs models inside the calling application, with bindings for Rust, Python, Swift and TypeScript.

### What is inference computing and how does it work?

The README does not explain inference computing in general. It shows how uzu performs it: you create an Engine from an EngineConfig, resolve a model by identifier, download it, open a chat session, and send messages to get replies.

### Is an LLM an inference model?

The README does not address this question directly. It lists llm among the repository topics and its quick-start example loads alibaba:qwen3.5:0.8b:mirai:mirai-m:4, a chat model, through the same engine and session API used for any supported model.

### What is an inference server?

The README does not define an inference server, and it does not document a server mode, an HTTP interface or a container image for uzu. Every sample constructs the engine in-process and opens a chat session directly.

## Sources

- [License: MIT](https://github.com/trymirai/uzu/blob/main/LICENSE)
- [Project website](https://trymirai.com)
- [README](https://github.com/trymirai/uzu/blob/main/README.md)
- [Releases](https://github.com/trymirai/uzu/releases)
- [trymirai/uzu on GitHub](https://github.com/trymirai/uzu)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/trymirai-uzu
