# zeraix puts a local model, a file editor and an agent in one Electron window

> Zeraix is a desktop workspace that runs GGUF models through llama.cpp-based runtimes, gives agents file, terminal and browser tools behind four approval modes, and points at a separate inference engine called Imparo that it does not yet ship. The install is a .dmg or .exe rather than a package, and the stable channel table disagrees with the sentence printed under it.

**zeraix/zeraix** — Open-source local AI workspace — advancing on-device inference.

- Repository: https://github.com/zeraix/zeraix
- Website: https://zeraix.com
- Stars: 445 · Forks: 13
- Language: TypeScript
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/zeraix-zeraix

## The stable channel says v2.3.1, the line under it says 2.0

The quick start table lists one row: Stable, v2.3.1, with a macOS `.dmg` and a Windows `.exe`. The sentence immediately underneath says something different, that 2.0 is the current stable release including the features introduced during the 2.0 beta cycle. Both cannot describe the same channel. The published releases add a third data point: v2.3.2 went out on 2026-09-30, one day after v2.3.1 and the same day as the last push, and the root `package.json` already reads 2.3.2. Practical advice follows from that: treat the releases page as the source of truth for a download, and expect the `main` branch to be ahead of any installer, which is stated outright in the quick start. Earlier optimization results are collected separately in MODEL_SYSTEMS.md, with a CHANGELOG.md and the releases page carrying the rest of the history.

## The install is a .dmg or .exe, not a package manager entry

Zeraix is a private package (`"private": true` with `main` pointing at `electron/main.mjs`), so there is no `npm install` path to the application itself. On macOS the requirement is 13 or newer on Apple Silicon: open the `.dmg` and drag Zeraix into Applications. Windows is x64 only, 10 or 11, through the `.exe` installer. Memory guidance is blunt, 16 GB or more recommended, with some smaller models working on 8 GB and longer contexts needing more. If your operating system shows a security warning, the README tells you to confirm the installer came from this repository before continuing. First run then goes through Model Library, which detects hardware, lets you review a recommended model's memory and disk requirements, and downloads both the model and its runtime before you can select the running model in chat and send it a message. Sandbox resources arrive later, when something needs them.

## Four approval modes decide what an agent may do

Default, Full trust, Manual approval and Plan mode are the four settings on the chat composer, and they govern how tool calls are approved. Plan mode investigates and proposes; Default reviews protected actions as a task progresses. Sub-agents can be delegated to, inspected while they run, and stopped individually. File reading, searching, editing, diff inspection, terminal commands and browser tools all sit behind the same approval layer, which is the Rust Agent Runtime's execution and permission layer. An optional QEMU sandbox provides an isolated environment, but only for supported commands, so the honest reading is that you still have to check whether a given command lands on the host or in the sandbox. Built-in file operations share that same Rust execution and permission layer rather than having a separate path, which keeps one set of rules to reason about.

## llama.cpp runs inference today, Imparo is the research track

The README's hero section points at Imparo, described as an open-source LLM inference engine that adapts to hardware and workload, with its own source, updates and performance reports. Further down, the desktop is stated to use llama.cpp-based runtimes for general local inference, and Imparo is said not to be the bundled default engine yet. Three components are therefore in play: the desktop workspace, the Rust Agent Runtime that executes agents and tools, and the separate Imparo engine project. The tuned paths are narrower still, covering model-specific memory planning, mapped weights, MoE pooling, speculative decoding and persistent KV reuse on Apple Silicon. Those results are specific to tested Apple Silicon configurations, and the documentation says so directly: Windows users should use what their runtime documents, because Apple Silicon numbers do not establish Windows performance.

## Six curated models, everything else is an import

The catalogue names six models with their supported profiles: Qwen3.6-35B-A3B and Gemma 4 26B-A4B as mixture-of-experts entries with vision and MTP, Qwen Bonsai 27B as a compact dense model with optional DSpark speculative decoding, Gemma 4 12B and E4B for vision and audio, and LFM2.5-2.6B as a lightweight text model with tool calling. Availability depends on the runtime and platform, and usable context length depends on your memory. Past that curated set, community GGUF models come with repository search, quantization selection, memory estimates and context or KV settings, and you can point the app at a custom OpenAI-compatible endpoint. Architecture and tool-calling support follow the runtime, which is the part worth checking before you commit to a quantisation. Custom models and generation engines can also be edited in place, so a curated entry is a starting point rather than a ceiling.

## A TAP reporter made a red CI run unreadable, twice

The scripts section carries an unusually honest comment. The default test command is `node --test` over `test/**/*.test.mjs`, and the CI variant switches to the spec reporter because, in the words of the comment, the default TAP reporter prints a failure inline thousands of lines up, so a CI log tail shows only the counts, and twice a red run had been unreadable without hunting for the one line that matters. Development setup has a matching note: `electron:dev` runs two preparation steps inline rather than in a pre-hook, because pnpm does not run pre and post scripts by default and the sidecar is not optional, since without it the app has no file tools:

```bash
node scripts/ensure-runtime.mjs
node scripts/ensure-skin-engine.mjs
```

That script then starts `next dev` and waits on port 3000 before launching Electron, and `rebuild:native` exists for rebuilding `node-pty` after an Electron upgrade.

## Offline means local, and only after the downloads

The offline claim has a boundary and the documentation states it. Local inference and local tools work offline once setup is done, while web search, remote MCP servers, generation services and cloud models each need their own network connections. Setup itself is not offline: the first run downloads model and runtime assets, sandbox resources are downloaded when needed, and memory and context limits depend on the selected model, configuration and hardware. Bundled document Skills cover Word documents, PDFs, presentations and spreadsheets with local processing, and the media library previews images, video, audio, PDFs and text. Extensions come through local or remote MCP servers, a plugin catalogue, and custom models and generation services you can edit in place. Ongoing work is a feature of its own: conversations and memory stay local, task goals are tracked, recurring workflows keep run history and notifications, and after an interrupted session a recovery notice tells you what was running so you can decide how to continue.

## Where a desktop workspace is the wrong tool

Zeraix bundles a lot to get one window: Electron, a Next.js renderer, a Rust sidecar, node-pty, and optionally QEMU for sandboxing. If all you need is to serve a model to your own code, the llama.cpp runtimes it wraps are the smaller dependency, and if you want to go further in that direction the Imparo project is where the inference work is heading. A cloud assistant is the other comparison, and the difference is the one Zeraix is built around: local models need no account, no subscription and no API key, while search, remote tools and generation services need credentials you were going to supply anyway. The cost of that privacy is hardware. Sixteen gigabytes, Apple Silicon or Windows x64, and an installer you download rather than a package you pin.

## Conclusion

zeraix fits someone who wants one window for local inference, file work and agent tasks on a 16 GB Apple Silicon or Windows machine, and is willing to accept that Windows gets none of the tuned inference paths. Skip it if you only need a model server, since the desktop brings Electron, Next.js, a Rust sidecar and a QEMU sandbox with it. Before installing, read the release page rather than the stable channel table, which points at v2.3.1 while the newest published tag is v2.3.2.

## FAQ

### What are Zeraix's system requirements?

macOS 13 or newer on Apple Silicon, or Windows 10/11 x64. Memory guidance is 16 GB or more, though some smaller models support 8 GB systems, and longer contexts need more. Tuned inference paths such as mapped weights and speculative decoding are documented for Apple Silicon only.

### Which version of Zeraix should I download?

Check the releases page rather than the stable channel table. That table points at v2.3.1 for macOS and Windows, the sentence below it calls 2.0 the current stable release, and the newest published tag is v2.3.2, with the main branch possibly ahead of any installer.

### Does Zeraix need an account or API key to run local models?

No. The local core requires no Zeraix account, subscription or API key. Cloud models, web search, remote MCP servers and generation services do need their own network connections, and a custom OpenAI-compatible endpoint needs its own credentials.

### Is Imparo the inference engine Zeraix ships with?

Not yet. The desktop currently uses llama.cpp-based runtimes for general local inference, and the README states that Imparo is not yet the bundled default engine in Zeraix. The Rust Agent Runtime is a separate component that handles agent execution and tools.

## Sources

- [License: Apache-2.0](https://github.com/zeraix/zeraix/blob/main/LICENSE)
- [Project website](https://zeraix.com)
- [README](https://github.com/zeraix/zeraix/blob/main/README.md)
- [Releases](https://github.com/zeraix/zeraix/releases)
- [zeraix/zeraix on GitHub](https://github.com/zeraix/zeraix)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zeraix-zeraix
