# A local coding agent that uses XML instead of JSON because the model is small

> The readme is candid about what this is, calling it a proof of concept, and the one design decision worth arguing about is stated with a reason: small models handle XML more reliably than JSON function calling. Everything else, the four model sizes, the forty-round agent loop and the offline promise, follows from running a 3 GB model on a laptop.

**ammaarreshi/gemma-chat** — Local AI chat + coding agent for Apple Silicon, powered by Gemma 4 via MLX / Supports Ollama

- Repository: https://github.com/ammaarreshi/gemma-chat
- Stars: 1,425 · Forks: 257
- Language: TypeScript
- License: MIT
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/ammaarreshi-gemma-chat

## The premise is privacy, and the readme states it three times

Getting it running is a clone and a development command, on macOS with Apple Silicon, a supported Python range and Node 20 or later:

```bash
git clone https://github.com/ammaarreshi/gemma-chat-public.git
cd gemma-chat-public
```

```bash
npm install
npm run dev
```

The header says vibe code without the internet, and expands it into three clauses: no API keys, no cloud, no wireless required. The idea section asks what you could build from an airplane or a cabin with no signal, and then says the plainer version, without sending your code to someone else's server. That framing matters because it tells you what kind of project this is competing with. It is not competing on capability with a hosted assistant; it is competing on where the inference happens. Everything downstream follows from that. The model is small, because a small model fits in the memory of a laptop. The tool protocol is XML, because a small model is more reliable with it. The agent loop is capped, because a small model will keep going in circles if you let it. And the app is an Electron shell with a live preview, because you want to see the thing build while it builds, which is the actual appeal of this genre. The readme's own framing, a proof of concept for fully offline local-first coding with a small open model, is the most accurate sentence in it, and it is worth holding on to when the marketing language in the rest of the document gets enthusiastic.

## Four model sizes, and the hardware table is really a capability table

The model table is the most decision-relevant part of the readme, and it is laid out as size, recommended or not, and best for. Four variants are offered. The smallest, at about one and a half gigabytes, is for fast question answering and simple tasks. The recommended one, at about three gigabytes, is described as the balance between speed and capability. A mixture-of-experts variant at about eight gigabytes is for stronger reasoning and needs sixteen gigabytes of memory or more. The largest, at about eighteen gigabytes, is for maximum quality and needs thirty-two or more. Read as a table, that is a quality ladder whose rungs are memory. That is the honest framing: on this project you are not choosing a model so much as choosing how much of your machine's memory to give up and how much latency to accept. The recommended default at three gigabytes is a small model, and a small model is what produces the two design decisions the readme explains at length. So the model table and the architecture section are the same document. If your machine has thirty-two gigabytes you can run the largest and the same code will feel considerably better, and if you have eight you are running the smallest and should expect it to struggle with a multi-file project.

## XML as a tool protocol, and why that is a real engineering choice

The under-the-hood section explains the agent loop and then justifies its wire format. The loop works like this: in build mode each assistant turn streams tokens from the local server, XML action blocks are parsed out of the stream as they arrive, each is executed, and the results are fed back for the next turn, up to forty rounds per user message. And the reason for XML rather than JSON function calling is stated in one sentence: small models handle XML more reliably than JSON function calling. That is a correct and specific claim, and it is the kind of thing you only learn by trying. A model with a few billion parameters has learned to emit well-formed JSON from a schema far more often than it emits a well-formed XML element with correct nesting and closing tags, and a malformed JSON object breaks a parser while a malformed XML block can be caught and skipped. The example shows the shape: an action element with a name attribute, then child elements for the path and the content, with the file content inline. Two consequences follow. First, the parser is a custom one and the readme names it, living in a tools file alongside the tool definitions and the system prompts. Second, an agent that can write files and run shell commands, driven by a small model, is a meaningful capability to hand to something running on your laptop, and the round cap plus the sandboxed per-conversation workspace are the two boundaries the readme puts around it.

## Live streaming, a 450 millisecond flush, and a preview that reloads

The streaming behaviour is described with a specific number, which is the kind of detail that tells you the feature was actually built rather than specified. As the model generates file content, partial writes are flushed to disk on an interval of roughly four hundred and fifty milliseconds, and the preview frame reloads in real time so you watch the page assemble itself. Two design choices are visible in that sentence. Flushing on a timer rather than per token means the disk sees a write every interval even if the model is producing a large file, which is what makes the preview feel live. Reloading the preview on a flush rather than on a completed action means you see half-written files, which is the aesthetic of the genre and also a source of flash. The architecture listing confirms the mechanism: a canvas component with preview, code and file tabs for build mode, and a workspace file in the main process described as a per-conversation workspace plus a static file server. So the preview is an embedded frame loading from a local static server rooted at the conversation's sandbox directory, which is why the workspace isolation matters. Each conversation gets its own directory and its own server, so two chats cannot see each other's files, and that is the boundary between a local assistant and a local assistant that will occasionally write to the wrong place.

## Voice in the browser, auto-provisioning a Python runtime, and a signed build

Three operational details round out the picture. Voice input is described as local speech-to-text via an in-browser implementation of a speech model running through WebAssembly, and the architecture confirms it with a dedicated library file in the renderer and a WebAssembly note in the stack table, so the microphone path never leaves the machine either, which is the point of the project. Setup is described as zero configuration, and the mechanism is that on first launch the app detects a Python installation, creates a virtual environment, installs the runtime, and downloads the model. That is a friendly story and a heavy first run: a multi-gigabyte download plus a compiled dependency installed from source, on a machine where the compiler for that dependency is the thing that has to be right. The build story is the third item. A distributable command produces a signed disk image in a directory, and the readme says you can share it directly and recipients drag it to applications. Signed by whom is the question, since an unsigned desktop app triggers a security warning on macOS, so this claim is worth verifying before you distribute anything. The repository also carries separate type-check configurations for the node side and the web side, which is the right structure for a desktop app and more than most projects at this stage bother with.

## Conclusion

Use gemma-chat if the point is that your code must not leave your machine, since the readme's premise is no API keys, no cloud and no network after a one-time model download, and on Apple Silicon with enough memory the larger variants are usable. Do not expect it to match a hosted coding assistant, because the recommended model is a few gigabytes of parameters and the readme calls the whole thing a proof of concept. Four things to verify. How much memory your machine has, since the readme gives a variant at about eight gigabytes needing sixteen or more and a largest one at about eighteen gigabytes needing thirty-two or more, so the quality ceiling is a hardware ceiling. That you are on Apple Silicon with a specific Python and Node range, since the runtime is a framework for that hardware and a version mismatch in the local runtime is the most likely first failure. How you feel about the tool protocol, because the agent executes shell commands and writes files from a model's XML output, which is a real capability running on a model small enough to make mistakes. And whether the signing story is solid, since the readme says a distributable build produces a signed disk image, and a self-signed artefact is a different thing. The licence is MIT, there are no published releases, and the last push was on 2026-04-28.

## FAQ

### What does gemma-chat do?

It is an Electron app that runs a small open model locally on Apple Silicon as both a chat assistant and a coding agent. In build mode it writes multi-file projects into a sandboxed workspace with a live preview, and in chat mode it can use tools including web search, URL fetching, a calculator and shell commands. The readme calls it a proof of concept for fully offline local coding.

### Why does gemma-chat use XML for tool calls instead of JSON?

The readme states that small models handle XML more reliably than JSON function calling. Actions are emitted as XML elements with a name attribute and child elements for arguments, and a custom parser reads them out of the token stream as they arrive.

### Which models does gemma-chat offer and what do they need?

Four sizes. About one and a half gigabytes for fast question answering, about three gigabytes recommended as the balance of speed and capability, about eight gigabytes for stronger reasoning needing sixteen or more gigabytes of memory, and about eighteen gigabytes for maximum quality needing thirty-two or more. Switching between variants is described as happening on the fly.

### What are the requirements to run gemma-chat?

macOS on Apple Silicon, a specific range of Python versions, and Node 20 or later. First launch auto-detects Python, creates a virtual environment, installs the runtime and downloads the model, and the readme notes that the model download is around three gigabytes.

### How does gemma-chat keep separate conversations isolated?

Each conversation gets its own sandboxed filesystem workspace served by a local static HTTP server, which is also what backs the live preview. The architecture describes a workspace file in the main process handling the per-conversation directory and the static file server, and the build mode writes multi-file projects into that sandbox.

### What licence is gemma-chat released under?

MIT. The repository publishes no GitHub releases, the manifest reads version 0.1.0, and the last push to the main branch was on 2026-04-28. Credits go to the model, the runtime framework and the browser speech library.

## Sources

- [ammaarreshi/gemma-chat on GitHub](https://github.com/ammaarreshi/gemma-chat)
- [Issues](https://github.com/ammaarreshi/gemma-chat/issues)
- [License: MIT](https://github.com/ammaarreshi/gemma-chat/blob/main/LICENSE)
- [README](https://github.com/ammaarreshi/gemma-chat/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ammaarreshi-gemma-chat
