# TurboFieldfare: Gemma 4 26B-A4B inference in about 2 GB of RAM on Apple Silicon

> A model-specific Swift and Metal runtime that streams MoE experts from SSD instead of loading the full 14.3 GB checkpoint, aimed at 8 GB Macs. It is macOS 26 and arm64 only, and it is not a general inference engine.

**drumih/turbo-fieldfare** — Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook

- Repository: https://github.com/drumih/turbo-fieldfare
- Stars: 6,761 · Forks: 431
- Language: Swift
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/drumih-turbo-fieldfare

## The memory problem TurboFieldfare attacks

A 26-billion-parameter model at 4-bit weights is roughly 14.3 GB on disk. On a 16 GB or 24 GB Mac that is already uncomfortable once the OS and the KV cache take their share, and on an 8 GB MacBook Air it simply does not fit. TurboFieldfare's answer is to stop treating the checkpoint as something that must be resident. According to the README, the runtime keeps the shared 1.35 GB core and the FP16 KV cache in memory, then streams only the experts needed for each token from SSD. The stated memory figure is about 2 GB of weights plus a 4K KV cache, which is what makes an 8 GB machine a supported target rather than a curiosity.

The audience is narrow and clearly stated: Apple Silicon Macs, with an 8 GB M2 MacBook Air as the validated target. This is not a tool for someone with a GPU server. It is for the person who already owns the laptop and wants a 26B model on it without buying RAM. The trade is explicit. You spend disk and latency instead of memory. The repository does not claim the trade is free, and the benchmark table shows what it costs: 5.1 to 6.3 tok/s decode on an 8 GB M2 MacBook Air, and 31 to 35 tok/s on a 24 GB M5 Pro. The README calls the measured result a reference point, not a performance ceiling, and notes that prompt length, generated length, page-cache state and hardware all affect throughput. Page-cache state deserves emphasis: because experts come off SSD, a warm cache and a cold one are different experiences, and the README does not promise otherwise.

## How the streaming MoE runtime is put together

Gemma 4 26B-A4B is a mixture-of-experts model with 26B total parameters and about 3.88B active per token. That gap is the whole opportunity. A dense 26B model would need every weight for every token; an A4B model needs a small routed subset. TurboFieldfare exploits the sparsity at the storage layer, keeping the always-used core resident and pulling routed experts on demand.

The quantisation scheme is documented rather than hidden: MLX affine 4-bit with group size 64, an 8-bit router, and 4-bit shared and routed experts. The router runs at higher precision than the experts it selects, which is the sensible ordering, since a misrouted token wastes the whole expert fetch. The runtime, the streaming installer, the CLI and the Mac app are all Swift and Metal. The README is direct that this is model-specific rather than a wrapper around MLX or llama.cpp, and the repository layout supports that reading: Sources/ holds the Swift targets, Scripts/ and Tests/ sit alongside them, and the documentation includes docs/SYSTEM_DESIGN.md and docs/OPTIMIZATION_JOURNEY.md. The README also points to an experiment inventory summarising 103 measured results across kernels, caching, I/O, prefill and decode. That record is the most useful artefact for judging whether a design decision was tested or assumed, and it is the first thing I would read after the quick start.

The package exposes six products: TurboFieldfare (the library and Metal kernels), TurboFieldfareMac (the app), TurboFieldfareDecodeService (a one-shot local model and Metal owner used by the app), TurboFieldfareCLI, TurboFieldfareServer (a loopback OpenAI-compatible Chat Completions server), and TurboFieldfareRepack (the streaming installer and install verifier). The README states that all of them use the same .gturbo model directory but that only one model-owning product should run at a time. That constraint is architectural, not advisory: the decode service exists precisely so a single process owns the model and the Metal device.

## Installing TurboFieldfare and generating a first completion

The README's quick start is a clone, a release build, and the app binary. On the first run, Swift Package Manager downloads and builds the Swift packages the tokenizer needs, so expect the first build to take longer than later ones.

```bash
git clone https://github.com/drumih/turbo-fieldfare.git
cd turbo-fieldfare
swift build -c release
.build/release/TurboFieldfareMac
```

The README is explicit that you should build the complete package so the app and its sibling decode service are both available; building only the app target leaves you without the process that owns the model. When the app opens, choose Download and let TurboFieldfare fetch and repack the pinned model, which is about 15 GB. Then choose Load Model, type your prompt, and press Generate. The download is the long pole, and the README lists enough free storage for the roughly 14.3 GB text model installation as a requirement, plus an internet connection for the first install.

Generation defaults are temperature 0.2, Top-K 64 and Top-P 0.95. Setting temperature to 0 gives deterministic greedy output. Those are the numbers to change first if the model is wandering or repeating. For a scripted path, the CLI product provides instruction chat and raw completion against the same .gturbo directory, and the server product exposes a loopback OpenAI-compatible Chat Completions endpoint. The README does not give CLI flags or server invocation examples in the text available here, so read docs/OPENAI_SERVER.md before wiring anything to the server. Two operational rules come straight from the README: run only one model-owning product at a time, and do not expect the app or CLI to execute tools. The loopback server accepts function-tool declarations and returns model-produced tool calls, but the client authorises and executes them.

## Requirements that rule people out

The platform requirements are the first filter, and they are strict. macOS 26 with Metal 4, Xcode 26 and Swift 6.2 or newer, and an Apple Silicon Mac. The package is arm64-only, and the README states plainly that older macOS and Metal versions are not supported. If you are on macOS 15 or on an Intel Mac, there is no degraded mode to fall back to. This is a project that moved with the platform rather than one that carries compatibility shims, and that choice is visible in the badges and the requirements list.

Images are the second filter, and the split is finer than the headline suggests. Vision support installs as a companion pack beside the text model, about 1.1 GB, and once installed the app, CLI and server all accept images. Without it they report that image support is unavailable, and the text runtime is unaffected. The constraint: the image tower requires an M2 or newer Apple Silicon Mac, while text-only inference remains available on M1. So an M1 owner gets the 26B text model but not the vision path. The README does not document a workaround for that, and docs/SYSTEM_DESIGN.md is where it says the cost of the tower on an 8 GB machine is discussed.

The third limitation is the one that decides whether this is the right tool at all. Throughput on the validated 8 GB M2 target is single-digit tokens per second. That is fine for a prompt you walk away from and poor for interactive chat, agent loops, or anything that regenerates long outputs. The README itself warns that the model can repeat itself or give incorrect answers and that important results should be checked. Combine that with a 4K KV cache budget and the practical context window is modest. If your work needs long documents in context or fast turnaround, the memory saving is not worth the latency.

## How TurboFieldfare differs from MLX and llama.cpp

The obvious alternatives are MLX and llama.cpp, and the difference is not speed or polish. It is where the model lives. Both of those projects are general runtimes: they load a model's weights into unified memory and execute it, with quantisation formats and backends chosen per model family. TurboFieldfare inverts that. It is written for one model, Gemma 4 26B-A4B, and its central mechanism is streaming experts from SSD so the full checkpoint never has to be resident. The README states this directly: TurboFieldfare is model-specific rather than a wrapper around MLX or llama.cpp.

That trade cuts both ways. A general runtime can run a dozen model families and will gain new ones as they ship. TurboFieldfare runs the model it was built for, and the engineering effort visible in the repository (the repack installer, the decode service, the experiment inventory) is spent on making that one model fit in 2 GB rather than on breadth. If Gemma 4 26B-A4B is the model you want and RAM is the binding constraint, the general runtimes do not offer the same answer, because their answer is to load the weights. If you want to try a different model next month, TurboFieldfare is the wrong shape of tool. One caveat worth stating: the README does not document what happens if the .gturbo directory is incomplete or corrupted beyond the existence of TurboFieldfareRepack as an install verifier, so treat the installer's verification step as the thing to run rather than something to skip.

## Maintenance, licence and the cost of staying current

The repository is not archived, and the last push was on 2026-09-08. Release 0.8.0 landed the same day, with 0.7.2 on 2026-09-04 and 0.7.1 on 2026-08-28. That cadence over a few weeks is consistent with an actively developed project, and the version numbers are pre-1.0, which is the honest signal here: interfaces and the .gturbo layout can still move between minor releases.

The upgrade cost is dominated by the model, not the code. The installer fetches and repacks a pinned model of about 15 GB, so a change to the repacking format or the pinned revision means another large download rather than a quick rebuild. The README does not document an in-place upgrade path or a rollback for the model directory, and it does not describe how a previously installed .gturbo directory is migrated across versions. That is a real gap for anyone planning to keep this on a laptop with limited free space. Budget the disk before you commit, and keep the installer around for re-verification.

The licence is Apache-2.0, which is permissive and includes an explicit patent grant. The repository also carries a THIRD_PARTY_NOTICES.md, which is where the obligations for bundled dependencies live, and the model weights come from Google under their own terms rather than the repository licence. Those are two separate licences covering two separate things, and the README does not restate the model terms. Read both before shipping anything built on this; that is a reading task, not legal advice.

## Conclusion

Adopt TurboFieldfare if you have an Apple Silicon Mac, macOS 26 with Metal 4, roughly 14.3 GB of free storage and a workload that tolerates single-digit tokens per second on an 8 GB M2. Do not adopt it if you need CUDA, x86, a general model runner, or the M1 image tower. Before installing, confirm the macOS 26 and Xcode 26 requirements, check free disk space, and read docs/BENCHMARKS.md for the M2 and M5 numbers so the throughput you get is the throughput you expected.

## FAQ

### What is TurboFieldfare?

It is a Swift and Metal runtime for running Gemma 4 26B-A4B inference on Apple Silicon Macs in about 2 GB of RAM. It keeps the shared 1.35 GB core and the FP16 KV cache in memory and streams only the experts needed for each token from SSD. The README describes it as model-specific rather than a wrapper around MLX or llama.cpp.

### What does the name "Fieldfare" mean?

The README does not explain the name in the text available here. The logo is described as a fieldfare inside a segmented cache ring, which suggests the bird and the cache-ring motif are the reference, but the repository does not state an origin for the name.

### Can TurboFieldfare run on an 8 GB MacBook Air?

Yes. The README names an 8 GB M2 MacBook Air as the validated target, with about 2 GB of weights and a 4K KV cache, and reports 5.1 to 6.3 tok/s decode on that machine. The requirements are macOS 26 with Metal 4, Xcode 26 and Swift 6.2 or newer, and the package is arm64-only.

### Does TurboFieldfare support images?

Images are supported through a vision tower that installs as a companion pack of about 1.1 GB beside the text model. Once installed, the app, CLI and server all accept images; without it they report that image support is unavailable and the text runtime is untouched. The image tower requires an M2 or newer Apple Silicon Mac, while text-only inference remains available on M1.

### Can the TurboFieldfare OpenAI-compatible server execute tool calls?

No. The loopback server accepts function-tool declarations and returns model-produced tool calls, but the client authorises and executes them. The Mac app and CLI support user and model messages plus optional system guidance and do not expose or execute tools.

## Sources

- [drumih/turbo-fieldfare on GitHub](https://github.com/drumih/turbo-fieldfare)
- [Issues](https://github.com/drumih/turbo-fieldfare/issues)
- [License: Apache-2.0](https://github.com/drumih/turbo-fieldfare/blob/main/LICENSE)
- [README](https://github.com/drumih/turbo-fieldfare/blob/main/README.md)
- [Releases](https://github.com/drumih/turbo-fieldfare/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/drumih-turbo-fieldfare
