mac-code's 30 tok/s headline is a 2-bit model, not the 4-bit one it benchmarks
mac code — Claude Code, but it runs on your Mac for free. 35B AI agent at 30 tok/s via Apple Silicon flash-paging. $0/month.
At a glance
- What is it?
- A Mac-only workspace for running local models, built on two ideas: a fast path that fits a heavily quantized model entirely in RAM, and a research path that streams feed-forward weights from SSD so a full 4-bit model that does not fit can still run. The two paths differ by a factor of twenty in speed and by a full bit of quantization, and the README's own numbers make that trade visible.
- Who is it for?
- There are two different projects here and they should be judged separately. The fast path is a standard llama.cpp setup with a two-bit quantized mixture-of-experts model that fits in memory, which is useful today and delivers the headline number honestly for what it is.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 177 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The 30 tok/s number belongs to a two-bit model
Two configurations of the same 35B model appear at opposite ends of this project's own documentation, and they are not equally good.
The fast one is a 2-bit quantization of a mixture-of-experts model, ten point six gigabytes, that fits entirely in RAM on a 16 GB Mac mini M4 and runs at 30 tokens per second. That is the recommended quick start, and it is the configuration the project description points at when it advertises a 35B agent at 30 tokens per second.
The research one is described as full four-bit quality, and it runs between one and five tokens per second depending on the mechanism. Every quality entry in the streaming results table reads full 4-bit.
So the headline speed comes from a model compressed to two bits per weight, and the quality claim comes from a model at four bits. Those are different artefacts, and the pages on which they appear do not say so.
The mechanism behind that trade is the project's actual contribution and it is a good one. Feed-forward weights, which are the bulk of any transformer, are streamed from SSD one layer at a time, used for a single matrix multiplication, then discarded, while attention weights, embeddings, norms and the KV cache stay pinned in memory. Memory therefore never grows no matter how large the model is.
One model, three mechanisms, three speeds and two memory figures
The compatibility table at the top lists Qwen3.5-35B-A3B three separate times on the same machine, and the numbers do not reconcile with the tables further down.
Thirty tokens per second with the two-bit quantization fitting in RAM. Then 5.4 tokens per second for a four-bit version through something called Expert Sniper, holding 8.7 GB. Then 1.54 tokens per second for the quantized-for-memory version through Flash Streaming.
Further down, the streaming results table gives that same model as 19.5 GB total, 8.7 GB of RAM and 5.4 tokens per second. Then the section heading for running the agent gives 1.54 tokens per second and 1.42 GB of RAM.
So the one row that claims 1.42 GB of memory and the table that claims 8.7 GB for the same model and the same machine are two different claims, and neither is explained. The 27B numbers do reconcile, at 0.18 tokens per second and 5.5 GB in both places.
None of this makes the tool wrong. It means a reader comparing rows has to work out which mechanism and which memory figure belongs together, and the documentation does not do that work for them.
A performance table whose quality claim rests on five prompts
The streaming results section opens with an emphatic claim: every number below was measured on a 16 GB Mac mini M4, and nothing is estimated. The methods sentence underneath is where the caveat lives.
Across five varied prompts from cold start. And quality verified, citing a specific numeric answer that came out right and some correct Python, with the word and others.
So the speed column is measured and the quality column is five prompts. For a project whose entire argument is that streaming weights from storage produces output good enough to use, that is a thin basis, and it is the number a sceptical reader should push on first.
The tuning parameter is at least named with its failure mode. The mechanism combines a cache-aware routing bias, a co-activation prefetch and a right-sized least-recently-used cache. The bias value of 1.0 is described as the universal safe maximum, and the claim is that quality degrades at 1.5 on both models. Having a stated cliff rather than a vague sensitivity is good practice.
The generalisation is wider than the measurements. Everything was measured on one machine, while the compatibility table extends the claims to any Mac with 8 GB or 16 GB, and the pull quote asserts the method works on any 16 GB Apple Silicon Mac.
What actually streams, and why the mixture-of-experts models are faster
The mechanism is described as splitting the model by access pattern, which is the right mental model.
Pinned in RAM, four to six gigabytes: attention weights, embeddings, norms and the KV cache. Loaded once and kept. Streamed from SSD on every token: the feed-forward weights, loaded layer by layer, used for one matrix multiplication, then thrown away, so memory stays flat.
The per-token loop is written out as four steps. Run attention from RAM, which is instant. Load that layer's feed-forward weights from SSD, between 165 and 221 megabytes depending on the layer. Run the multiplication on the GPU. Discard the weights.
For mixture-of-experts models the second step is where the speed comes from. Instead of loading the whole expert tensor, it loads only the active experts for the token, eight of them, around 14 megabytes, rather than all 256. That is the stated reason mixture-of-experts models are roughly ten times faster than dense ones under this scheme, and it lines up with the table, where the two mixture-of-experts rows are at 4.3 and 5.4 tokens per second against 0.15 and 0.18 for the dense models.
The trade is bandwidth for memory. The method does not make a big model fast; it makes it possible.
Three runtimes: llama.cpp, MLX with a split step, and mlx-sniper
The repository reaches for a different tool in every band of model size, and the switching points are not stated as such.
The fast path is llama.cpp, installed through a package manager, serving a GGUF file fetched from a community re-quantization of the Qwen models. Server flags are spelled out: flash attention on, a context size of 12288 for the 35B and 65536 for the 9B, both key and value caches quantized to four bits, all GPU layers offloaded, reasoning off, four threads, and the server bound to the loopback interface on port 8000. The 9B example drops the single-slot flag the 35B one sets.
The dense streaming path is a different stack. It installs an MLX library and transformers, downloads a community MLX conversion, runs a split script, and then runs a separate streaming script. The split step is a prerequisite, not an optimisation.
The largest models use a third tool, one that downloads and runs a model directory directly, and the disk table runs from a 30B at seventeen gigabytes to a 235B at roughly a hundred and thirty, with headroom to about a hundred and forty. The documentation suggests an external NVMe drive for those, and gives the two commands for it.
So the answer to whether this runs on your Mac depends on which of three stacks you are willing to set up.
The headline capability needs a preprocessing step before it runs at all
Flash streaming is presented as the thing that lets you run models that genuinely do not fit in RAM at full quality. The line under the command is the catch:
It requires pre-built stream files, and the split and rebuild tools live in the research directory.
So the model files this method needs are not downloadable. You produce them locally, once, from a full model you download anyway. For the dense 27B the sequence is explicit: install two packages, snapshot-download roughly sixteen gigabytes, run a split script, then run the streaming script. There is no note about how long the split takes or how much extra disk it needs at peak, and the disk table covers the post-processed sizes rather than the working set.
The research directory is described as a journey, with each file representing a step, which is a reasonable way to present an experimental directory and also a warning that the scripts are not a supported interface.
One number in there needs its label read carefully. The batched union-of-experts prototype reports 5.1 tokens per second, which sits in the same range as the generation numbers above it, but the documentation states plainly that this is verification speed for checking draft tokens and not generation speed, and that it is useful for speculative decoding. So it cannot be compared with anything else in the README.
Both recommended installs use --break-system-packages
The quick start begins with a package manager install and then this:
pip3 install rich ddgs --break-system-packagesThe same flag appears again on the dense streaming path, for the MLX library and transformers. Both are labelled as the way to install.
The flag exists to override the protection newer Python distributions put on installing into the system interpreter, so the documented setup deliberately bypasses it. On a machine where that protection is on, which is most current Python installations on macOS, the alternative would be a virtual environment, and the documentation does not mention one.
There is also no dependency file to inherit a sane answer from. The repository root has no requirements file and no packaging manifest; the only dependency declarations are the install lines in the documentation plus whatever setup.sh does, and setup.sh is described as a one-command install covering llama.cpp, the model download and the configuration.
The dependencies themselves are small and sensible, a terminal formatting library and a search backend for the agent, plus the inference engine and the model weights.
No licence file, and an agent file the documentation never names
Two housekeeping details worth catching before you read the rest of this project as a description of itself.
The repository root has no licence file, and the platform's own metadata records the licence as unknown. For code you are being invited to run locally, install packages for, and drive with your own API keys, that is the first thing to settle.
Second, the root also contains an instructions file for coding agents that appears nowhere in the documentation. The Files section describes the production agent, a simple chat script, a monitoring dashboard for the inference server, a setup script, an example configuration and a small web interface with its own server and page. It does not mention the agent instructions file that sits beside them.
The agent itself is small and legible, which is a virtue. Slash commands cover agent mode as the default, with search, shell and chat; a raw streaming mode with no tools at all; a quick web search; session statistics; conversation reset; and exit. The separation between the tool-using mode and the raw mode is the kind of boundary that makes an agent script auditable.
The file listing for the streaming research directory begins and then ends mid-row, so the scripts named in that section are the only description of them.
Editorial conclusion
There are two different projects here and they should be judged separately. The fast path is a standard llama.cpp setup with a two-bit quantized mixture-of-experts model that fits in memory, which is useful today and delivers the headline number honestly for what it is. The research path is the interesting one: streaming feed-forward weights from SSD so a full four-bit model runs on a machine it does not fit in, which works and is slow, in the range of one to five tokens per second on the machine everything was measured on. Before you rely on either, three things. Read the quality claim for what it rests on, since a table of measurements is validated by five prompts from cold start. Expect to run a split step before the streaming path works at all, since the model files it needs are built locally rather than shipped. And note that the measurements are all from one Mac mini configuration, while the compatibility table generalises to any Mac with 8 or 16 GB, which is a wider claim than the evidence supports.
Frequently asked questions
What does walter-grace/mac-code actually do?
It gives a Mac a local agent backed by llama.cpp, plus a research path that streams feed-forward weights from SSD so a full 4-bit model can run on a machine it does not fit in. Attention weights, embeddings, norms and the KV cache stay pinned in RAM while the bulk of the model is loaded per layer, used for one matrix multiplication and discarded, so memory stays flat.
How fast is walter-grace/mac-code, and on what?
Every measurement in the documentation was taken on a 16 GB Mac mini M4. The fast path, a two-bit quantization of a 35B mixture-of-experts model that fits in RAM, runs at 30 tokens per second. The full 4-bit streaming results range from 0.15 tokens per second for a dense 32B to 5.4 for a mixture-of-experts 35B.
Does walter-grace/mac-code need pre-built model files?
For the flash streaming path, yes. The documentation says pre-built stream files are required and points at the split and rebuild tools in the research directory. For the 27B dense path the sequence is to download roughly sixteen gigabytes, run a split script, and then run the streaming script.
What licence is walter-grace/mac-code under?
None is declared. There is no licence file in the repository root and the platform records the licence as unknown. For code intended to be run locally against your own keys, the terms need to be settled with the author before use.
How do I install walter-grace/mac-code?
There is no dependency file to install from. The quick start uses a package manager to install the inference engine and then pip3 with the break-system-packages flag to install a formatting library and a search backend. The same flag is used on the dense streaming path for the MLX library and transformers. A setup script in the repository root is described as a one-command install covering the engine, the model download and the configuration.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/walter-grace-mac-code)