# Flash-MoE: Running a 397B Parameter Model on a Laptop

> Flash-MoE is a C, Objective-C and Metal inference engine that streams a 397B parameter Qwen3.5 MoE model from SSD on a 48GB MacBook Pro. It is a research artifact with a narrow hardware target, not a general serving stack.

**danveloper/flash-moe** — Running a big model on a small laptop

- Repository: https://github.com/danveloper/flash-moe
- Stars: 4,182 · Forks: 513
- Language: Objective-C
- License: not declared
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/danveloper-flash-moe

## The problem Flash-MoE solves: a 397B model that does not fit in RAM

A 397 billion parameter Mixture-of-Experts model does not fit in 48GB of unified memory, and the README is explicit that the weights are not all resident. The engine keeps only the non-expert weights mapped and streams expert weights from disk. The README states the entire 209GB model streams from SSD through a custom Metal compute pipeline, with no Python and no frameworks.

The target user is narrow. The documented machine is a MacBook Pro with an Apple M3 Max, 48GB unified memory, and a 1TB SSD measured at 17.5 GB/s sequential read running macOS 26.2. The README reports 4.36 tokens per second with 4-bit experts and full tool calling, and notes that 2-bit quantization reaches 5.74 tokens per second but produces malformed JSON. Anyone without that class of hardware is reading the paper, not running the binary.

## How the SSD expert streaming pipeline works

The model has 60 transformer layers: 45 GatedDeltaNet linear attention layers and 15 standard full attention layers. Each layer holds 512 experts, and K=4 are activated per token plus one shared expert, with a hidden dimension of 4096. Only those active experts are read from disk, roughly 6.75MB each at 4-bit.

The README describes the per-layer pipeline as a fixed sequence. CMD3 from the previous layer is already in flight. CMD1 computes attention projections and the delta-net on GPU (1.22ms), the CPU flushes results, CMD2 handles o_proj, norm, routing and the shared expert (0.55ms), the CPU does softmax and topK routing (0.003ms), parallel pread pulls the K=4 experts (2.41ms), and CMD3 encodes the expert forward pass plus combine and norm without waiting. The README gives 4.28ms as the average per-layer time at 4-bit.

The design decision that stands out is what the README calls Trust the OS. There is no custom expert cache. The OS page cache, around 35GB, manages expert data with standard LRU, and the README states it reaches roughly 71% hit rate naturally. The authors report that every custom cache they tried (Metal LRU, malloc cache, LZ4 compressed cache) was slower because of GPU memory pressure or overhead. A second constraint is hardware-level: on Apple Silicon, SSD DMA and GPU compute share the memory controller, so the README argues the serial GPU to SSD to GPU pipeline is optimal rather than a compromise.

## Installing Flash-MoE and running a first prompt

The README points at the metal_infer directory for the engine. The build uses the Makefile in that directory, and the 4-bit path expects a packed_experts/ directory to already exist. The README does not describe downloading prebuilt weights, so treat weight preparation as a step you have to work out from the scripts in the repository.

Build the engine first:

```bash
cd metal_infer
make
```

Then run a short generation. The README gives this exact invocation, and you should see token-by-token output in the terminal:

```bash
./infer --prompt "Explain quantum computing" --tokens 100
```

If you want the per-layer timing breakdown the README mentions, add the timing flag:

```bash
./infer --prompt "Hello" --tokens 20 --timing
```

For interactive use there is a separate binary. The README describes chat.m as an interactive chat TUI with tool calling:

```bash
./chat
```

The 2-bit path is a different flag on the same binary, and the README warns it breaks tool calling:

```bash
./infer --prompt "Explain quantum computing" --tokens 100 --2bit
```

Weight preparation is not covered by a single documented command. The repository contains extract_weights.py, which the README says creates model_weights.bin from safetensors, and repack_experts.py for 4-bit expert packing. The README does not document the arguments either script takes.

## Where Flash-MoE breaks: 2-bit output, memory contention, and missing docs

The clearest failure mode is documented by the authors themselves. At 2-bit quantization the model emits a backslash where a quote belongs in JSON, so tool calling becomes unreliable. The README calls 4-bit the production configuration and labels the 2-bit rows as not suitable for tool use. If your workload depends on structured output, the faster configuration is off the table.

The second limitation is the memory controller. The README states that even small background SSD DMA causes disproportionate GPU latency spikes through memory controller arbitration, and that a prefetch attempt (F_RDADVISE) netted 0% because SSD DMA slowed the GPU by 73%. This is not a bug to fix; it is a property of the hardware the engine is tuned for.

The documentation has gaps that matter before you invest time. The README does not document rollback, versioning, or a supported upgrade path, and there are no retrieved releases. The licence is not stated in the repository metadata available, which is a real blocker for anyone evaluating redistribution. The README also does not document how to obtain the base model weights or how long weight preparation takes.

## Flash-MoE compared with MLX and llama.cpp style serving

The related searches include mlx and flash moe mlx, which is a fair comparison point. MLX is Apple's array framework with a Python-first interface and a model conversion pipeline; Flash-MoE is the opposite approach. It is a single ~7000-line Objective-C file plus ~1200 lines of Metal shaders, with a C BPE tokenizer chosen because the README reports 180ms startup versus 3500ms for the alternative, a 20x difference.

The trade-off is flexibility. With MLX you get a maintained framework, Python tooling, and a broader model zoo. With Flash-MoE you get a hand-tuned pipeline for one model family on one hardware profile, and you inherit the responsibility for weight extraction and packing. The README's experiment log in results.tsv lists 58 discarded approaches, including LZ4 expert compression at minus 13%, temporal expert prediction at minus 18% with a 25% hit rate, and an MLP routing predictor at 31% accuracy. That log is the honest signal here: the current design is the survivor of a lot of dead ends, and the dead ends are documented rather than hidden.

## Maintenance cost and licence status

The last push to the repository was on 2026-03-19, which is roughly six months before today. There are no retrieved releases, so there is no tagged version to pin and no changelog to read. The repository is not archived, but a single-author research artifact with no releases means you should expect to read the source rather than the docs when something changes.

The practical upgrade cost is weight repacking. Moving between 4-bit and 2-bit requires repack_experts_2bit.py according to the README, and the 4-bit path needs repack_experts.py. Any change to the expert layout means regenerating those artifacts, which are large. Budget disk space accordingly: 209GB for 4-bit and 120GB for 2-bit per the results table.

The licence is not identified in the repository metadata available, so the redistribution terms are unknown. That is not a legal conclusion, just a gap you have to close with the repository owner before shipping anything built on this code.

## Conclusion

Adopt Flash-MoE if you have an Apple Silicon Mac with enough unified memory and SSD bandwidth to stream a 209GB 4-bit expert set, and you want to read or modify the Metal kernels rather than call an API. Do not adopt it if you need a supported, portable serving stack, CUDA or Linux, or reliable tool calling at 2-bit. Before committing, verify that packed_experts/ and model_weights.bin exist in the layout the Makefile expects, that your SSD read bandwidth is in the range the README describes, and that your output stays valid JSON at the quantization you pick.

## FAQ

### What is Flash-MoE?

It is a pure C, Objective-C and Metal inference engine that runs the Qwen3.5-397B-A17B Mixture-of-Experts model on a MacBook Pro with 48GB RAM, streaming the 209GB 4-bit weight set from SSD. The README reports 4.36 tokens per second with full tool calling.

### Is GLM 4.7 a flash MoE model?

The README describes Flash-MoE as an inference engine for Qwen3.5-397B-A17B, a 397 billion parameter Mixture-of-Experts model with 512 experts per layer and K=4 activated per token. It does not mention GLM 4.7.

### Does Flash-MoE run on Ubuntu or Linux?

The README describes a macOS stack built on Metal compute shaders and Apple Silicon unified memory, with macOS 26.2 on an M3 Max as the documented machine. No Linux or CUDA path is described.

### How much disk space does Flash-MoE need for its weights?

The README's results table lists 209GB on disk for the 4-bit expert configuration and 120GB for the 2-bit configuration. The non-expert weights are a separate 5.5GB model_weights.bin file that is mmap'd.

### Does Flash-MoE use a custom expert cache?

No. The README describes a Trust the OS principle where the OS page cache manages expert data with standard LRU, reaching roughly 71% hit rate naturally. Every custom cache the authors tried was slower.

## Sources

- [danveloper/flash-moe on GitHub](https://github.com/danveloper/flash-moe)
- [Issues](https://github.com/danveloper/flash-moe/issues)
- [README](https://github.com/danveloper/flash-moe/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/danveloper-flash-moe
