# FreeToken: an edge-native MoE serving engine for consumer GPUs

> FreeToken is an Apache-2.0 Python runtime from FlashML that runs frontier-scale Mixture-of-Experts models on laptops and gaming desktops by treating GPU, CPU and host memory as one elastic inference platform. The engine is credible; the packaging story is not yet finished.

**FlashML-org/FreeToken** — FreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast and efficiently.

- Repository: https://github.com/FlashML-org/FreeToken
- Website: flashml.ai
- Stars: 13,930 · Forks: 1,373
- Language: Python
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/flashml-org-freetoken

## What FreeToken actually solves, and for whom

A frontier open-weight MoE checkpoint is large. FreeToken's README describes the target as "290B+ frontier MoE models" running locally "on your gaming PC". The gap it addresses is not model access, since the weights are already public, but the serving layer: how to keep a model that does not fit in VRAM answering requests at interactive speed when the only hardware available is one consumer GPU plus system RAM.

The intended user is a developer running coding or tool-calling agents against a local endpoint. The README names Codex, Claude Code, OpenCode, OpenClaw and DeepSeek Harness as clients, and the project exposes OpenAI- and Anthropic-compatible APIs so those tools can point at it without modification. That framing matters: FreeToken is not a research harness for measuring throughput, it is a serving engine meant to sit behind an agent that makes many short calls.

It is not for someone who wants a managed endpoint, nor for a team that needs a supported distribution. The pyproject classifiers list Linux and NVIDIA CUDA only, and the project is self-described as Beta.

## The mechanism: bandwidth-adaptive CPU-GPU co-execution

MoE models activate a small fraction of their experts per token, so the working set at any moment is much smaller than the checkpoint. FreeToken's design leans on that. The README lists a "bandwidth-adaptive CPU-GPU co-execution ($q^\star$ policy)", meaning the engine decides per layer how much expert computation to place on the GPU versus the CPU based on measured bandwidth rather than a fixed split. The arXiv paper title, "Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution", confirms that this policy is the central contribution.

Three supporting pieces appear in the README. Full-layer double-buffered prefill streaming overlaps weight movement with computation during the prompt phase. A global LRU expert cache keeps recently used experts resident, and pyproject.toml names the kernel behind it: "slot_cache: the device-side LRU admission kernel behind the MoE expert cache", shipped in flashlib==0.3.0. FTW is the project's own weight format, with a separate docs/ftw-hotfix.md for repairing older checkpoints.

The second mechanism is caching at the context level. Semantic anchor checkpoints for recurrent state and KV caches let an agent edit its context (a tool call, a thinking block) without recomputing everything before the edit. For an agent loop that appends and rewrites context constantly, that is the difference between a usable and an unusable round trip.

The third is memory management. The README states VRAM can be re-allocated at runtime between expert caches and KV memory "without engine restarts or weight reloading". Reloading a large checkpoint is slow enough that avoiding it is a real feature, not a nicety.

## Installing FreeToken and running a first model

The README offers two paths. The desktop app for Windows or Linux is downloaded from flashml.ai and sets up the engine with a GUI. The CLI path installs the Python package, and the README recommends uv.

```bash
uv pip install "freetoken[accel]"
```

The accel extra is what pulls in the accelerated runtime. The README does not enumerate what the extra contains, so treat the install as the point where you find out whether your CUDA toolchain is compatible.

If you prefer to build from source, the README gives this sequence. It creates a virtual environment, activates it, and installs the package in editable mode with the same extra.

```bash
git clone https://github.com/FlashML-org/FreeToken.git && cd FreeToken
uv venv && source .venv/bin/activate
uv pip install -e ".[accel]"
```

A source build is not a formality here. setup.py calls _check_toolchain(), which runs check_nvcc_matches_torch() before configuring extensions, and raises a RuntimeError if CUDA_HOME is unset because the pinned_tensor extension links against cudart. The build also compiles a CPU MoE executor for a moe-backend cpu option, using avx512bf16 and avx512f target attributes with a runtime __builtin_cpu_supports dispatch. Expect the build to fail loudly on a mismatched toolkit rather than produce a subtly wrong binary.

After install, the README points to docs/quickstart.md and docs/cli.md for the actual invocation, and docs/models.md for the supported checkpoints, which the README lists as DeepSeek-V4-Flash, Qwen3.6-35B-A3B and GLM-5.2. The README does not print a launch command, so the CLI reference is the file to read before your first run.

## Where FreeToken will disappoint you

The packaging is the weakest part of the story. pyproject.toml pins apache-tvm-ffi==0.1.13.post3 and flashlib==0.3.0 exactly, and its own comment says the version ranges "are the contract; there is no lockfile". Two exact pins plus a torch window of >=2.11,<2.12 means a conflict with an existing environment is likely, and there is no lockfile to resolve it for you. The build-system comment is explicit that a torch mismatch links the C++ extensions against the wrong libtorch, which is the kind of failure that surfaces at runtime rather than install time.

Platform support is narrow. The classifiers list POSIX Linux and NVIDIA CUDA, with RTX 30, 40 and 50 series called out in the README. There is no macOS classifier and no non-NVIDIA accelerator classifier. If your laptop is an Apple Silicon machine, this project's own metadata does not claim to support you.

The documentation set is also uneven. The README links install, quickstart, models, CLI and an FTW hotfix page, but it does not document rollback, and it does not state memory requirements per model. Deciding whether a given checkpoint fits your card requires reading docs/models.md and finding out empirically.

Finally, Beta is the project's own classification. The nightly channel exists alongside tagged releases, which suggests the surface is still moving.

## How FreeToken differs from llama.cpp and vLLM

The obvious local-inference comparison is llama.cpp, and the difference is architectural. llama.cpp is built around GGUF quantization and a CPU-first execution model with optional GPU offload; FreeToken is built around MoE expert routing, a device-side LRU expert cache, and a policy that splits each layer between CPU and GPU according to measured bandwidth. For a dense model on a CPU, llama.cpp is the natural choice. For a sparse MoE checkpoint whose experts do not all fit in VRAM, FreeToken's design targets exactly that case.

The server-side comparison is vLLM. vLLM assumes a datacenter GPU with enough VRAM for the weights and optimizes throughput and batching across many concurrent requests. FreeToken assumes the opposite: one consumer card, weights partly in host memory, and a latency budget for interactive agent calls. The README's own acknowledgment lists vLLM, SGLang, FlashInfer, flash-linear-attention, LightLLM and llama.cpp as projects it learned from or reused code from, and names mini-sglang as its deepest inspiration. So this is not a from-scratch design; it is a re-targeting of serving-engine ideas toward a single-machine, memory-constrained deployment.

The practical consequence: if your model fits comfortably in VRAM, FreeToken's offload machinery buys you little and you are better served by a conventional server. FreeToken earns its place when the model does not fit.

## Maintenance, releases and what the Apache-2.0 licence means here

The repository is not archived and the last push was on 2026-09-16, the same day v0.1.3 was tagged. v0.1.2 came on 2026-08-19 and a nightly build on 2026-09-09, so the release cadence over that window is roughly monthly for tagged versions with a rolling nightly in between. That is a young but currently moving project, and the nightly channel is the one to watch if you need a fix that has not been tagged.

Upgrade cost is the part to plan for. With no lockfile and two exact pins in pyproject.toml, upgrading FreeToken means re-resolving the whole dependency set, and the torch window will pull you forward whenever you cross a boundary. If you build from source, every upgrade re-runs the nvcc-versus-torch check, so a CUDA toolkit upgrade and a FreeToken upgrade are coupled events. Budget for a rebuild, not a version bump.

The licence is Apache-2.0, declared as an SPDX string in pyproject.toml with license-files = ["LICENSE"]. That is a permissive licence with an explicit patent grant, and it is compatible with commercial deployment. The README's acknowledgment section lists SGLang, vLLM, FlashInfer, flash-linear-attention, LightLLM and llama.cpp as sources of reused code, all of which are themselves permissively licensed, but if you redistribute a modified FreeToken you should read the LICENSE file and those upstream notices rather than assume the single SPDX string covers everything. This is not legal advice.

## Conclusion

Adopt FreeToken if you have a single NVIDIA RTX 30, 40 or 50 series card and you want to run a large MoE checkpoint locally behind an OpenAI- or Anthropic-compatible endpoint, and you are willing to build from source when the wheel does not fit your CUDA setup. Do not adopt it if you need macOS, non-NVIDIA accelerators, or a support contract, because the classifiers list Linux only and the dependency ranges are floors rather than a lockfile. Before committing, verify three things: that your target model appears in docs/models.md, that the torch range in pyproject.toml (torch>=2.11,<2.12) matches the CUDA build you already have, and that the install.md path works on your machine rather than the desktop download. The repository's own pin on apache-tvm-ffi==0.1.13.post3 and flashlib==0.3.0 tells you how tightly the runtime is coupled to specific kernel releases.

## FAQ

### What is FreeToken?

FreeToken is an edge-native Mixture-of-Experts serving engine from FlashML that runs frontier open-weight MoE models on personal and consumer hardware. It exposes OpenAI- and Anthropic-compatible APIs so coding and tool-calling agents can point at it.

### What are tokens and how do they work?

The repository does not explain tokenization; it is a serving engine that consumes and produces tokens through its API, and the README does not cover token mechanics.

### Which models does FreeToken support?

The README lists DeepSeek-V4-Flash, Qwen3.6-35B-A3B and GLM-5.2, across MXFP4, NVFP4, FP8 and BF16 quantization formats. The README points to docs/models.md for the full list.

### What hardware does FreeToken require?

The README states native support for NVIDIA RTX 30, RTX 40 and RTX 50 series GPUs, scaling across consumer laptops, gaming desktops and workstation GPUs. The pyproject classifiers list POSIX Linux and NVIDIA CUDA, and there is no macOS classifier.

### How do I install FreeToken from the command line?

The README recommends uv and gives uv pip install "freetoken[accel]", or a source build that clones the repository into a uv virtual environment and installs with the accel extra. A source build requires CUDA_HOME to be set.

## Sources

- [FlashML-org/FreeToken on GitHub](https://github.com/FlashML-org/FreeToken)
- [Issues](https://github.com/FlashML-org/FreeToken/issues)
- [License: Apache-2.0](https://github.com/FlashML-org/FreeToken/blob/main/LICENSE)
- [README](https://github.com/FlashML-org/FreeToken/blob/main/README.md)
- [Releases](https://github.com/FlashML-org/FreeToken/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/flashml-org-freetoken
