# FastLLM: a C++ inference engine with no PyTorch underneath

> An Apache licensed inference engine that reads Hugging Face weights directly, spreads dense and MoE layers across GPU, CPU, NUMA and disk, and serves OpenAI and Anthropic compatible endpoints.

**ztxz16/fastllm** — fastllm是后端无依赖的高性能大模型推理库。同时支持张量并行推理稠密模型和混合模式推理MOE模型，任意10G以上显卡即可推理满血DeepSeek。双路9004/9005服务器+单显卡部署DeepSeek满血满精度原版模型，单并发20tps；INT4量化模型单并发30tps，多并发可达60+。

- Repository: https://github.com/ztxz16/fastllm
- Stars: 5,093 · Forks: 502
- Language: C++
- License: Apache-2.0
- Published: 2026-10-07 · Updated: 2026-10-07 · Language: en
- Canonical page: https://hysenlabs.com/projects/ztxz16-fastllm

## No PyTorch between the weights and the GPU

The repository description is in Chinese and makes a precise claim: a high-performance large model inference library with no backend dependency, supporting tensor-parallel inference for dense models and MoE inference for mixture models, capable of running a full DeepSeek on any GPU above 10 GB, and citing 20 tokens per second at single concurrency for a full precision DeepSeek on dual 9004 or 9005 servers, 30 for INT4, and 60 or more at higher concurrency.

Those numbers are the author's own claims and they are specific enough to be useful or to be wrong, and this repository offers no benchmark data to check them against. There are no releases published on GitHub and no benchmark file in the tree. What is in the tree is `docs/benchmark.md`, linked from the README, which is where any real evidence lives.

The architectural claim is easier to verify and more interesting. The README states that the core runtime is implemented in C++ and does not depend on PyTorch, supports dense and MoE models, and offers CUDA, ROCm, CPU, NUMA and disk-hybrid inference plus multi-card tensor parallelism. GitHub reports the language as C++, the license as Apache-2.0, and the default branch as `master`. The tree backs the claim up: `main.cpp`, `src/`, `include/`, `CMakeLists.txt`, `third_party/` and `pyfastllm/`.

## Layer placement is the feature, not a footnote

Almost everything unusual in this engine comes from one decision: it does not require a model's weights to fit in video memory. The README describes deploying ordinary layers and expert layers separately across CUDA, CPU, NUMA or disk, and combining several kinds of hardware in proportions you choose. That is aimed at machines with limited VRAM and generous host memory or SSD, which is a large fraction of real machines running models at home or in a lab.

The rest of the core capability list follows from it. Tensor parallelism including an odd number of cards, dynamic batching, streaming output, Paged KV Cache, prefix caching, chunked prefill and CUDA Graph. Speculative decoding is offered through several paths: MTP, built-in or external DSpark, and DFlash2, for the models that support it.

On formats and precision, the README lists Hugging Face Safetensors, the FastLLM export format, AWQ and partial GGUF, with FP16, BF16, FP8, NVFP4, MXFP4, INT4 and K-Quant paths chosen per model and hardware. The backends are built-in CPU, CUDA and ROCm operators, with optional Triton operators and a path for custom Python model graphs and other accelerator backends.

One line in the README deserves to be read twice, because it is the honest one: operator support is not identical across models, quantization formats and hardware backends, so verify accuracy, VRAM use and throughput on your target combination before deploying.

## Which models the current mainline covers

The README has dropped earlier models from the front page and now lists what the development mainline focuses on, in a table: Qwen covering Qwen4-Exp, Qwen3.8-Flash-Next and Qwen3.5 through 3.8; DeepSeek covering V4 and V4-Flash with sparse attention and built-in DSpark; Kimi covering K3 with KDA or MLA attention and CUDA, NUMA, CPU or GPU experts plus disk experts; GLM covering 5, 5.3-Flash and quantized KV-cache CPU inference; and others listed as Dots3-Note, Laguna, HY-V3, Step3.5 and 3.7, MiniMax-M2 and Gemma4.

The table also records what each family does differently, which is the detail that matters if you are choosing an engine. Qwen4-Exp does not load vision weights. Qwen3.8-Flash-Next can load MTP weights on demand with a flag to enable speculative decoding. GLM-5.3-Flash is listed with paged cache support, and GLM-5.2 with quantized KV cache for CPU inference.

Older model compatibility is explicitly not deleted but relocated: the README says it can still be looked up in the supported models list under `docs/models.md`, with the version log at `docs/version.md` as the authority on recent adaptations and limits. A reader evaluating this engine should go to that version log rather than assume the table is exhaustive, because for a project moving this quickly it is a moving target.

## Five commands cover nearly everything

Installation on Linux with an NVIDIA GPU is a single pip line into a virtual environment, which the README recommends creating:

```bash
python -m pip install -U ftllm
```

Windows with NVIDIA uses the same command, with one precondition: if the first install reports a missing DLL, a separate dependency wheel has to be installed first. Linux with an AMD GPU goes through the ROCm install and compile instructions in `docs/rocm.md` rather than a wheel. CPU-only, unusual architectures and other accelerators build from source with CMake. The README also warns that a Conda environment can produce dynamic library conflicts and suggests a clean `venv` instead.

The verification step uses a deliberately small model, Qwen3-0.6B, and the README is explicit that it is only a quick download for a smoke test and does not represent the current mainline models. The daily commands then look like this:

```bash
ftllm server Qwen/Qwen3-0.6B
ftllm run Qwen/Qwen3-0.6B
ftllm webui --api_base http://127.0.0.1:8080/v1
```

The API server listens on `0.0.0.0:8080` by default, the WebUI listens on `127.0.0.1:1616` and connects to a server you already started, and `ftllm launch` with no arguments starts a browser-based management page on `127.0.0.1:8000`. The terminal deployment wizard is `ftllm tui` and the benchmark tool is `ftllm bench`, which takes a device and token counts.

## Serving OpenAI, Anthropic and everything above them

The API server is the part most teams will actually use, and it accepts a local path as readily as a repository ID:

```bash
ftllm server /data/models/my-model --device cuda
```

A named local model with an explicit bind address and key looks like this:

```bash
ftllm server /data/models/my-model \
  --model_name local-model \
  --host 0.0.0.0 --port 8080 \
  --api_key local-key
```

Calling it is a standard OpenAI Chat Completions request, which means existing clients and SDKs need no modification:

```bash
curl http://127.0.0.1:8080/v1/chat/completions \
  -H "Authorization: Bearer local-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "local-model",
    "messages": [{"role": "user", "content": "你好，请介绍一下 FastLLM。"}],
    "stream": false
  }'
```

The server also provides the OpenAI Responses API and the Anthropic Messages API. Where the model supports it you get separated reasoning content, tool calling, streaming responses, cache hit statistics, server-side sampling parameters and startup progress events.

One operational warning is easy to miss. The README states that listening on anything other than localhost uses unencrypted HTTP and should only be done on a trusted network, and that exposing a port to the public internet additionally requires firewall and cloud security group changes plus NAT port forwarding, which the launcher will not detect for you.

## The launcher, the studio and a directory agent with real permissions

The `ftllm launch` page is more than a model browser. It can download models from ModelScope, save launch configurations, preview the resulting command, and then host either a `ftllm server` or a chat `ftllm webui` for you. When you add a local model it recommends tensor parallel, MoE hybrid and N-gram storage parameters based on model structure, weight size and your GPU, memory and NUMA topology, and you can re-run or clear that analysis.

It closes the launcher and the downloads and model processes it was hosting with it, shares its configuration file with the terminal wizard, and ships a light and dark theme. Customisation goes further than themes: the interface supports modifying the page, studio features, status bar and global skin through natural language conversation, with generated content saved under `~/.fastllm/plugins` so core code is never edited.

The part worth thinking hardest about is the directory agent. A working directory can be set with a root path that defaults to the home directory, and the README is explicit that a directory agent can modify files and execute commands, so it should only be exposed to trusted users. There is a flag to disable it entirely, which also disables directory browsing and saved tasks while leaving ordinary chat working. Code analysis and web search run on a separate Pi agent runtime, installed on Linux x86-64 through the studio interface as `ftllm-agent-runtime==0.3.3`, roughly 43 MB, bundling the runtime plus `rg` and `fd`, with a builtin single-turn path available as a fallback.

## Two repositories in one: the modern CLI and an older webui build

There is a genuine tension inside this repository that a reader should resolve before deploying. The README documents a current generation of entry points: `ftllm run`, `ftllm server`, `ftllm webui`, `ftllm launch`, `ftllm tui` and `ftllm bench`, with the server on 8080 and the WebUI on 1616. The `Dockerfile` and `docker-compose.yaml` in the tree belong to an earlier generation. They build a single binary and run it as the container's main command:

```yaml
command: /fastllm/build/webui -p /models/chatglm2-6b-int8.flm -w ./example/webui/web
```

with port 8081 inside the container mapped to 11234, a model mounted from `./models/`, a CUDA 12.1 devel base image and a build step compiling for architectures 80, 86, 89 and 90. The README's current text also states the opposite architecture: the WebUI does not load a model in its own process and requires a running OpenAI compatible API Server first, which is a different design from a single binary doing both.

So there are two facts, and neither cancels the other. The pip package and the README describe a multi-process, OpenAI compatible design. The container files describe a self-contained webui binary from an earlier layout. The practical rule is to trust whichever path you actually installed: if you installed from PyPI, follow the README commands and use `ftllm server`; if you built the container, expect the compose file's port mapping and single command, and check `docs/` for whether that layout is still maintained.

GitHub reports the last push on 2026-09-22, 5071 stars, 500 forks and 340 open issues, and the repository is not archived. There is an `AGENTS.md`, a `ToolCall-SGLang-PLAN.md`, a `desktop/` directory and an `README_EN.md`, which together suggest a project actively reorganising its own tooling as much as its inference paths.

## Conclusion

The case for FastLLM is narrow and clear: a machine with more host memory or SSD than VRAM, running a mixture of expert models where you need to choose precisely which layers live where, and an appetite for building against source rather than installing a wheel. What the repository does not settle is whether your specific model, quantization and hardware combination has working operators, which the README warns about in its own words. Install into a clean virtual environment, smoke test with the small Qwen model the README uses, and read the deployment guide for your model in `docs/` before committing a production configuration.

## FAQ

### Does FastLLM require PyTorch?

No. The README states the core runtime is implemented in C++ and does not depend on PyTorch, and GitHub lists the repository language as C++. Built-in CPU, CUDA and ROCm operators are provided, with optional Triton operators as an alternative.

### Which models does FastLLM support?

The current development mainline covers Qwen4-Exp and Qwen3.8-Flash-Next, Qwen3.5 through 3.8, DeepSeek V4 and V4-Flash, Kimi K3, GLM 5 and 5.3-Flash, and others including Dots3-Note, Step3.5 and 3.7, MiniMax-M2 and Gemma4. Earlier models are listed in the supported models file under docs.

### Can FastLLM run a model that does not fit in VRAM?

That is the design centre of the project. The README describes placing ordinary layers and expert layers separately on CUDA, CPU, NUMA or disk and combining hardware types in proportions you choose, for machines with limited VRAM and generous host memory or SSD.

### What API does the FastLLM server expose?

It exposes OpenAI Chat Completions, the OpenAI Responses API and Anthropic Messages, so existing clients work without changes. Where a model supports it you also get separated reasoning content, tool calling, streaming and cache hit statistics.

### How do I install FastLLM?

On Linux or Windows with an NVIDIA GPU, use python -m pip install -U ftllm inside a virtual environment. Windows may need a dependency wheel first if the install reports a missing DLL. AMD GPUs follow the ROCm build instructions, and CPU-only or other architectures build from source with CMake.

### Is the directory agent in the launcher safe to enable?

It is powerful by design and the README says a directory agent can modify files and execute commands, so it should be enabled only for trusted users. There is a flag to disable the directory agent entirely, which turns off directory browsing and saved directory agent tasks while leaving ordinary chat working.

## Sources

- [Issues](https://github.com/ztxz16/fastllm/issues)
- [License: Apache-2.0](https://github.com/ztxz16/fastllm/blob/master/LICENSE)
- [README](https://github.com/ztxz16/fastllm/blob/master/README.md)
- [ztxz16/fastllm on GitHub](https://github.com/ztxz16/fastllm)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ztxz16-fastllm
