Colibri runs a 744B MoE model off disk in C, and promises nothing about speed while promising everything about semantics
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine, pure C, zero deps, experts streamed from disk. Tiny engine, immense model.
At a glance
- What is it?
- Colibri is a pure C inference engine with no engine dependencies that treats VRAM, RAM, and NVMe as a single placement hierarchy, so models from 744B to 2.8T parameters run on a consumer machine with experts streamed from storage. It ships nine model families behind one front end, and its own policy is unusually explicit: no service level on speed, and an absolute rule that a memory shortage may never quietly change the model's precision or router behaviour.
- Who is it for?
- Colibri earns a place on the desk of a researcher who wants to watch a frontier mixture-of-experts model actually run, measure where a turn goes, and change the code that does it, or of a developer locked out of a frontier model by hardware.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly C, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
VRAM, RAM, and NVMe are one hierarchy, not three capped tiers
The central idea is that a mixture-of-experts model does not need its experts resident to run. Colibri treats video memory, system memory, and NVMe storage as placement tiers for the same weights, so a limited amount of fast memory changes throughput rather than changing what the model is. The claim in the project's own words is that the hierarchy is not limited by tier capacity, and the second core technique spells out the mechanism: a JIT for weights, where measured routing heat drives a per-layer LRU, a learned pinned hot-store, and one-layer-ahead prefetch, instead of loading every expert for every token. The dashboard's profiling page shows what that looks like end to end. A recorded CPU-only run of Qwen3.6 spent 19.0 seconds of wall time on 36 prompt and 55 generated tokens, reached 2.9 tokens per second, and overlapped 11.4 seconds of disk service with compute.
No service level on speed, and a hard rule on semantics
The project publishes a policy that is unusual enough to quote, because most inference tools promise the opposite. Colibri is described as deliberately a place to test aggressive systems ideas, so there is no service level agreement on speed, while there is a hard guarantee on semantics. Concretely: experiments must earn their place through reproducible end-to-end measurements, and the default policy never silently changes model precision or router semantics. Insufficient fast memory may reduce speed, but it must not quietly redefine the model. That is the contract to hold a deployment to. If an answer changes because a machine ran short of RAM, that is a bug against the stated policy rather than a performance trade-off, and the rest of the writing reinforces it, with memory, latency, and correctness properties presented as properties and not as a blanket throughput claim.
The weight JIT wins on repeat workloads and admits where it loses
The interesting material is in the caveats, because they are attached to the same bullets as the claims. The weight JIT is described as winning on repeatable workloads, with two named failure modes. History can overfit, meaning a pinned hot-store tuned to yesterday's traffic can be wrong for today's, and lookahead can lose on some hosts, meaning the one-layer-ahead prefetch is a net cost rather than a saving on particular machines. Both are described as remaining measurable policies rather than promises, which is the framing to hold on to. The same honesty appears in the speculation work, where native multi-token prediction and grammar-forced drafts are measured end to end and can simply be disabled when the acceptance rate does not repay the verification cost. The stated bar for adoption is a controlled end-to-end A/B, not a microbenchmark.
Direct I/O and dual-SSD striping are described as drive dependent
Storage is treated as part of the engine rather than as an unavoidable cost, and the techniques named are batched expert unions, overlapped reads and compute, O_DIRECT, and weighted dual-SSD striping. The stated aim is to attack the streaming path instead of pretending storage latency is free, and the two caveats are specific. O_DIRECT is called out as drive dependent, which in practice means the win exists on some hardware and not on others and cannot be assumed from a data sheet. Weighted dual-SSD striping still needs broader end-to-end community A/Bs, so the project is asking for measurements rather than claiming a settled result. Heterogeneous execution sits alongside this, with CPU, CUDA, Metal, NUMA memory, and partial or full expert residency sharing one runtime, combined according to the machine, with the caveat that the profitable combination depends on compute, bandwidth, residency, and workload.
Nine model families, one C file each, one front end
Model coverage is the widest part of the project. Nine families run today, each implemented in one C file and all reached through the same three commands. They are GLM-5.2 and GLM-5.3 at 744B, GLM-5.3-Flash at 321B with vision, Inkling at 975B, Kimi K3 at 2.8T, DeepSeek V4 Flash at 284B, DeepSeek V4.1 Flash at 552B with vision, Qwen3.8-Flash-Next at 125B plus a 51B n-gram component, Qwen3.6 at 35B-A3B, and OLMoE at 7B. The front ends are a terminal chat, a service, and a browser dashboard, invoked as coli chat, coli serve, and coli web. The stated practical consequence is accessibility: running a 744B parameter model on hardware you already own, watching every expert fire, and being able to change the code that does it, rather than renting intelligence behind an API.
Zero engine dependencies does not mean zero dependencies
The pure C, zero dependency claim covers the engine, and the surrounding tooling is a separate matter. The repository ships a Python package named colibri-engine that requires Python 3.10 or newer and registers a single console script named coli. That package carries three optional dependency groups, and the one that matters for verification is called oracle:
[project.optional-dependencies]
convert = [
"numpy",
"huggingface_hub",
]
oracle = [
"torch>=2.0",
"transformers>=4.40",
"safetensors",
]
bench = [
"tokenizers",
"datasets",
]So a caller who wants the oracle or benchmark path is installing PyTorch and Transformers, which is a very different footprint from a dependency-free engine. The package is also classified as a beta for science and research audiences on Linux, macOS, and Windows, which is a fair signal about what kind of software this is.
The build hands every target to the c/ directory
The Makefile is a one-line dispatcher, and reading it tells you where the real work lives. Every target in the file forwards to a sub-make in the C source directory:
.PHONY: all glm deepseek-v4 portable test check cuda-test clean install uninstall
all glm deepseek-v4 portable test check cuda-test clean install uninstall:
$(MAKE) -C c $@The target names are informative on their own. Two are named after specific model families, glm and deepseek-v4, so the build knows about particular weights. One is named portable, which implies a build for machines that cannot assume CUDA or Metal. check and test are separate, and cuda-test is distinct from both, so the GPU path is verified on its own terms. The repository layout backs this up, with a c/ directory for the engine, a web/ directory for the dashboard, a desktop/ directory, a docker/ directory, a docs/ directory, a site/ directory for the published page, and a flake.nix with a lock file for reproducible environments.
Brio mode generates nothing at all and reports entropy instead
One feature in the dashboard is worth describing on its own because it inverts the usual interaction. Brio mode takes the same model and tells it to stop writing. You give it a document plus the only answers it is permitted to pick, and instead of generating text it reads the probability of each candidate, generates nothing at all, and reports an entropy value that indicates when it is not confident. The worked example in the documentation is a compliance style question answered at 99.9 percent with an entropy of 0.005, using 4 tokens read and 0 generated. That is a different contract from a chatbot, and it is a useful one when a wrong answer is more expensive than no answer. The Brain page sits beside it, drawing a measured expert atlas of GLM-5.2 as a cortex, with 13,260 characterised experts across ten regions, and it is explicit that position is measured routing affinity rather than a learned embedding.
Editorial conclusion
Colibri earns a place on the desk of a researcher who wants to watch a frontier mixture-of-experts model actually run, measure where a turn goes, and change the code that does it, or of a developer locked out of a frontier model by hardware. It does not serve an application that needs a latency guarantee, because the project states outright that there is no service level on speed and that its prefetch and routing policies can lose on some hosts, and it does not serve a minimal install, because the zero dependency claim covers the C engine while the Python tooling around it optionally pulls in PyTorch. Before adopting it, read the semantics policy and hold the project to it, check that your storage can sustain the streaming path, since the direct I/O and striping work is described as drive dependent, and pin a version, since the engine ships as 0.0.x style point releases that change behaviour between tags.
Frequently asked questions
how to install colibri ai
The build is driven from a Makefile whose targets forward to the C source directory, with an install target among them. The Python-facing package is colibri-engine, which requires Python 3.10 or newer and registers a console script named coli, so the chat, serve, and web front ends are then invoked as ./coli chat, ./coli serve, and ./coli web.
how to use colibri
Three front ends share the same engine and the same model families. coli chat opens a terminal session, coli serve exposes the engine as a service, and coli web opens the browser dashboard with a chat dock, Brio mode, the Brain expert atlas, and a profiling page.
how to setup colibri
Setup is a build rather than a wizard. The Makefile exposes targets named all, glm, deepseek-v4, portable, test, check, cuda-test, clean, install, and uninstall, and every one of them delegates to a sub-make in the c directory. A flake.nix and lock file are included for reproducible environments.
how to use colibri app
The application surface is the web dashboard started with ./coli web, redesigned in version 1.12.0. It is a workspace with a dock for the chat, Brio mode, the Brain page and its expert atlas, and a Profiling page showing where each turn spent its time by phase.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/justvugg-colibri)