colibri
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine, pure C, zero deps, experts streamed from disk. Tiny engine, immense model.
What it solves
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine, pure C, zero deps, experts streamed from disk. Tiny engine, immense model.
The project context on this page is free to read. The original GitHub repository remains the source of truth; sign in only when you want to save or join the discussion.
Community notes