kimi-k3-in-c: running a 2.78T-parameter model on one CPU in 8.24 GB of RAM
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
At a glance
- What is it?
- FareedKhan-dev's kimi-k3-in-c is a 176 KB C99 inference engine that streams a 1.56 TB checkpoint off disk, keeps the dense trunk resident, and multiplies routed experts straight out of 4-bit. It is a research artifact, not a chat server.
- Who is it for?
- Adopt kimi-k3-in-c if you want to read or modify a from-scratch CPU inference engine, or if you need to demonstrate that a 2.78T-parameter checkpoint can produce byte-identical output at 8 GB and at 128 GB. Do not adopt it as a serving layer for an application: the README's own laptop row is 26.5 s per token, and the engine has no chat template.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 8 days ago.
- What is it written in?
- Mainly C, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem kimi-k3-in-c actually solves
A 2.78-trillion-parameter checkpoint is 1.56 TB on disk. The usual answer is a cluster. The README's framing is narrower: the model should run on whatever machine you already own, and more memory should buy speed rather than change the answer. The engine is aimed at people who want to see how far that idea goes on commodity hardware, and at engineers who want a readable implementation of the memory layout decisions involved. It is not aimed at anyone who needs throughput. The README's own table lists 26.5 s per token on an 8 GB laptop and 5.6 s per token on a 128 GB+ workstation. Those numbers describe a demonstration, not a service.
Where the bytes live: trunk resident, experts streamed
The architecture described in the README splits the model by residency. The dense trunk stays in memory to a depth you choose. The 1.45 TB of routed experts are never resident and are multiplied directly out of their packed 4-bit form. That single decision is what makes an 8 GB budget possible at all: the resident working set is small, the model sits on NVMe underneath, and pipes move data between them. The README states the consequence plainly: the same model runs in 8 GB and in 224 GB and produces byte-identical output at every budget between. The engine is 176 KB of portable C99 with no BLAS and no framework, and the repository carries tests that need no model weights, which is the practical reason the build is verifiable before you download 1.56 TB.
Building the engine and running your first prompt
The Makefile is the entry point. The README says the quick start is clone, build and verify in about a minute, with no model. The documented targets are make for the engine binary, make test for every test that needs no weights, make bench for kernel microbenchmarks, make portable to build without -march/-mcpu=native, and make debug, make asan, make ubsan for instrumented builds. On macOS/arm64 the Makefile notes that plain make works but needs Homebrew's libomp for OpenMP, because Apple Clang ships no OpenMP runtime; the platform block detects and wires it up. Windows builds under MSYS2's MinGW64 environment, and the Makefile is explicit that you open the "MSYS2 MinGW x64" shell specifically, not the plain MSYS2 shell, so cc and make resolve to the native-Windows-target toolchain. The Python tooling in pyproject.toml is separate from the engine: numpy and torch are pinned for fixture generation and reference-vs-C conformance, and tiktoken lives in the dev group as the reference tokenizer, never as a runtime dependency of the C engine.
make
make testThe first command produces bin/k3. The second runs the weight-free test suite, which the Makefile calls the gate that must stay green. If either fails, the checkpoint is irrelevant.
The README's first runnable example passes a model directory, a trunk directory, a preset, a tokenizer path, a prompt and a generation length. The preset selects the memory budget; the README's laptop preset is the 8 GB row.
./bin/k3 ~/k3model --trunk ~/k3trunk --preset laptop \
--tok ~/k3model --prompt "The capital of France is" --gen 8 --incrementalThe documented output for that invocation is a continuation beginning " Paris.", then a summary line reading 8 tokens in 261.5 s at 32.69 s/token average, and a peak RSS of 8.24 GB. The README notes that this is a base model, so what follows the prompt is a continuation rather than a reply, and there is no chat template. The second documented example uses the server preset with a Python prompt and reports 28 tokens in 299.3 s at 10.69 s/token and a peak RSS of 127.92 GB. The README attributes the higher clock in the first capture to a slower drive, since the two demos are the original captures.
What the engine does not do
The most important limitation is stated in the README itself: this is a base model, so there is no chat template. If your application needs a conversational assistant, this engine gives you a continuation and nothing else. The second is speed. At 26.5 s per token on an 8 GB machine, a 200-token answer is over an hour of wall clock, and the README's own numbers show that adding RAM only moves you to 5.6 s per token at 128 GB+. The third is that the first three memory rows still read the model from disk on every step, so a slower drive is slower there; the 128 GB+ row keeps everything in memory and stops waiting on the disk. A machine with a large RAM budget but a slow disk lands in the worst of both. The README also does not document rollback, checkpoint conversion, or a way to resume a partially generated sequence, and the repository layout does not show a server or an HTTP interface.
How this differs from llama.cpp and ggml-based runners
llama.cpp and the ggml family are the obvious comparison, and the difference is in what each one optimizes. A ggml-based runner typically converts a checkpoint into its own quantized format and then serves from memory, with a broad model-format surface and a server mode. kimi-k3-in-c does the opposite: it reads the checkpoint's own headers, keeps the dense trunk in memory to a chosen depth, and streams the routed experts from disk in their packed 4-bit form on every step. That is why the same binary spans 8 GB and 224 GB with byte-identical output, and it is also why the laptop row is disk-bound. If you want a drop-in runner for many models with an OpenAI-compatible endpoint, this is the wrong tool. If you want to read a small codebase that shows the memory accounting for a mixture-of-experts model at this scale, the trade is the other way.
Licence, maintenance and the cost of keeping up
The project is Apache-2.0, with a NOTICE file in the repository, which is the usual arrangement for a permissively licensed project that wants attribution preserved. Apache-2.0 also carries an explicit patent grant, which matters more than usual here because the engine implements specific attention and quantization schemes. Nothing in the README suggests the licence restricts commercial use; as always, read LICENSE and NOTICE yourself rather than treating this as legal advice. On maintenance: the repository is not archived and the last push was on 2026-08-26, which is recent. Two releases exist, v0.1.0 on 2026-08-02 and v1.0.0 on 2026-08-07. The upgrade cost is dominated by the checkpoint, not the code. The README attributes an 8x reduction in per-token math, a 3.9x faster follow-up question in a chat, and roughly half the cost for long prompts to v1.0.0, so the version you build against changes the clock. The Python side pins numpy and torch exactly rather than as ranges, which the pyproject.toml comment explains is deliberate so the next environment resolves to the same versions.
Editorial conclusion
Adopt kimi-k3-in-c if you want to read or modify a from-scratch CPU inference engine, or if you need to demonstrate that a 2.78T-parameter checkpoint can produce byte-identical output at 8 GB and at 128 GB. Do not adopt it as a serving layer for an application: the README's own laptop row is 26.5 s per token, and the engine has no chat template. Before committing, verify three things: that your checkpoint and trunk directories match the layout the binary expects, that your drive is fast enough for the preset you pick, and that the preset's RAM budget is actually available, since the measured peak RSS for the server preset is 127.92 GB.
Frequently asked questions
What is kimi-k3-in-c?
It is a portable C99 inference engine for a 2.78-trillion-parameter Kimi K3 model, with no BLAS, no framework and no GPU. The README describes it as 176 KB of engine code that runs the model on a single CPU, with a measured peak RSS of 8.24 GB in its laptop example.
What GPU is required to run Kimi K3 with this engine?
None. The README's headline is zero GPUs, and the platform badge lists Linux x86-64 as the reference, with macOS/arm64 and MSYS2 MinGW64 builds also described in the Makefile. The constraint is RAM and disk speed, not a graphics card.
Can I run Kimi K3 locally with kimi-k3-in-c?
Yes, that is the point of the project. The README shows a laptop preset running in 8.24 GB of peak RSS, with the model streaming off disk on every step, and states that the output is byte-identical from the smallest machine to the largest.
What are the uses of Kimi K3?
The README does not describe downstream applications. It shows the model completing prompts such as "The capital of France is" and "def fibonacci(n):", and notes that it is a base model with no chat template, so output is a continuation rather than a reply.
What is Kimi K3 in Ollama?
The README does not describe an Ollama integration for this project. The repository is a standalone C99 binary built with make, and no Ollama packaging or model registry entry appears in the documented setup.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/fareedkhan-dev-kimi-k3-in-c)