Model or dataset
gavamedia/deltafin avatar
gavamedia/deltafin

Deltafin's headline is fidelity, and its benchmark is 3.4 seconds per token

Run full Kimi K3 on a single device. And an OpenAI-compatible API server for local chat and coding agents.

828 stars99 forksRustNOASSERTION

At a glance

What is it?
A Rust workspace that runs the full 2.8 trillion parameter Kimi K3 on one machine, with draft models allowed to guess and K3 checking every token. The page is emphatic that nothing is pruned, and the number it leads with is 0.29 tokens per second on an M1 Max.
Who is it for?
Decide what you are buying before you build. If you want K3's weights exactly as Moonshot released them, this is a serious attempt and the provenance argument is coherent, and you accept three and a half seconds per token on a laptop.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 58 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The headline number is 3.4 seconds per token

The page is built around one promise: the full, never-pruned model, all 16 routed experts, with K3 as the sole authority for every token, on consumer hardware. Then the benchmark, on an M1 Max laptop, reads 0.2901 token per second, which is 3.447 seconds for a single token. That is the fastest figure in the document and it is the one the page puts first, so the mission statement and the measurement are the same sentence. For scale, the same page says K3 targets infrastructure on the scale of 16 nodes and roughly 4.8 TB of aggregate VRAM, and contrasts a 15,000 dollar home setup with two million dollars of hosted hardware. The project is not hiding the ratio. It is arguing that the ratio is the point, since every gain is described as research rather than as a product improvement.

The gains fall from 830 percent to 1.9 in six days

The historical benchmark list is the most informative table on the page, and nobody framed it as one. On 27 July 2026 the figure is 0.0141 tokens per second. The next day 0.1311, described as 829.8% higher throughput. On 30 July 0.2660, up 102.9%. On 2 August 0.2847, up 7.0%. And then the newest bullet, 0.2901, described as 1.9% higher throughput than the last update. So the project improved by more than eight times in a day, doubled again, and is now gaining under two percent a change. That is the shape of a curve flattening out, presented as a list of wins. One more detail: every historical entry carries a date and the newest one does not, so the most important number in the document is also the only one you cannot place in time.

Small models guess, and K3 checks every guess

The fidelity argument is specific about how speed is obtained. A resident model is allowed to guess ahead of K3, and nothing reaches the caller without K3's sign-off, so the output is K3's own. The default install pulls a third-party draft model from the Hugging Face hub, Inferact's Kimi-K3-DSpark, which takes 6.635 GiB on disk and about 4.49 GiB when admitted at runtime, and the project notes that it avoids materialising DSpark's redundant copy of K3's embedding table. A second draft model, Qwen, is an optional add-on for raw continuation, 4.337 GiB more, installed by its own command. The one measured result given for it is a 17-token completion that ran 2.7 times faster with the same output identifiers. Identical identifiers is the right check, and it is the check that lets the quality claim survive the speedup.

The build replaces curl-sys with a crate in the repository

The workspace is five crates, not one: the main binary, a bootstrap crate, a curl FFI crate, a native build crate and an xtask runner, with the binary as the default member. The interesting part is a patch applied to the whole dependency graph:

toml
[patch.crates-io]
# Deltafin keeps curl-rust's reviewed FFI surface but replaces curl-sys's
# generic helper-driven discovery/build script with a bounded, Rust-only
# system-libcurl selector for the supported native targets.
curl-sys = { path = "native/deltafin-curl-sys-direct" }

So the system library is not found by upstream's build script. The reason given is portability and boundedness, and the trade is that you now depend on a hand-written FFI selection layer rather than a widely reviewed one. The same care appears in the build command, which uses `--locked`, and in the fact that a lock file is committed at the top level. For a project that downloads a terabyte of weights, having a reproducible dependency graph is not a small detail, and this is the part of the design that gets it.

The release profile trades unwinding and parallelism for latency

Four release settings, each with a stated reason:

toml
[profile.release]
codegen-units = 1
lto = "fat"
panic = "abort"
strip = "symbols"

One codegen unit so the optimiser can inline across scheduler and configuration boundaries, fat link time optimisation for the same reason, and symbols stripped to keep the binary small. The fourth is the one with consequences beyond size. Aborting rather than unwinding is right at a C, ATen, Metal or CUDA boundary, where an unwind across the frame is unsafe or slow. It also means a panic anywhere in the scheduler takes the process down, and that any component expecting to catch a failure in Rust, a worker supervisor, a retry wrapper, a plugin boundary, cannot. The comment explains the choice and the cost is implicit. The toolchain floor is Rust 1.85 on edition 2024, so this is not a codebase that builds on an older image without thought.

Upgrade remembers the build, so a CUDA switch is manual

The upgrade path is unusually careful and worth reading twice. It fetches what it needs and rebuilds the binary, and it leaves models, converted weights and caches alone; it never re-runs setup and never re-downloads the model. It requires a clean, non-diverged branch, and it records how the binary was built, so a build made for an NVIDIA target stays that way rather than quietly falling back to a CPU build on the next upgrade. Anything unexpected stops the upgrade. It also ignores build environment variables on purpose, which means switching from CPU to CUDA, or relocating a LibTorch tree, is something you do by hand once: run the locked release build yourself with the new variables, and that becomes the recorded setup. The design treats a silent fallback to a slower build as the worst outcome, which is a defensible priority for something whose headline number is already measured in seconds per token.

The Python install is a migration note, and there is no release

The project used to be Python, and the page still carries a note for people arriving from the old version: check the working tree, continue only when the status is empty, then pull with `--ff-only`, rebuild and upgrade. Preserving or committing uncommitted work is left to the reader, with the reasoning that an upgrade procedure should not guess. That note is now the only trace of the earlier implementation; the repository itself is Rust end to end, with a manifest, a lock file, a native tree, docs, research and tools. There are no GitHub releases, and the install path has no prebuilt binary at all, which means the documented first step is a locked source build on your own machine before a 1.7 TB model download, or 215 GB if you take the streaming route. One smaller thing: the repository's own licence metadata comes back unclassified while the workspace manifest declares MIT and a LICENSE file sits at the root, and the page does not say whether the file's text matches the manifest.

Editorial conclusion

Decide what you are buying before you build. If you want K3's weights exactly as Moonshot released them, this is a serious attempt and the provenance argument is coherent, and you accept three and a half seconds per token on a laptop. If you want something usable in a conversation or inside an agent loop, the same page names the projects it regards as competitors and the reason it rejects them, a roughly 3-bit re-encoding of the expert bank, which is the direction to look instead. Two practical notes before you start: there is no release and no prebuilt binary, so the first step is a locked source build against a 1.85 toolchain, and the licence is declared in the workspace manifest and unclassified in the repository metadata, so read the LICENSE file rather than trusting either label.

Frequently asked questions

How fast does Deltafin run Kimi K3 on a laptop?

The benchmark given is on an M1 Max laptop at 0.2901 token per second, which is 3.447 seconds per token, described as 1.9% higher throughput than the previous update. The historical list runs from 0.0141 tokens per second on 27 July 2026 up to that figure, with intermediate steps of 0.1311, 0.2660 and 0.2847.

How much disk does Deltafin need for Kimi K3?

The full model is 1.7 TB, downloaded by a setup command marked optional but fastest. A streaming setup starts at 215 GB by installing the resident model and fetching exact experts on demand, which runs much more slowly until a cache has built up. The page also links requirements documentation for anything the install does not cover.

Does Deltafin prune or quantise Kimi K3?

The page states nothing is pruned and all 16 routed experts run, with K3 as the sole authority for every token, and criticises other projects for re-encoding the expert bank down to about 3 bits. Speed comes from draft models guessing ahead while K3 checks each guess, and the one measured draft result is a 17-token completion that ran 2.7 times faster with identical output identifiers.

How do I upgrade Deltafin without losing my models?

The upgrade command rebuilds the binary and leaves models, converted weights and caches alone, never re-running setup and never re-downloading the model. It requires a clean, non-diverged branch, records how the binary was built so a CUDA build does not fall back to CPU, ignores build environment variables, and stops on anything unexpected.

How do I install Deltafin?

Clone the repository, build with cargo build --locked --release, then either run the full setup, which downloads the 1.7 TB model, or the stream setup, which starts at 215 GB. A separate command installs the optional Qwen draft model. There are no GitHub releases, so the build from source is the install path.

Official sources

  1. gavamedia/deltafin on GitHub
  2. Issues
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/gavamedia-deltafin.svg)](https://hysenlabs.com/projects/gavamedia-deltafin)