Model or dataset
brontoguana/krasis avatar
brontoguana/krasis

Krasis: a hybrid LLM runtime for large MoE models on consumer NVIDIA GPUs

Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware

521 stars33 forksC++NOASSERTION

At a glance

What is it?
Krasis runs multi-hundred-billion-parameter mixture-of-experts models on VRAM-limited NVIDIA cards by keeping prefill and decode on the GPU while HCS moves hot and cold experts between VRAM and system RAM. The install path is short; the cache build and the licence are the parts to think about first.
Who is it for?
Adopt Krasis if you have an Ampere-or-newer NVIDIA GPU, enough system RAM for the quantized expert cache plus the HCS backing store, and a workload where a 300B-class MoE model at tens of tokens per second beats a smaller dense model that fits in VRAM. Skip it if you are on AMD, Apple Silicon or CPU-only hardware, if you need a permissive licence for a closed product, or if you cannot accept a first run that builds caches under ~/.krasis before producing useful speed.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 16 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The VRAM wall Krasis is built to climb

A 300B-parameter mixture-of-experts checkpoint does not fit in 32 GB of VRAM, and neither does its attention working set once you add a KV cache. The usual answers are to rent a bigger card, to serve a smaller dense model, or to accept CPU inference at a few tokens per second. Krasis takes a fourth route: keep the compute-heavy stages on the GPU and treat VRAM as a cache for experts rather than as the place the whole model must live.

The README frames the target narrowly. Krasis is an LLM runtime for running large MoE models on NVIDIA consumer GPUs, built around fast GPU prompt processing, GPU-executed decode, and HCS expert residency management so models much larger than VRAM can still run locally. The audience is therefore someone with one consumer or workstation NVIDIA card, a lot of system RAM, and a workload where a large sparse model is worth the setup cost. It is not a drop-in replacement for a small dense model on a laptop, and the README does not pretend otherwise: it lists CUDA and Ampere-or-newer as requirements and says input models should be BF16 safetensors from Hugging Face or another local safetensors source.

The published benchmark table shows the shape of the trade. On an RTX 5090 with 32 GB, Krasis reports 973.8 tok/s prefill and 10.04 tok/s decode for Qwen3.5-397B-A17B, and 1,852.2 tok/s prefill with 41.87 tok/s decode for Nemotron-3-Super-120B-A12B. Those are internal engine measurements from the project's own table, not independent numbers. The pattern matters more than the digits: a 397B model on a 32 GB card is slow but running, while a 30B-class MoE reaches 151.76 tok/s decode. Sparse activation is what makes the arithmetic possible at all.

How HCS, prefill and decode divide the work

The runtime is split into two languages for a reason. Python handles the launcher, setup and model loading, while the performance-sensitive path is Rust/CUDA: orchestration, CUDA kernels, cached quantized weights and measured VRAM budgeting. The README states plainly that the current runtime is no longer the early Python-hot-path prototype. That is a meaningful architectural claim, and the Cargo.toml backs it up with a cudarc dependency, a cuda feature enabled by default, and a cdylib target that maturin builds into the wheel.

HCS is the piece that makes the hybrid story work. The README describes it as managing hot and cold expert residency between VRAM and CPU RAM, with GPU-executed decode on top. Recent changes listed for this release line include measured startup calibration, prompt-conditioned reload, a dynamic recency tail, per-stage budgets, soft-tier reload caps, and safe eviction and reload paths. Read together, that is a caching policy tuned at startup rather than a fixed split: the runtime measures what fits, then decides which experts stay resident and which are fetched.

Around that sit the VRAM safety systems: short and long prefill and decode calibration, measured scratch budgets, pressure detection, idle pressure drain, and hard exit protection before CUDA enters an unsafe OOM state. That last item is the honest part of the design. A runtime that deliberately overcommits VRAM needs a way to stop before the driver does, and the README treats that as a feature rather than hiding it.

The quantized caches are the other half. Krasis builds cached INT4/INT8 expert formats and HQQ attention caches under ~/.krasis, and supports compact KV cache modes: k6v6, described as the quality-oriented launcher default, and k4v4 for tighter VRAM budgets. AWQ and Polar4 are deprecated for new production runs. This is a runtime with opinions about which quantization path you should be on, and the README says so.

Install and first run on Linux or WSL2

The README points at v1.0.16 as the release to install, even though the repository has moved on to 1.0.21 release candidates. For Linux or WSL2 it gives a single installer command that pulls the prerelease channel:

bash
curl -sSf https://raw.githubusercontent.com/brontoguana/krasis/main/install.sh | bash -s -- prerelease

On Windows the README offers a native installer, KrasisSetup-1.0.16-win64.exe, which installs for the current user, bundles its own Python runtime, and adds Krasis to the Start Menu. Either way, the entry points come from pyproject.toml: krasis for the launcher, krasis-chat for the chat client, and krasis-setup for setup. Python 3.10 or newer is required.

After install, the launcher is where a first real run happens. The README describes an interactive launcher with a curated Hugging Face downloader flow, so a first session looks like this:

bash
krasis

Expect the first run to be slow. Krasis builds optimized local caches before it can serve the model at speed, and later runs reuse them. The caches land under ~/.krasis, and the README warns that disk usage must cover the source model plus the Krasis cache artifacts. If you are testing a 300B checkpoint, plan for both copies before you start the download.

Once a model is loaded, the runtime exposes an OpenAI-compatible API, so existing clients can point at the local server rather than at a hosted endpoint. Two maintenance commands are documented in the release notes: krasis update and krasis prerelease. For monitoring during a run, the README points at a separate project, ktop.

Where Krasis is the wrong tool

The hardware requirement is the first hard boundary. Krasis targets NVIDIA GPUs with CUDA, Ampere and newer. There is no documented path for AMD, Apple Silicon, or CPU-only machines, and the README does not describe one. If your fleet is not NVIDIA, this project is not a candidate regardless of how the benchmarks read.

System RAM is the second constraint, and it is easy to underestimate. The README says RAM should be sized for the selected quantized cache and the HCS backing store, and that larger models need substantial RAM even when GPU VRAM is limited. HCS exists precisely because experts live in CPU RAM and move in and out of VRAM; the cheaper your RAM subsystem, the more that movement costs. A machine with a fast GPU and 32 GB of system RAM is a worse fit than the VRAM number alone suggests.

The third issue is quality drift under quantization. The README itself warns readers to see the per-family quality and tool-use limitations rather than assuming every quantized runtime is equally faithful, and it ships a separate STATS-QUALITY.md for that reason. A 397B model running at k4v4 is not the same model as the BF16 checkpoint, and the project does not claim it is. If your task depends on exact tool-call syntax or long-context recall, validate on your own prompts before you commit.

Finally, the release channel is worth reading carefully. The most recent releases are 1.0.21 release candidates, while the README still directs new users to v1.0.16. That gap is normal for a project moving fast, but it means the documented install path and the newest code are not the same artifact. Anyone who needs a stable pin should check the releases page rather than follow the top of the README.

Krasis against llama.cpp and vLLM

The obvious comparison is llama.cpp, and the repository's own topics list llama-cpp-alternative as a declared position. The difference is in where the effort goes. llama.cpp is built to run across a wide range of backends and hardware, with CPU inference as a first-class path and GPU offload as an accelerator. Krasis inverts that: the README describes full GPU prefill and GPU-executed decode as the design, with CPU RAM serving as storage for cold experts rather than as a compute device. If your machine has no usable CUDA GPU, llama.cpp covers ground Krasis does not attempt.

Against vLLM the split is different again. vLLM is a serving system built around high-throughput batched inference on hardware sized to hold the model. Krasis is aimed at the case where the model does not fit, and it accepts single-stream decode rates in the tens of tokens per second to get there. The published table shows HTTP round-trip figures that are lower than the internal decode numbers, which is what you would expect when a local client and server sit in the same loop.

The honest framing is that Krasis trades throughput headroom for reach. It lets a 32 GB card address a 397B MoE checkpoint. It does not promise that doing so is fast, and the benchmark table shows decode falling to 10.04 tok/s in exactly that case. Pick it when the model matters more than the latency.

Licence, upgrade cost and what the repository does not say

The licence situation needs care. The repository's licence metadata is reported as NOASSERTION, while pyproject.toml and Cargo.toml both declare SSPL-1.0, and pyproject.toml also carries the classifier License :: Other/Proprietary License. Those signals point in the same general direction but are not identical, and the classifiers list Development Status :: 5 - Production/Stable alongside a project whose newest releases are release candidates. Treat the SSPL-1.0 declaration in the build files as the operative statement and read the LICENSE file in the repository root before you ship anything. This is not legal advice, and the SSPL is not a permissive licence in the way MIT or Apache-2.0 are.

Upgrade cost is dominated by caches and calibration rather than by the wheel itself. The release notes describe cache build and rebuild support for HQQ attention, and the README says first run is slower because Krasis builds optimized local caches. A version bump that changes the cache format or the calibration logic can therefore invalidate work under ~/.krasis and force a rebuild, which on a multi-hundred-billion-parameter checkpoint means both time and disk. The repository layout is consistent with that reading: there is a CHANGELOG.md, a TESTING.md, a benchmarks directory, and separate STATS-BENCHMARKS.md and STATS-QUALITY.md files, so the project expects you to check before upgrading rather than assume compatibility.

What the repository does not document is equally worth noting. The README gives no rollback procedure for a bad upgrade, no uninstall instructions, and no guidance on running multiple models concurrently. The maintenance commands krasis update and krasis prerelease are named in the release notes but their behaviour on a failed download is not described. The last push to the repository was on 2026-09-03, and the most recent release tag carries the same date, so the codebase is current; the gaps above are documentation gaps, not signs of abandonment.

Editorial conclusion

Adopt Krasis if you have an Ampere-or-newer NVIDIA GPU, enough system RAM for the quantized expert cache plus the HCS backing store, and a workload where a 300B-class MoE model at tens of tokens per second beats a smaller dense model that fits in VRAM. Skip it if you are on AMD, Apple Silicon or CPU-only hardware, if you need a permissive licence for a closed product, or if you cannot accept a first run that builds caches under ~/.krasis before producing useful speed. Verify three things before committing: that your exact model family appears in the validated list in the README, that the INT4/HQQ4/k4v4 path holds quality on your task by reading STATS-QUALITY.md, and that disk and RAM budgets cover the source safetensors plus the cache artifacts.

Frequently asked questions

What hardware does Krasis need?

The README states that Krasis currently targets NVIDIA GPUs with CUDA, including Ampere and newer architectures, and that the production HQQ attention and compact KV cache modes do not require FP8 support. System RAM should be sized for the selected quantized cache and the HCS backing store, so larger models need substantial RAM even when VRAM is limited.

How do I install Krasis on Linux or WSL2?

The README gives a single installer command that fetches the prerelease channel: curl -sSf https://raw.githubusercontent.com/brontoguana/krasis/main/install.sh | bash -s -- prerelease. Windows users are pointed at the KrasisSetup-1.0.16-win64.exe installer, which installs for the current user, bundles its own Python runtime, and adds Krasis to the Start Menu.

Why is the first Krasis run slow?

The README states that first run is slower because Krasis builds optimized local caches, and that later runs reuse those caches. Those cached INT4/INT8 expert formats and HQQ attention caches are written under ~/.krasis, so disk usage must cover the source model plus the cache artifacts.

What are the k6v6 and k4v4 modes in Krasis?

They are compact KV cache modes. The release notes describe k6v6 as the quality-oriented launcher default and k4v4 as the option for tighter VRAM budgets, alongside BF16 KV depending on the memory and quality target you want.

What licence is Krasis released under?

Both pyproject.toml and Cargo.toml declare SSPL-1.0, while the repository's licence metadata is reported as NOASSERTION and pyproject.toml carries the classifier License :: Other/Proprietary License. Read the LICENSE file in the repository root before relying on any of those signals.

Official sources

  1. brontoguana/krasis on GitHub
  2. Issues
  3. README
  4. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/brontoguana-krasis.svg)](https://hysenlabs.com/projects/brontoguana-krasis)