ik_llama.cpp: a llama.cpp fork for CPU inference, extra quants and CUDA
llama.cpp fork with additional SOTA quants and improved performance
At a glance
- What is it?
- ik_llama.cpp is a C++ fork of llama.cpp with additional quantization types and different CPU and CUDA kernels. It is aimed at people running quantized models on their own hardware, and it asks more of them than upstream does.
- Who is it for?
- Adopt ik_llama.cpp if you run quantized models on an AVX2-or-better CPU or a Turing-or-newer Nvidia GPU and want the extra quantization types and kernels this fork carries. Do not adopt it if your only hardware is ROCm, Vulkan, Metal, an AVX-only CPU or an older Nvidia card: the README says those issues will not be resolved unless you contribute the backend work yourself.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What ik_llama.cpp actually is, and who it is for
This repository began as a fork of llama.cpp in June 2024 and, per the README, was last synced with upstream in August 2024. That single sentence explains most of the project's character. It is not a thin wrapper or a plugin. It is a diverged C++ codebase that tracks model architectures and quantization research on its own schedule, and it carries features that later appeared in llama.cpp, with the README naming MLA, quant repacking, fused delta-net (called GDN upstream), tensor parallel, MTP and DFlash as examples.
The audience is narrow and specific. You are the target user if you run quantized GGUF models locally and care more about what runs on your particular CPU or GPU than about breadth of hardware support. The README is blunt that the only fully functional and performant backends are CPU with AVX2 or better (or ARM_NEON or better) and CUDA on Turing or newer. Everything else is explicitly out of scope for the current contributors.
That is a real trade-off, not a marketing line. Mainline llama.cpp tries to run everywhere. This fork tries to run well on a subset and says so. If your machine falls outside that subset, the fork has nothing to offer you.
How the quant repacking and backend split work
The mechanism that matters most is tensor repacking, exposed through the -rtr option. The README states that -rtr repacks all tensors left in RAM into row-interleaved format while the model loads. Row-interleaved layouts are what let the quantized matrix multiplication kernels run efficiently, and the fork uses them to get its performance on CPU.
The catch is that repacking is not free of consequences for hybrid setups. Because not every quantization type has a CUDA row-interleaved implementation, a repacked tensor that stays in RAM will have its matrix multiplications done on the CPU even when the GPU would have been faster. The README calls out k-quants specifically, listing K2_K, Q3_K, Q4_K, Q5_K and Q6_K as lacking a CUDA row-interleaved implementation. For MoE models with all or some experts on the CPU, the README's warning is direct: do not use -rtr unless you know what you are doing, because the typical outcome is lower prompt processing speed.
A second mechanism worth knowing is the AVX-512 path. The README points to docs/build.md and its section on CPU build flags for AVX-512, which activates the IQK quantized GEMM kernels, referred to as the HAVE_FANCY_SIMD path. Without those flags, the README says a vanilla Release build silently falls back to the AVX2 path on AVX-512 hardware. Silent fallback is the kind of default that costs you performance without an error message, so it is worth checking the build documentation rather than assuming -DGGML_NATIVE=ON is enough.
There is also a known interaction between split mode graph and partial GPU offload. The README notes that some users reported issues when combining graph parallel with --cpu-moe, --n-cpu-moe or tensor overrides, and suggests adding -cuda graphs=0 if responses come out incoherent.
Building ik_llama.cpp on Linux and starting the server
The README gives a clone, a Debian/Ubuntu package list, a CPU build, a CUDA build and a run command. Start by cloning the repository and installing the build dependencies.
git clone https://github.com/ikawrakow/ik_llama.cpp
cd ik_llama.cppOn Debian or Ubuntu, the README lists these packages. Other distributions need the equivalent packages adapted by the reader.
apt-get update && apt-get install build-essential git libcurl4-openssl-dev curl libgomp1 cmakeThe CPU build uses CMake with native optimizations enabled, then a Release build across all cores.
cmake -B build -DGGML_NATIVE=ON
cmake --build build --config Release -j$(nproc)For CUDA, install the Nvidia drivers and the CUDA Toolkit first, then add the CUDA flag to the configure step.
cmake -B build -DGGML_NATIVE=ON -DGGML_CUDA=ON
cmake --build build --config Release -j$(nproc)After downloading a .gguf file to a directory of your choice, the README's example starts the server with a 4096-token context. The README uses /my_local_files/gguf as the example directory and Qwen_Qwen3-0.6B-IQ4_NL.gguf as the example model.
./build/bin/llama-server --model /my_local_files/gguf/Qwen_Qwen3-0.6B-IQ4_NL.gguf --ctx-size 4096To offload layers to the GPU, the README adds -ngl 999 to the same command. The server then answers on port 8080 at 127.0.0.1, where the README says you can open a browser and chat or point a program at the API endpoints.
./build/bin/llama-server --model /my_local_files/gguf/Qwen_Qwen3-0.6B-IQ4_NL.gguf --ctx-size 4096 -ngl 999If you prefer containers, the README lists images on ghcr.io under ghcr.io/ikawrakow/ik-llama-cpp with cpu-swap, cpu-server, cpu-full, cu12-swap, cu12-server and cu12-full tags, and points to docker/README.md for customization.
docker pull ghcr.io/ikawrakow/ik-llama-cpp:cpu-serverThe README also links a step-by-step guide for a successful Windows build at docs/build.md. If you are searching for an ik_llama cpp Windows build, that document is where the project sends you.
Backend support is the fork's sharpest limitation
The README does not hedge here. It states that issues related to ROCm, Vulkan, Metal, old Nvidia GPUs and AVX CPUs will not get resolved unless the person filing them rolls up their sleeves and helps bring the backend up to speed. The stated reason is contributor bandwidth, not technical impossibility. That is an honest answer, and it is also a hard boundary: if your hardware is on that list, this fork is the wrong tool and mainline llama.cpp is the right one.
Model compatibility has a similar edge. The README warns against quantized models from Unsloth that have _XL in their name, saying they are likely not to work with ik_llama.cpp. A later clarification narrows the problem to _XL models that contain f16 tensors, which the README calls never a good idea in the first place, and states that all others are fine. The practical reading is that a specific subset of published quants will fail, and the failure is attributed to how those files were produced rather than to the fork's loader.
The -rtr warning is a limitation in the same category. It is an option that helps CPU-bound inference and hurts hybrid MoE inference, and the README expects you to know which situation you are in before you pass it. There is no automatic detection described.
The AVX-512 fallback is the quietest of the three. A build that succeeds and runs is not necessarily a build that uses the fast path, and the README's own wording is that it silently falls back. Silent is the operative word.
How it differs from mainline llama.cpp
The comparison is not about features in the abstract. It is about where each project spends its maintenance effort. Mainline llama.cpp is the upstream that this fork diverged from in June 2024 and last synced with in August 2024. Upstream carries the broad backend matrix: CUDA, ROCm, Vulkan, Metal, SYCL and CPU paths, maintained by a much larger contributor group. The fork carries a narrower set, and the README says plainly that the current regular contributors do not have the bandwidth to work on all the backends available in llama.cpp.
What the fork offers in exchange is quantization research that lands here first. The README's list of features that appeared in this repository before llama.cpp includes MLA, quant repacking, fused delta-net, tensor parallel, MTP and DFlash. The model support list in the README is long and reads as a changelog of merged pull requests, from LLaMA-4 through DeepSeek-V3, Kimi-2, GLM-4.5 and its later variants, Qwen3, Qwen3-VL, Qwen3-Next and Qwen3.5-MoE, grok-2, Seed-OSS and others. If you follow new architectures closely and want them in a local runtime early, that list is the argument for this fork.
The honest framing is that you are choosing between breadth of hardware support and proximity to new quantization and kernel work. Neither choice is wrong. They just fail differently. Upstream fails by not having a new quant yet. This fork fails by not running on your GPU.
Maintenance, licence and the cost of staying current
The repository is not archived, and the last push was on 2026-09-23. The most recent release listed is t0002, dated 2025-07-22. That gap between a tagged release and continuous commits is worth noting: the project's activity is in the branch, not in release artifacts, so anyone who pins to a release tag is pinning to something from July 2025.
The licence is MIT, which is the same licence family as the upstream project this was forked from. MIT permits commercial and private use, modification and redistribution provided the copyright notice and permission notice are preserved. That is a summary of the licence identifier, not legal advice; if you are redistributing binaries or embedding the runtime in a product, read the LICENSE file in the repository and get your own counsel on notice requirements.
The upgrade cost is where this fork differs most from a normal dependency. Because it was last synced with upstream in August 2024, you cannot treat it as a drop-in that inherits upstream fixes. Bug fixes and backend improvements made in llama.cpp after that date are not automatically present here, and the README's stance on non-CPU, non-CUDA backends means you should not expect them to arrive. Upgrading within the fork means pulling main, rebuilding, and re-checking your flags, particularly the AVX-512 flags and whether -rtr still suits your workload. If you need a version you can pin for a year and forget, this is not that project.
Editorial conclusion
Adopt ik_llama.cpp if you run quantized models on an AVX2-or-better CPU or a Turing-or-newer Nvidia GPU and want the extra quantization types and kernels this fork carries. Do not adopt it if your only hardware is ROCm, Vulkan, Metal, an AVX-only CPU or an older Nvidia card: the README says those issues will not be resolved unless you contribute the backend work yourself. Before committing, verify three things: that your target model loads, that the quantization type you want has the backend implementation you plan to run it on, and whether you need -rtr at all, since the README warns it forces row-interleaved tensors onto the CPU and can slow prompt processing for MoE models with experts left in RAM.
Frequently asked questions
What is ik_llama.cpp?
It is a fork of llama.cpp started in June 2024 that offers additional SOTA quantization types and, in many cases, better performance, according to the README. It was last synced with upstream in August 2024, so it has diverged since then.
How do I install ik_llama.cpp?
Clone the repository, install the build packages listed for Debian or Ubuntu, then configure and build with CMake using -DGGML_NATIVE=ON, adding -DGGML_CUDA=ON for a CUDA build. The README also lists prebuilt ghcr.io images under ghcr.io/ikawrakow/ik-llama-cpp for Docker or Podman.
How do I use ik_llama.cpp?
Download a .gguf model file, then run ./build/bin/llama-server with --model pointing at it and --ctx-size 4096, adding -ngl 999 to offload layers to the GPU. The README says the server is then reachable at http://127.0.0.1:8080 for chatting or for API calls.
How does ik_llama.cpp compare with vLLM?
The README does not compare the two. What it does state is that ik_llama.cpp runs GGUF models locally through llama-server, with CPU (AVX2 or better) and CUDA (Turing or newer) as the only fully functional and performant backends, and it asks that issues about other backends not be filed.
Does ik_llama.cpp run locally?
Yes. The README's quickstart downloads a .gguf file to a local directory, starts ./build/bin/llama-server with --model and --ctx-size 4096, and then points the browser at http://127.0.0.1:8080. No hosted service is involved in that flow.
What is llama-cpp used for?
In this repository, it is used to run quantized GGUF language models locally for chat and for programmatic access through the llama-server API endpoints on port 8080. The README's examples cover both CPU-only and GPU-offloaded inference.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ikawrakow-ik-llama-cpp)
Community notes