# KTransformers: CPU-GPU Heterogeneous Inference for Large Language Models

> KTransformers is a research framework for running large language models on hardware that would normally lack sufficient GPU memory. It offloads model experts to CPU and DRAM while keeping the compute-intensive parts on GPU, enabling inference of 671B-parameter MoE models on consumer hardware.

**kvcache-ai/ktransformers** — A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations

- Repository: https://github.com/kvcache-ai/ktransformers
- Website: https://kvcache-ai.github.io/ktransformers/
- Stars: 19,541 · Forks: 1,580
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/kvcache-ai-ktransformers

## What KTransformers Solves: Running 671B-Parameter Models on Consumer Hardware

Large Mixture-of-Experts (MoE) models like DeepSeek-V3 and DeepSeek-R1 have hundreds of billions of parameters but activate only a fraction of them per token. This sparsity makes them candidates for CPU offloading: the inactive experts sit in DRAM, and only the active ones are loaded to GPU VRAM at inference time.

KTransformers exploits this pattern. The README describes the project as focused on efficient inference and fine-tuning of large language models through CPU-GPU heterogeneous computing. The February 2025 update documented support for DeepSeek-R1 and V3 on a single 24GB VRAM GPU with 382GB DRAM, with speeds up to 28x faster than a naive CPU-only approach. A 24GB VRAM card alone cannot hold a 671B model; KTransformers makes the inference possible by treating CPU DRAM as a slower secondary memory tier.

The project is classified as a research framework. The pyproject.toml does not assign a development status classifier. Release v0.7.1 was published on 2026-09-15 and the last push was on 2026-09-23.

## Installing KTransformers

KTransformers requires Python 3.11 or later and runs on Linux. The pyproject.toml classifiers list only POSIX Linux as the operating system, though the README mentions a Windows native support entry from August 2024 in its update log.

The package is named ktransformers on PyPI. The setup.py shows the main package depends on kt-kernel at the same version, with optional extras for SFT fine-tuning and SGLang integration:

```bash
pip install ktransformers
```

For fine-tuning:

```bash
pip install ktransformers[sft]
```

For SGLang integration:

```bash
pip install ktransformers[sglang]
```

The SFT extra installs transformers-kt and accelerate-kt, which are KTransformers-specific forks of the Hugging Face transformers and accelerate libraries. The sglang extra installs sglang-kt. A Docker directory is present in the repository for containerized deployment.

The README documentation is split into two entry points: kt-kernel/README.md for inference and doc/en/SFT/KTransformers-Fine-Tuning_Quick-Start.md for fine-tuning. The install.sh script at the repository root also supports setup.

## Inference Architecture: AMX, AVX, and Expert Scheduling

The kt-kernel component provides the CPU-optimized kernels for heterogeneous inference. The README describes AMX and AVX512/AVX2 optimized kernels for INT4 and INT8 quantized inference. Intel AMX (Advanced Matrix Extensions) is the fastest path, followed by AVX512, then AVX2 as a fallback added in March 2026 for CPUs without AMX or AVX512.

The CPU-GPU Expert Scheduling feature, added in January 2026, allows dynamic routing of MoE experts between GPU VRAM and CPU DRAM based on a scheduling policy. This is distinct from static offloading: instead of always keeping specific experts on CPU, the scheduler moves experts based on usage patterns.

A 3-layer prefix cache added in June 2025 stores reusable KV cache across GPU VRAM, CPU DRAM, and disk. The README describes this as GPU-CPU-Disk prefix cache reuse, allowing long-context inference to reuse cached prefixes without recomputing them.

For quantization, the framework supports INT4, INT8, BF16, FP8, and block-FP8 formats. The FP8 GPU kernel support was added in February 2025 specifically for DeepSeek-V3 and R1.

## Supported Models

KTransformers documents day-zero support for major MoE model releases. The README update log shows support for:

- DeepSeek-V3, DeepSeek-R1, DeepSeek-V4-Flash
- Kimi-K2, Kimi-K2-Thinking, Kimi-K2.5
- GLM-5, GLM-5.2, GLM-5.3-flash (with 1M-token context and multimodal input)
- Qwen3-Next
- MiniMax-M2.1, MiniMax-M2.5, MiniMax-M3
- LLaMA 4 (experimental)
- Mixtral 8x7B and 8x22B

Day-zero support means the team adds support at or near the time a model is publicly released. GLM-5.3-flash is noted in the August 2026 update as bringing 1M-token context and multimodal input to consumer GPUs. The README does not document the minimum VRAM required for each model; the DeepSeek-V3 and R1 tutorials are the most detailed, documenting 24GB VRAM as the single-GPU path.

The model support pattern reveals a constraint: KTransformers explicitly targets MoE architectures. Dense models like LLaMA 4 appear in the list but are marked experimental. The expert offloading optimization only yields large memory savings when the model has sparse MoE routing. For a dense model, the heterogeneous approach does not reduce VRAM requirements the same way, and a standard llama.cpp run would likely be simpler to configure.

## Fine-Tuning with LLaMA-Factory Integration

The SFT (Supervised Fine-Tuning) capability was added in November 2025 via an integration with LLaMA-Factory. The README links to a Fine-Tuning Cookbook that covers hardware checks, installation, BF16/FP8/INT8 recipes, LoRA and full fine-tuning, resource planning, and troubleshooting.

LoRA fine-tuning was extended in August 2026 to support compatible AVX512 x86 CPUs including AMD servers without requiring AMX. Block-FP8 LoRA fine-tuning was added in August 2026. Full-parameter BF16 fine-tuning with complete checkpoint saving was added in July 2026.

Kimi K2.5 and K2.6 LoRA fine-tuning with native RAWINT4 routed experts was the September 2026 addition. Fine-tuning runs on the same CPU-GPU heterogeneous hardware as inference, so the same DRAM requirements apply. The README does not publish minimum DRAM figures for the fine-tuning path.

RL-DPO fine-tuning with LLaMA-Factory was added in December 2025, extending the fine-tuning surface beyond supervised methods. The quick start entry point is doc/en/SFT/KTransformers-Fine-Tuning_Quick-Start.md rather than the top-level README, which means the fine-tuning documentation is separate from the inference documentation and has its own setup flow.

## Hardware and Platform Constraints

The primary development target is Intel x86-64 hardware with AMX or AVX512 support. AVX2-only support was added as a fallback for CPUs like AMD servers, but the README notes this as a separate tutorial entry point, suggesting it is less optimized than the AMX path. The March 2026 update introduced the AVX2 backend specifically for hardware without AMX or AVX512 instructions.

GPU support extends beyond NVIDIA. The README documents ROCm support for AMD GPUs (added March 2025), Intel Arc GPU support (added May 2025), and Ascend NPU support (added October 2025). DeepSeek-V4-Flash on a single Ascend NPU with CPU expert offload is documented in a tutorial from August 2026.

The pyproject.toml classifiers list only Linux. macOS support is not documented in the README or setup files. Teams on macOS should not assume compatibility without testing.

Multi-GPU inference was added in August 2024. Multi-concurrency for serving multiple simultaneous requests was added in April 2025. The balance-serve tutorial under doc/en/ covers the multi-concurrency setup. These features extend KTransformers from a single-user research tool toward a shared inference server, though the SGLang integration is the recommended path for production multi-user serving according to the README roadmap link.

The repository also documents an AutoDL cloud deployment path for users who want to test without local hardware. The README mentions an integrated KTransformers plus AutoDL plus LlamaFactory pipeline for fine-tuning and inference on rented cloud instances.

## KTransformers vs llama.cpp and Ollama

llama.cpp handles CPU inference for smaller quantized models efficiently and runs on macOS, Linux, and Windows with minimal dependencies. Ollama wraps llama.cpp in a server with a management interface. Both tools prioritize broad compatibility and ease of setup.

KTransformers targets a different use case: very large MoE models that require heterogeneous compute to fit in available hardware. The README notes a 3x to 28x speedup compared to a naive approach for DeepSeek-R1 and V3 on 24GB VRAM plus 382GB DRAM. That hardware profile (high-capacity DRAM, Intel AMX CPU, NVIDIA GPU) is not a typical consumer desktop.

The SGLang integration (sglang-kt extra) connects KTransformers to the SGLang serving framework, which handles batched multi-user inference. The README links to an SGLang roadmap issue from October 2025 and a blog post about the integration. llama.cpp and Ollama do not target this high-end heterogeneous server configuration.

For the Qwen3 model family, the README documents a tutorial for Qwen3-Next added in September 2025, and a separate mention of Ktransformers Qwen3 5 appears in the related search data. The repository also has a tutorial directory with per-model guides under doc/en/, which means the practical starting point for a new model is its specific tutorial rather than the generic kt-kernel README. Each tutorial specifies the exact hardware configuration, quantization format, and context length that was validated.

## Conclusion

KTransformers is for researchers and engineers who need to run large MoE models on hardware with limited GPU memory, particularly those working with DeepSeek, GLM, Kimi, or Qwen3 model families. It is not a drop-in replacement for llama.cpp or Ollama for general-purpose local inference of smaller models. Before adopting it, verify that your CPU supports the AMX, AVX512, or at minimum AVX2 instruction sets, since the performance difference between these backends is substantial and the AVX2 path is described as a fallback. Confirm also that you are running Linux and that your DRAM capacity matches the requirement documented in the per-model tutorial for the model you intend to run.

## FAQ

### What is KTransformers used for?

KTransformers enables CPU-GPU heterogeneous inference for large language models, particularly large Mixture-of-Experts models like DeepSeek-V3, DeepSeek-R1, and GLM variants. It offloads inactive model experts to CPU DRAM at inference time, allowing inference on GPUs with limited VRAM that could not otherwise hold the full model in memory.

### What is KTransformers?

KTransformers is a Python-based research framework from kvcache-ai focused on efficient inference and fine-tuning of large language models through CPU-GPU heterogeneous computing. It ships AMX and AVX512/AVX2 optimized kernels and integrates with LLaMA-Factory for supervised fine-tuning.

### What is the alternative to KTransformers?

For smaller quantized models on consumer hardware, llama.cpp and Ollama are the main alternatives, with broader platform support including macOS. For large MoE models requiring heterogeneous CPU-GPU inference, KTransformers targets a more specific hardware profile (high-DRAM Intel AMX systems) that llama.cpp does not optimize for.

## Sources

- [kvcache-ai/ktransformers on GitHub](https://github.com/kvcache-ai/ktransformers)
- [License: Apache-2.0](https://github.com/kvcache-ai/ktransformers/blob/main/LICENSE)
- [Project website](https://kvcache-ai.github.io/ktransformers/)
- [README](https://github.com/kvcache-ai/ktransformers/blob/main/README.md)
- [Releases](https://github.com/kvcache-ai/ktransformers/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kvcache-ai-ktransformers
