Model or dataset
Tiiny-AI/PowerInfer avatar
Tiiny-AI/PowerInfer

PowerInfer: running LLMs on a consumer GPU with hot and cold neurons

High-speed Large Language Model Serving for Local Deployment

9,815 stars604 forksC++MIT

At a glance

What is it?
PowerInfer splits LLM inference between GPU and CPU by preloading frequently activated neurons and computing the rest on the CPU. It is built for ReLU-sparse models on a single consumer GPU, and its own README limits how far that idea travels.
Who is it for?
Adopt PowerInfer if you run a ReLU-sparse model such as Falcon-40B, Llama2, ProSparse Llama2 or Bamboo-7B on one consumer NVIDIA GPU under Linux or Windows and want the hot/cold split. Do not adopt it for dense model weights or for macOS performance work: the README states that llama.cpp weights work but yield no performance gain, and that Mac is CPU only with no significant improvement.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 142 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

The problem PowerInfer targets: a single consumer GPU, a model that does not fit

Serving a large language model locally usually forces a choice. You either buy server-grade hardware, or you accept that most of the model runs on the CPU and generation crawls. PowerInfer is aimed at the second case and tries to avoid the crawl. The README frames the target as a personal computer with a single consumer-grade GPU, and the design rests on one observation about how LLM inference behaves: neuron activation follows a power-law distribution. A small set of neurons, which the project calls hot neurons, fire across almost every input. The rest, the cold neurons, depend on the specific input.

That split is what makes a hybrid engine worth building. If a predictable subset of the model is always needed, that subset can live in GPU memory permanently, and the unpredictable remainder can be computed on the CPU without pulling weights across the PCIe bus on every token. The README states the goal directly: reduce GPU memory demands and CPU-GPU data transfers. The audience is anyone with one graphics card and a model too big for its VRAM.

How the hot and cold neuron split actually runs

The mechanism has three cooperating parts. First, an adaptive predictor decides, at inference time, which neurons are likely to activate. Second, neuron-aware sparse operators compute only the activated subset rather than the full dense matrix. Third, the placement policy keeps hot-activated neurons preloaded on the GPU for fast access while cold-activated neurons are computed on the CPU.

The predictor is the part that decides whether the whole scheme pays off. A wrong prediction means either wasted GPU work or a cold neuron computed in the wrong place. The README describes the predictor as adaptive, which implies it adjusts rather than using a fixed offline list, but it does not document the prediction algorithm, its accuracy, or what happens on a misprediction. That is a real gap if you are evaluating this for production serving rather than a demo.

The repository layout shows the engine inherits a large amount of ggml infrastructure: ggml.c, ggml-cuda.cu, ggml-metal.m, ggml-opencl.cpp, ggml-quants.c and ggml-alloc.c all sit at the top level, alongside llama.h and a llama.cpp directory. The build also carries build.zig, flake.nix and Package.swift, so CMake is not the only entry point, though the README documents CMake as the pre-requisite.

Installing PowerInfer and running a first model

The README lists CMake as a pre-requisite and points to its own Installation section for the rest. The repository also ships a requirements.txt for the Python-side conversion tooling, which pins numpy, sentencepiece and transformers and installs the local gguf-py and powerinfer-py packages.

Start by installing the Python conversion dependencies from the repository root:

bash
pip install -r requirements.txt

After that, the README directs you to the Setup and Installation section for the CMake build, then to Model Weights and Inference. The repository provides convert.py, convert-dense.py and convert-hf-to-powerinfer-gguf.py for turning model weights into the format the engine expects. The README does not spell out the exact arguments for these scripts, so read the scripts and the Installation section rather than guessing flags.

Once a model is available, the README states that most of the examples/ directory can be used the same way as llama.cpp, including server and batched generation. That means the existing shell and batch scripts in examples/, such as chat-13B.sh and chat-13B.bat, are the closest thing to a documented first run. The README does not give an exact invocation for them, so treat the script contents as the source of truth.

The sparse model requirement is the real constraint

PowerInfer is not a general-purpose replacement for a dense inference engine. The README is explicit that it is compatible with ReLU-sparse models and lists what it supports: Falcon-40B, the Llama2 family, the ProSparse Llama2 family and Bamboo-7B. The performance story depends on activation sparsity. The project's own TurboSparse work sparsified Mistral and Mixtral to nearly 90% sparsity, and the README notes that TurboSparse-Mixtral activates only 4B parameters.

The backward compatibility note is where the limitation becomes concrete. PowerInfer can load llama.cpp model weights, but the README states there will be no performance gain. So if your model is dense and you were hoping the hot/cold split would speed it up, the engine will run it without the benefit that justifies using PowerInfer at all. You would be carrying the extra predictor and sparse-operator machinery for nothing.

Platform coverage is the second constraint. The README lists x86-64 CPUs with AVX2 under Linux and Windows, with or without NVIDIA GPUs, and Apple M chips on macOS as CPU only. It adds that the project does not optimize for Mac and that the performance improvement there is not significant. A Metal backend for sparse inference is listed as coming soon, not as available. If you are on Apple silicon and want the speedup, this is the wrong tool today.

PowerInfer compared with llama.cpp and Ollama

The comparison the README makes is with llama.cpp. The two share lineage: PowerInfer's examples/ directory mirrors llama.cpp's, and it can consume llama.cpp weights. The difference is the execution strategy. llama.cpp runs the model as a dense computation, splitting work across CPU and GPU by layer or by tensor. PowerInfer splits by neuron, keeping the consistently activated ones resident on the GPU and routing the rest to the CPU. On a single RTX 4090 running Falcon(ReLU)-40B-FP16 with 24 GB of VRAM, the README reports up to an 11.69x speedup over llama.cpp, with both engines fully utilizing VRAM.

Ollama is a different kind of comparison, and the README does not make it. Ollama is a model runner and packaging layer: it manages downloads, serves an HTTP interface and handles the model lifecycle. PowerInfer is an inference engine with a specific optimization for sparse activation. The honest framing is that they occupy different layers. If you want the easiest path to running a quantized dense model, an engine-agnostic runner is the practical choice. If you specifically want to serve a 40B sparse model on one 24 GB card and are willing to build from source, PowerInfer's hot/cold split is the reason to pick it.

Maintenance, licensing and what the repository does not promise

The repository is not archived, and the last push was on 2026-05-11. The README's news entries run through 2026/1/5, when the team announced Tiiny AI Pocket Lab, described as a pocket-size computer running GPT-OSS-120B at int4 locally at 20 tokens/s. Earlier entries cover PowerInfer-2 for smartphones, the TurboSparse models and AMD ROCm support added on 2024/5/17. There are no retrieved releases, so there is no versioned release channel to track; you would be following the main branch.

The licence is MIT, which is permissive and places few obligations on how you redistribute or modify the code. That is a statement about the licence text, not legal advice; if you are embedding PowerInfer in a product, read the LICENSE file and get your own counsel.

Upgrade cost is the part the repository does not answer. With no releases and no documented compatibility policy for model formats, moving to a newer commit means rebuilding and re-verifying your converted weights. The README does not document rollback, and it does not state a supported upgrade path. Budget for rebuilding from source rather than for a package upgrade.

Editorial conclusion

Adopt PowerInfer if you run a ReLU-sparse model such as Falcon-40B, Llama2, ProSparse Llama2 or Bamboo-7B on one consumer NVIDIA GPU under Linux or Windows and want the hot/cold split. Do not adopt it for dense model weights or for macOS performance work: the README states that llama.cpp weights work but yield no performance gain, and that Mac is CPU only with no significant improvement. Before committing, check the CMake pre-requisites in the README, confirm your CPU has AVX2, and verify that the model you intend to serve is one of the listed sparse families.

Frequently asked questions

How does PowerInfer compare with Ollama?

They sit at different layers. Ollama is a model runner and packaging layer, while PowerInfer is an inference engine built around a hot/cold neuron split for ReLU-sparse models. The README does not compare the two directly.

What is PowerInfer?

It is a C++ CPU/GPU LLM inference engine that exploits activation locality, preloading hot neurons onto the GPU while cold neurons are computed on the CPU. The README describes it as designed for local deployment on consumer-grade hardware.

Which models can PowerInfer serve?

The README lists Falcon-40B, the Llama2 family, the ProSparse Llama2 family and Bamboo-7B. It also states that llama.cpp model weights are supported for compatibility, but without any performance gain.

Does PowerInfer work on macOS?

The README says PowerInfer has been tested on Apple M chips on macOS with CPU only, and that the project does not optimize for Mac, so the performance improvement is not significant there. A Metal backend for sparse inference is listed as coming soon.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. Tiiny-AI/PowerInfer on GitHub
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tiiny-ai-powerinfer.svg)](https://hysenlabs.com/projects/tiiny-ai-powerinfer)
Community notes

Community notes