gemma.cpp: A Minimal C++ Inference Engine for Gemma, PaliGemma and Research Work
lightweight, standalone C++ inference engine for Google's Gemma models.
At a glance
- What is it?
- gemma.cpp is Google's standalone C++ runtime for Gemma 2-3 and PaliGemma 2, built around a roughly 2K LoC core, Highway SIMD, and a training path. It is aimed at people who want to modify the model, not ship it.
- Who is it for?
- Adopt gemma.cpp if you need to read, modify or embed the Gemma forward pass in C++ and are comfortable building from source with Clang and CMake. Skip it if you need GPU inference, a stable packaged release, or a supported conversion path from Safetensors, since the README states that conversion is not yet open sourced.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
The gap gemma.cpp is trying to fill
Deployment-oriented C++ runtimes are tuned for serving, and Python ML frameworks hide the low-level computation behind a compiler. The README frames gemma.cpp as the space between those two: a direct implementation you can read end to end. The stated core is about 2K lines of code, with roughly 4K more lines of supporting utilities. That size is the product. If you want to change how attention is computed, or swap a quantization scheme, you are editing a small tree rather than tracing through a graph compiler.
The intended audience is narrow and the README says so: experimentation and research use cases, plus embedding in other projects with minimal dependencies. It covers Gemma 2, Gemma 3 and PaliGemma 2, and it is CPU-only. The README explicitly points production edge deployments at Python frameworks such as JAX, Keras, PyTorch and Transformers instead. Treat that as the project telling you where it does not want to be used.
What the engine actually contains
Inference runs on the CPU through the Google Highway library, which provides a single SIMD implementation that selects the instruction set at runtime. That is why the README can claim portability across Linux, Windows and OS X while still using vector code. Sampling is limited to TopK with temperature. There is a backward pass (VJP) and an Adam optimizer, which is unusual for an inference engine and signals that training experiments are in scope.
The numerics section is where most of the engineering sits. GEMM is mixed precision across fp8, bf16, fp32 and fp64, designed around BF16 instructions and able to emulate them. Runtime autotuning covers seven parameters per matrix shape. Weight compression is integrated directly into GEMM, including a custom fp8 format with 2 to 3 mantissa bits and tensor scaling, plus bf16, f32 and non-uniform 4-bit (NUQ). The README says new formats are easy to add, which fits the small-core design.
Infrastructure details matter if you plan to run this on a server rather than a laptop. Tensor parallelism is CCX-aware with a multi-socket thread pool. Disk I/O is either memory-mapped or parallel reads, chosen by a heuristic that the user can override. Model weights use a custom format with forward and backward compatible metadata serialization. Frontends are a C++ API with streaming for single and batched queries, a basic interactive command-line app, and basic Python bindings via pybind11.
Building gemma.cpp and running the 2B checkpoint
The README lists CMake, a Clang C++ compiler supporting at least C++17, and tar as prerequisites. On Windows, native builds need the Visual Studio 2022 Build Tools with the Clang/LLVM frontend, installed through winget. On WSL, the README warns to set the number of parallel build threads to 1, because a larger number may produce errors.
Weights come from Kaggle. Visit the Gemma-2 page, select the gemmaCpp variation, fill out the consent form, and you receive an archive.tar.gz. Extract it with tar, which yields files such as 2b-it-sfp.sbs and tokenizer.spm. The README strongly recommends starting with the gemma2-2b-it-sfp model, and notes that bfloat16 weights are higher fidelity while 8-bit switched floating point weights give faster inference.
tar -xf archive.tar.gzConfiguration uses CMake presets. The README shows configuring the build directory with cmake --preset make, then building with cmake --build --preset make and a -j value. If you are unsure of the thread count, plain make gemma still produces the ./gemma executable. If you reconfigure with different settings, the README says to clear the build directory first with rm -rf build/*.
cmake --preset make
cmake --build --preset make -j4The result is a ./gemma binary in the build directory. The README describes it as a basic interactive command-line app, so expect a prompt rather than a server. There is a separate API_SERVER_README.md at the repository root for anyone who needs the server shape, and examples/hello_world/ plus examples/simplified_gemma/ for smaller reference programs.
Where gemma.cpp is the wrong tool
The most concrete limitation is stated plainly: model conversion from Safetensors is not yet open sourced. You cannot take an arbitrary Gemma checkpoint off the Hugging Face Hub and feed it in. You are dependent on the gemmaCpp-specific artifacts published on Kaggle, and the README points to the ModelPrefix function in configs.cc for how model names map to sizes. If your workflow requires converting your own fine-tuned weights, that path is closed today.
CPU-only inference is the second boundary. There is no GPU backend in the feature list, and the README redirects production edge deployments to Python frameworks. On a 27B model, the practical ceiling is set by your CPU, memory bandwidth and the NUQ or fp8 compression you choose. The project gives you the knobs but does not promise serving throughput.
A third issue is release cadence. The most recent tagged release listed is v0.1.4 from 2025-03-25, following v0.1.3 on 2025-03-14 and v0.1.2 on 2024-04-05. The repository itself is not archived and its last push was on 2026-09-21, so work is happening, but the version numbers have not moved past 0.1.x. If your process requires pinning to a tagged release with a changelog, you will be tracking a branch, not a release train.
gemma.cpp versus llama.cpp and the Python stacks
The README names ggml, llama.c and llama.rs as the inspiration for the vertically integrated approach, and llama.cpp is the obvious comparison point because both are C++ runtimes with quantization and CPU inference. The difference is intent. llama.cpp is a general runtime covering many model families with a broad tooling surface for conversion and serving. gemma.cpp covers Gemma 2, Gemma 3 and PaliGemma 2 only, and spends its complexity budget on a small, modifiable core instead of model coverage. If you need to run a model that is not Gemma, gemma.cpp is not a candidate at all.
The Python comparison is the one the README makes itself. JAX, Keras, PyTorch and Transformers are the recommended production path, and they abstract the low-level computation through compilation. gemma.cpp does the opposite: it keeps the computation visible. The trade is that you get direct control over GEMM precision, compression formats and thread pool behavior, and you lose the ecosystem, the GPU support and the conversion tooling. The backward pass and Adam optimizer are the clearest signal of which side of that trade the project chose.
Maintenance, branches and licence
The README carries a note that active development happens on the dev branch and that pull requests should target dev rather than main, which is described as intended to be more stable. That is a two-branch model where the stable branch and the development branch can diverge. Anyone building from main and anyone building from dev may not be running the same code, so record which commit you built from.
The licence is Apache-2.0, and the repository also carries a LICENSE-BSD3 file, which suggests some components are under a BSD 3-clause licence. The README does not break down which files fall under which, so if you are redistributing a binary you need to read both files rather than assume a single licence covers the tree. This is a description of what is in the repository, not legal advice.
On upgrade cost, the custom weight format is described as having forward and backward compatible metadata serialization, which reduces the risk that a new build rejects your existing .sbs files. The README does not document rollback, and it does not describe a migration procedure between versions, so the safe assumption is that moving between commits is untested from the user's side.
Editorial conclusion
Adopt gemma.cpp if you need to read, modify or embed the Gemma forward pass in C++ and are comfortable building from source with Clang and CMake. Skip it if you need GPU inference, a stable packaged release, or a supported conversion path from Safetensors, since the README states that conversion is not yet open sourced. Verify first that a Kaggle gemmaCpp checkpoint exists for the model size you want, and confirm your toolchain meets the C++17 Clang requirement before configuring the build.
Frequently asked questions
Is Gemma free to use?
The gemma.cpp repository itself is published under Apache-2.0, and it also carries a LICENSE-BSD3 file. The README does not describe the terms that apply to the model weights, which are distributed through Kaggle behind a consent form.
How is Gemma at coding?
The README does not report coding benchmarks or any task-level evaluation results for the models. It describes the engine's features, such as TopK and temperature sampling, rather than model quality.
What is the Google Gemma model?
Gemma is Google's family of foundation models, and gemma.cpp is a standalone C++ inference engine for them. According to the README, the engine implements Gemma 2, Gemma 3 and PaliGemma 2, and points to ai.google.dev/gemma for more on the models themselves.
What is Llama.cpp used for?
The README does not describe llama.cpp's use cases; it only cites ggml, llama.c and llama.rs as inspiration for a vertically integrated implementation. gemma.cpp is a separate project that covers Gemma 2, Gemma 3 and PaliGemma 2 only.
What is gemma.cpp?
It is a lightweight, standalone C++ inference engine for Google's Gemma models, covering Gemma 2, Gemma 3 and PaliGemma 2. The README describes it as a minimalist implementation aimed at experimentation and research rather than production deployment.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/google-gemma-cpp)
Community notes