Model or dataset
ggml-org/llama.cpp avatar
ggml-org/llama.cpp

llama.cpp: LLM inference in C/C++ for local and cloud hardware

llama.cpp runs LLM inference in plain C/C++, serving models locally through a REST API and web UI with multimodal support in its llama-server.

129,881 stars23,899 forksC++MIT

At a glance

What is it?
llama.cpp runs LLMs and VLMs from a single C/C++ codebase across CPUs and GPUs, with a CLI and an OpenAI-compatible server. Here is how it installs, how the pieces fit, and where it stops being the right tool.
Who is it for?
Adopt llama.cpp when you want one C/C++ inference engine that covers CPU, Apple Silicon, CUDA, HIP, Vulkan, SYCL and more, and you are willing to manage model files yourself. Skip it if you want a desktop application with a model catalogue, or a high-throughput multi-tenant serving stack with built-in scheduling and metrics.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What llama.cpp actually solves

The project's stated goal is LLM and VLM inference with minimal setup and state-of-the-art performance on a wide range of hardware, locally and in the cloud. That sentence is doing more work than it looks. Most inference stacks assume a datacenter GPU and a Python process per model. llama.cpp assumes neither. It is a plain C/C++ implementation without any dependencies, and it is built on the ggml library, which is where the tensor operations and backend dispatch live. The repository keeps ggml as a submodule directory rather than vendoring a copy, so the two projects move together.

The audience follows from that. If you have a laptop with Apple Silicon, an x86 box with AVX2 or AVX512, a RISC-V board, an NVIDIA or AMD card, an Intel GPU, an Ascend NPU, or a Snapdragon device, the README lists a backend for you. If your model does not fit in VRAM, CPU+GPU hybrid inference is an explicit feature rather than a workaround. The trade-off is that you are closer to the metal than with a Python framework: you deal with GGUF files, quantization levels, and build flags.

GGUF files, quantization, and the ggml backend split

Two mechanisms determine almost everything about how llama.cpp behaves. The first is GGUF, the file format the project uses for models and the format its converters produce. The repository ships convert_hf_to_gguf.py, convert_lora_to_gguf.py, and convert_llama_ggml_to_gguf.py at the top level, plus a gguf-py package that the Python scripts import. A model that exists as Hugging Face weights becomes a single GGUF file before llama.cpp can load it.

The second is quantization. The README lists 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization, described as a way to get faster inference and reduced memory use. This is the main lever you have when a model is too large for the machine: drop the bit width until it fits, and accept the quality cost. Because GGUF carries the quantized weights, the choice is made at conversion or download time, not at runtime.

Underneath both sits ggml. The supported backend table is long: BLAS and BLIS for general hardware, CANN for Ascend NPU, CUDA for NVIDIA, HIP for AMD, MUSA for Moore Threads, Metal for Apple Silicon, OpenCL for Adreno, SYCL for Intel GPU, Vulkan and WebGPU for GPUs generally, ZenDNN for AMD CPUs, IBM zDNN for Z and LinuxONE, plus RPC, VirtGPU, and two entries marked In Progress (Hexagon for Snapdragon, OpenVINO for Intel CPUs, GPUs and NPs). The In Progress marker matters: those two are not finished, and the README says so.

Installing llama.cpp and running a first model

The README gives four routes: visit https://llama.app and follow the instructions, run with Docker per docs/docker.md, download pre-built binaries from the releases page, or build from source using docs/build.md. The Makefile at the repository root is not a build path anymore. It contains a deliberate error that tells you the Makefile build has been replaced by CMake and points at docs/build.md. If you find an old tutorial that says `make`, it is stale.

Once installed, the README's quick start is two commands. The first downloads a model directly from Hugging Face and starts an interactive session:

bash
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF

The `-hf` flag is what pulls the GGUF repository, so you do not need to fetch the file yourself first. The second command launches the server against the same model:

bash
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF

The README states that this is an OpenAI-compatible API server, and that it has a built-in web UI, which the screenshot table labels as the web UI against llama serve. For build-time options, docs/build.md is the reference; the README links backend-specific sections there by anchor, such as docs/build.md#cuda, #hip, #metal-build, #vulkan, #sycl, #cann, #musa, #webgpu, #zendnn, and #blas-build. Multi-GPU setups are covered separately in docs/multi-gpu.md, and performance troubleshooting has its own page at docs/development/token_generation_performance_tips.md.

Where llama.cpp is the wrong tool

The README does not document rollback for the server API, and it does not document an API versioning scheme. The server REST API is tracked in a GitHub issue rather than a specification document, and the lib llama API has its own issue as well. If you are building a product on top of the C API or the REST endpoints, that is a real constraint: you are depending on an interface whose stability story lives in an issue thread, not in a versioned contract. The README is silent on both.

There is also a versioning convention worth understanding before you pin anything. The releases listed are named b10679, b10678, b10677, all pushed on 2026-08-28. The README's release badge links to a query for tags matching `v0` and a separate query for `b` tags. Build numbers move quickly, and the project's own release process is documented in docs/release.md rather than in the README.

Finally, the backend table includes entries marked In Progress. If your hardware is Snapdragon or you were planning on OpenVINO for an Intel NPU, the README tells you the support is not finished. Choosing llama.cpp for an unfinished backend means accepting that the path may change.

llama.cpp compared with Ollama and vLLM

Ollama and vLLM come up constantly in searches alongside this project, and the difference is mostly about what layer you want to own. Ollama is a model runner with its own model management and a friendlier surface; llama.cpp gives you the inference engine and the file format, and leaves model acquisition, naming and lifecycle to you. The README's quick start does include `-hf` for pulling a model straight from Hugging Face, so the gap is narrower than it once was, but llama.cpp still does not present a curated catalogue the way a model-runner product does.

vLLM occupies a different position. It targets serving throughput with GPU-resident models in Python. llama.cpp's differentiator is breadth: the same C/C++ engine runs on Apple Silicon through Metal, on NVIDIA through CUDA, on AMD through HIP, on Intel through SYCL, on Adreno through OpenCL, and on plain CPUs through BLAS or BLIS, with CPU+GPU hybrid inference when the model exceeds VRAM. If your deployment spans a laptop and a server, that single engine is the argument. If your deployment is a rack of identical GPUs and you measure success in concurrent requests per second, llama.cpp's breadth buys you nothing you need.

A third option people search for is LM Studio, which is a desktop application rather than a library. Choosing between them is choosing between shipping an application and building on an engine.

Maintenance, upgrades, and the MIT licence

The repository is not archived, and the last push was on 2026-08-28. The three most recent releases, b10679, b10678 and b10677, were all pushed that same day, which tells you the build numbering advances in small increments. Upgrading therefore means deciding how often you want to move across build numbers, and the project documents its release process in docs/release.md rather than in the README.

The upgrade cost is not only the binary. The Python tooling has its own pinned dependency set. The pyproject.toml pins transformers to exactly 4.57.6, requires Python >=3.10,<3.15, and pins torch to 2.11.0 across platforms, with a separate CPU wheel index for Linux. requirements.txt is a set of includes and warns against adding packages there directly, telling contributors to avoid it and keep versions compatible across all top-level scripts. If you convert models yourself, those pins are part of your upgrade surface, not just the C++ build.

The licence is MIT, stated in the README badge and in the LICENSE file, and the pyproject.toml classifiers list the MIT License as well. MIT is permissive, but the project bundles third-party code with its own terms: the acknowledgements list cpp-httplib under MIT, nlohmann/json under MIT, and stb, miniaudio and subprocess.h as public domain. There is a licenses/ directory at the repository root and a vendor/ directory, which is where you would look before redistributing a binary. This is not legal advice; read those files yourself.

Editorial conclusion

Adopt llama.cpp when you want one C/C++ inference engine that covers CPU, Apple Silicon, CUDA, HIP, Vulkan, SYCL and more, and you are willing to manage model files yourself. Skip it if you want a desktop application with a model catalogue, or a high-throughput multi-tenant serving stack with built-in scheduling and metrics. Before committing, check the backend table in the README against your hardware, read docs/build.md for the flags that backend needs, and confirm on the releases page that a pre-built binary exists for your platform. If your deployment depends on the REST surface, read tools/server/README.md first, because the README does not document rollback or versioning for that API.

Frequently asked questions

What does llama.cpp do?

It performs LLM and VLM inference in plain C/C++, with the stated goal of minimal setup and strong performance across a wide range of hardware, locally and in the cloud. It is built on the ggml library and ships a CLI and an OpenAI-compatible server.

Is llama.cpp faster than Ollama?

The README makes no comparison with Ollama and publishes no benchmark figures, so there is no basis here for a speed claim. What the README does describe is the backend set llama.cpp targets, which includes CPU, Apple Silicon, CUDA, HIP, Vulkan and SYCL, and a performance troubleshooting page at docs/development/token_generation_performance_tips.md.

How do I install llama.cpp?

The README lists four routes: follow the instructions at https://llama.app, run with Docker per docs/docker.md, download pre-built binaries from the releases page, or build from source using docs/build.md. The root Makefile no longer builds anything and instead points you at docs/build.md.

How do I use the llama.cpp server?

The quick start runs `llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF`, which the README describes as launching an OpenAI-compatible API server with a built-in web UI. The server tool has its own documentation at tools/server/README.md.

What are llama.cpp models?

llama.cpp loads models in the GGUF format, which the project's Python converters produce from other formats. The repository ships convert_hf_to_gguf.py, convert_lora_to_gguf.py and convert_llama_ggml_to_gguf.py, along with a gguf-py package, and the README lists integer quantization from 1.5-bit up to 8-bit.

How do I install llama.cpp on Windows?

The README does not give Windows-specific steps. Its four install routes are the llama.app instructions, Docker via docs/docker.md, pre-built binaries from the releases page, and a source build via docs/build.md, and there is a winget workflow referenced in the README badges.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ggml-org-llama-cpp.svg)](https://hysenlabs.com/projects/ggml-org-llama-cpp)
Community notes

Community notes