# llama.cpp: seventeen backends, three Python manifests, and a Makefile that only errors

> A dependency-free C/C++ implementation of LLM and VLM inference, MIT licensed, with a nightly build stream that shipped three tagged builds in three hours. The Makefile is a signpost that aborts, the C++ build is CMake, and the Python conversion scripts are described by three different manifests that do not agree on a torch version.

**ggml-org/llama.cpp** — llama.cpp runs LLM inference in plain C/C++, serving models locally through a REST API and web UI with multimodal support in its llama-server.

- Repository: https://github.com/ggml-org/llama.cpp
- Website: https://llama.app
- Stars: 129,881 · Forks: 23,899
- Language: C++
- License: MIT
- Published: 2026-08-04 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/ggml-org-llama-cpp

## The Makefile's entire content is an error that points at the build guide

This is the whole file.

```make
$(error Build system changed:
The Makefile build has been replaced by CMake.
For build instructions see:
https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md)
```

Someone kept the file and made it fail on purpose. If you type `make` in a fresh clone, nothing builds and nothing is created, and instead of a compiler error you get a message naming CMake and a documentation URL.

That is a courtesy and a hazard at the same time. Every tutorial, Docker image, CI snippet and Stack Overflow answer written before the migration still says `make`, and a copy-paste of one of those now fails with something that looks like a broken checkout rather than an intentional guard. The real entry points are CMakeLists.txt and CMakePresets.json, with docs/build.md as the guide. The Makefile is retained only to convert a stale instruction into a pointer, which is the cheapest possible migration path for a project with this many external references.

## The Python side has three manifests, and one of them forbids edits

The C++ has CMake. The Python has something more complicated: pyproject.toml, a requirements.txt, and a directory of per-script requirement files behind the requirements.txt.

That file states its own rules at the top.

```bash
# These requirements include all dependencies for all top-level python scripts
# for llama.cpp. Avoid adding packages here directly.
#
# Package versions must stay compatible across all top-level python scripts.
```

Then it includes six files by reference, one per script: the legacy llama converter, the Hugging Face converter and its update script, the llama-to-ggml converter, the LoRA converter, and the tool bench.

So requirements.txt is a union, it is a generated view of six other files, and the rule is that you must edit the specific file, not the union. The compatibility rule is the expensive half. Because every top-level script shares one pinned set, a single conversion script that needs a newer release of a library cannot take it, and the fix has to be negotiated across six files. That is the price of a flat repository where the scripts sit at the root rather than in per-project environments.

## torch is pinned twice in one file, and the two pins use different indexes per platform

pyproject.toml declares its dependencies in the standard table and then, further down, declares torch again in a second tool's table.

```toml
dependencies = [
    'numpy (>=2.2.6,<3.0.0)',
    'sentencepiece (>=0.1.98,<0.3.0)',
    'transformers (==4.57.6)',
    'protobuf (>=4.21.0,<5.0.0)',
    'torch (>=2.6.0,<3.0.0)',
    'gguf @ ./gguf-py',
]
```

The second table pins torch to an exact 2.11.0, and splits the source by platform: PyPI for darwin and win32, and a different index at download.pytorch.org for linux, where the version carries a +cpu suffix. A third table, for another tool, points torch at that same index.

Three things follow. Your resolved torch is not the same artifact on Linux as on macOS or Windows, so a conversion script verified on one platform is not verified on another. The range and the exact pin live in the same file and can drift apart without either being wrong on its own terms. And transformers is held at a single exact version, which is the correct choice for a converter whose output must be bit-identical across machines and a hard constraint on anyone who wants a newer release in the same environment.

## Two release streams, and the nightly one produced three tags in three hours

The three most recent releases are b11268, b11269 and b11270, published at 00:46, 01:17 and 03:52 on the same day. Those are not version bumps, they are build numbers on a rolling stream, and the README carries separate links for the b stream and for the v0 tag stream.

So there are two ways to install a binary and they have completely different meanings. The v0 stream is the one to pin. The b stream is effectively a nightly, and the README's own quick start tells you to download pre-built binaries from the releases page, which is a single page carrying both. Nothing on that page distinguishes them for you.

The release process has its own document, docs/release.md, and the quick start also points at llama.app as a first option. Three separate routes to a binary, at three levels of reproducibility, and the difference between them is a tag prefix. Pin the tag, not the word latest, and note that the last push to master was on 2026-09-29, so master moves continuously rather than in releases.

## CPU plus GPU hybrid is how you run a model larger than your VRAM

One bullet in the feature list is the one that changes what you can attempt: CPU plus GPU hybrid inference, to partially accelerate models larger than the total VRAM capacity.

That is the answer to the question everyone asks first, which is why it is in a short list of features rather than in a document. It also has a cost that the one-line description does not carry. Offloading part of a model to the GPU and part to system RAM means the tokens-per-second number you get depends on a transfer you cannot see and the project cannot schedule around, and it depends on the shape of the model rather than on your hardware alone.

The documentation structure reflects that this is two separate problems. Building is docs/build.md, and performance is a different document, docs/development/token_generation_performance_tips.md, alongside docs/multi-gpu.md for spreading across devices. If you are at the stage where you are reading performance tips, you have already finished the build and started the second problem.

## Seventeen backends, and the only one marked In Progress covers Intel and NPUs

The backend table is a compatibility map rather than a feature list, and it is worth reading as one.

The targets span a lot more than GPUs. CANN targets an Ascend NPU. Hexagon targets Snapdragon. IBM zDNN targets IBM Z and LinuxONE. ZenDNN targets AMD CPUs. VirtGPU targets the VirtGPU APIR. OpenCL targets an Adreno GPU, which is a phone. WebGPU and BLAS and BLIS and RPC are marked as covering all devices.

Exactly one row is qualified, and it is the one you would want on a modern laptop: OpenVINO, marked In Progress, for Intel CPUs, GPUs and NPUs.

The qualification is the useful information. A device being absent from the table means unsupported, and a device whose row says In Progress means the backend exists but is not finished, and that is the difference between a build that fails at configure time and one that miscompiles or miscounts. The RPC backend is also notable for what it implies: it lives at tools/rpc, so a machine with no accelerator at all can serve inference to one that has.

## Vendored public-domain code is what makes the multimodal subsystem work

The acknowledgements section is a licence inventory, and reading it tells you what the binary actually contains.

The HTTP server behind llama-server is a single-header library under MIT. The image format decoder used by the multimodal subsystem is stb, under public domain, and so is the audio format decoder in the same subsystem, miniaudio. JSON handling across the tools and examples is nlohmann/json under MIT, and process launching is a public-domain single-header solution for C and C++.

llama.cpp itself is MIT, and vendor/ and licenses/ exist to hold this inventory. So the build you produce is licence-clean in the ordinary sense, every vendored component is permissive or public domain, and there is no copyleft in the list.

The practical consequence is that you inherit the list, not just the licence. If your own policy requires an inventory of third-party code in a shipped binary, this is the manifest, and it names the exact components rather than leaving you to find them. The second-order point is which subsystems they serve: the multimodal ones are the public-domain decoders, so anything you build that accepts images or audio pulls those in.

## Conclusion

Use llama.cpp when you need inference on hardware no framework supports well, when the model has to stay on the machine, or when you want an OpenAI-compatible server from a single binary, and reach for a pre-built release rather than a build from source unless you need a backend the binaries do not carry. Do not plan on `make` working, because the Makefile is a deliberate error stub, and do not treat the Python conversion scripts as a library you can depend on, because their version constraints are declared in three files and already disagree about torch. Before you pin anything, check which of the two release streams you want, because the rolling b-numbered tags are hours old and the v0 tags are the stable line. If you are converting a model, read the requirements comment first, since it forbids adding packages to requirements.txt and requires every top-level script to stay on compatible versions.

## FAQ

### What does llama-cpp do?

It is a plain C and C++ implementation of LLM and VLM inference with no dependencies, built on top of the ggml library, aimed at running models locally and in the cloud across a wide range of hardware. It supports integer quantization from 1.5-bit to 8-bit, CPU plus GPU hybrid inference for models larger than the total VRAM capacity, and seventeen hardware backends including CUDA, Metal, Vulkan, SYCL and WebGPU.

### How do I install llama.cpp?

Four routes are offered: follow the instructions at llama.app, run it with Docker using the project's Docker documentation, download pre-built binaries from the releases page, or build from source with the build guide. Note that the repository's Makefile only prints an error saying the Makefile build has been replaced by CMake, so a `make` instruction from an older tutorial will not work.

### How do I use the llama.cpp server?

The `llama serve` command launches an OpenAI-compatible API server and pulls a model directly from Hugging Face, for example `llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF`. It has a built-in web UI, and the server REST API and the lib llama API are both tracked as separate issues in the repository. The HTTP server is built on the single-header cpp-httplib library.

### What are llama-cpp models?

The tools take GGUF files, and the quick start downloads one by name from Hugging Face with the `-hf` flag, for example ggml-org/Qwen3.5-0.8B-GGUF. The project supports 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit and 8-bit integer quantization for reduced memory use, and ships Python conversion scripts that turn Hugging Face, LoRA and legacy llama formats into GGUF.

### How do I use llama.cpp on a Mac?

Apple silicon is described as a first-class target, optimized through ARM NEON, Accelerate and the Metal frameworks, and the Metal backend's target device is Apple Silicon. On the x86 side there is AVX, AVX2, AVX512 and AMX support, and on RISC-V there is RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support, so the instruction set you get depends on the machine.

## Sources

- [Official documentation](https://llama.app)
- [Official README](https://github.com/ggml-org/llama.cpp#readme)
- [Project repository](https://github.com/ggml-org/llama.cpp)
- [Release notes](https://github.com/ggml-org/llama.cpp/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ggml-org-llama-cpp
