Model or dataset
abetlen/llama-cpp-python avatar
abetlen/llama-cpp-python

llama-cpp-python: Python bindings for llama.cpp, and when the build is the hard part

Python bindings for llama.cpp

10,637 stars1,460 forksPythonMIT

At a glance

What is it?
llama-cpp-python wraps llama.cpp in a ctypes layer, a high-level completion API and an OpenAI-compatible server. The package works, but installing it with the right GPU backend is where most of the effort goes.
Who is it for?
Adopt llama-cpp-python when you want llama.cpp's GGUF inference reachable from Python without writing a C extension, and when you are willing to own the cmake build flags for your hardware. Skip it if you need a managed multi-tenant serving stack with per-request isolation and an operations story; the bundled server is a single-process FastAPI app, and the README does not document rollback or upgrade procedures for the compiled extension.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 7 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What llama-cpp-python actually wraps

The project is a Python package around @ggerganov's llama.cpp library. The README lists what it provides: low-level access to the C API through a ctypes interface, a high-level Python API for text completion, and an OpenAI compatible web server. That layering matters more than it sounds. The ctypes layer means there is no compiled Python extension to import; the shared library built from llama.cpp is loaded and called through foreign function signatures. The high-level API sits on top of that, and the server sits on top of the high-level API.

Who is this for? Two groups. First, Python developers who want local GGUF model inference inside an existing Python process, without running a separate inference daemon. Second, teams that want an OpenAI-shaped HTTP endpoint on their own hardware, because the server exposes code completion, function calling, vision and multiple models. If you are already committed to a hosted API and have no reason to run weights locally, this package adds build complexity for nothing.

The ctypes layer, the completion API and the FastAPI server

Data flow is straightforward. A GGUF model file is loaded by the underlying llama.cpp code. The Python high-level API takes a prompt, applies the chat template through jinja2 (a declared dependency in pyproject.toml), calls into the C API, and streams tokens back. The ctypes interface is what makes the boundary crossing cheap and also what makes error reporting less friendly: failures inside the native library surface as Python-side exceptions with limited context.

The server is a separate optional install. pyproject.toml defines a server extra pulling uvicorn, fastapi, pydantic-settings, sse-starlette, starlette-context and PyYAML. That is a small, conventional stack, and it means the server is not a distributed system. It is one process holding the model in memory. The README advertises multiple model support in the server configuration, and the repository has an examples/server directory, but nothing in the documentation describes request queueing, admission control or memory reclamation when several models are loaded at once. Treat multi-model mode as a convenience for switching between small models, not as a serving architecture.

Installing llama-cpp-python, including the CUDA and Metal wheel paths

The base install compiles llama.cpp from source as part of the pip install, so a C compiler is required: gcc or clang on Linux, Visual Studio or MinGW on Windows, Xcode on macOS. Python 3.8 or newer.

bash
pip install llama-cpp-python

If the build fails, the README says to add --verbose to see the full cmake build log. That is the first thing to try, because almost every failure at this stage is a cmake configuration problem rather than a Python packaging problem.

There is also a pre-built wheel with basic CPU support, which avoids the compiler entirely:

bash
pip install llama-cpp-python \
  --extra-index-url https://abetlen.github.io/llama-cpp-python/whl/cpu

For GPU acceleration you either pass cmake options through CMAKE_ARGS or use a pre-built wheel. The CUDA wheel path requires a specific CUDA version, a compute capability floor and a Python version of 3.10, 3.11 or 3.12. The index suffix selects the CUDA build:

bash
pip install llama-cpp-python \
  --extra-index-url https://abetlen.github.io/llama-cpp-python/whl/cu121

On macOS, Metal has its own index and requires macOS 11.0 or later with Python 3.10 to 3.12:

bash
pip install llama-cpp-python \
  --extra-index-url https://abetlen.github.io/llama-cpp-python/whl/metal

If no wheel fits, build from source with the backend flag. The Makefile in the repository shows the same pattern for CUDA, OpenBLAS, Metal, Vulkan, Kompute, SYCL and RPC:

bash
CMAKE_ARGS="-DGGML_CUDA=on" pip install llama-cpp-python

A first real use is the example layout the repository ships: examples/high_level_api for the completion API, examples/low_level_api for the ctypes surface, examples/server for the HTTP endpoint, and examples/hf_pull for pulling a model from Hugging Face. Start with the high-level example before touching the server.

Where the build flags and wheel matrix bite

The genuine limitation is not the Python API. It is the build. Every acceleration backend is selected at compile time through CMAKE_ARGS, and the resulting artifact is tied to that choice. A wheel built for cu124 will not help you on a machine with CUDA 11.8, and the pre-built CUDA wheels list explicit constraints: CUDA 11.8, 12.1, 12.2, 12.3, 12.4, 12.5, 13.0 or 13.2, with compute capability 6.0 through 8.9 for the 11.8 wheels, 6.0 or newer for CUDA 12, and 7.5 or newer for CUDA 13. Anything outside that set means compiling from source, which means a working toolchain and a cmake build that takes real time.

This is also the wrong tool in a specific case: if you need a stable binary artifact you can pin and promote across environments without a compiler on the target host, llama-cpp-python's source path fights you. The wheel indexes exist, but they cover a fixed matrix, and the README does not document a fallback policy for hardware outside it. The Makefile's build.debug target exists for diagnosing this, and build.debug.extra adds address sanitizer flags, which tells you the maintainers expect native-level debugging to be part of the workflow.

Compared with Ollama, which hides the build entirely

The obvious alternative for the same job is Ollama. The difference in approach is where the complexity lives. Ollama ships a standalone runtime and a model registry, and you talk to it over HTTP; you do not compile a backend, and you do not manage a Python binding to a C library. llama-cpp-python instead puts llama.cpp inside your process, which means you can call the completion API directly, pass your own callbacks, and avoid a network hop and a second daemon.

The trade is real in both directions. With llama-cpp-python you inherit the build matrix, the CUDA and compute capability constraints, and the fact that upgrading the package may mean recompiling against a newer vendored llama.cpp. The repository carries vendor/llama.cpp as a submodule and the Makefile has update.vendor to pull its master branch, so the binding tracks upstream rather than freezing it. With Ollama you inherit its model management and its release cadence instead. If your application is Python and the model call is one function among many, the in-process binding is the better fit. If your application is not Python, or you want the runtime to be somebody else's operational problem, it is not.

Licence, release cadence and the cost of upgrading

The licence is MIT, declared in both LICENSE.md and pyproject.toml. That is permissive and imposes no copyleft obligation on your own code. It does not resolve the licensing of the model weights you load, and it says nothing about the licences of the optional server dependencies, which you should check separately if you redistribute a bundled environment.

On maintenance: the repository is not archived, and the last push was on 2026-09-20. Releases are split by backend. The three most recent are v0.3.35-vulkan, v0.3.35-rocm72 and v0.3.35-metal, all dated 2026-08-17. That naming is the upgrade story in miniature. There is no single version that covers every backend, so an upgrade means picking the tag that matches your hardware, and the changelog is the place to check what moved. Because llama.cpp is vendored as a submodule, a package upgrade can change inference behaviour even when your own code does not change. Pin the version, and test prompt outputs after any bump.

The README does not document a rollback procedure for the compiled extension, and it does not describe how to keep two backend builds side by side in one environment. If you need that, plan for separate virtual environments per backend.

Editorial conclusion

Adopt llama-cpp-python when you want llama.cpp's GGUF inference reachable from Python without writing a C extension, and when you are willing to own the cmake build flags for your hardware. Skip it if you need a managed multi-tenant serving stack with per-request isolation and an operations story; the bundled server is a single-process FastAPI app, and the README does not document rollback or upgrade procedures for the compiled extension. Before committing, verify that the pre-built wheel index for your backend matches your CUDA version and GPU compute capability, since a mismatch sends you back to a source build.

Frequently asked questions

What is llama-cpp-python used for?

It provides Python bindings for the llama.cpp library, giving low-level access to the C API through ctypes, a high-level Python API for text completion, and an OpenAI compatible web server. It also lists LangChain and LlamaIndex compatibility in the README.

How do I install llama-cpp-python?

Run pip install llama-cpp-python, which builds llama.cpp from source and requires a C compiler plus Python 3.8 or newer. If the build fails, the README says to add --verbose to see the full cmake build log.

How do I install llama-cpp-python with CUDA support?

Either set CMAKE_ARGS="-DGGML_CUDA=on" before pip install, or use a pre-built wheel from the CUDA index. The wheels require a specific CUDA version, a compute capability floor, and Python 3.10, 3.11 or 3.12.

How do I install llama-cpp-python on Windows?

The same pip install applies, but the build requires Visual Studio or MinGW as the C compiler. The README shows setting CMAKE_ARGS through $env:CMAKE_ARGS in PowerShell, and there is a pre-built HIP Radeon wheel index for Windows.

How do I install llama-cpp-python on a Mac?

Xcode provides the required compiler. For Metal acceleration, either set CMAKE_ARGS="-DGGML_METAL=on" before installing or use the metal pre-built wheel index, which requires macOS 11.0 or later and Python 3.10 to 3.12.

How do I run the llama-cpp-python server?

The server is an optional extra that pulls uvicorn, fastapi and related packages from pyproject.toml. The README points to the server documentation for code completion, function calling, vision support and multiple models, and the repository includes an examples/server directory.

Official sources

  1. abetlen/llama-cpp-python on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/abetlen-llama-cpp-python.svg)](https://hysenlabs.com/projects/abetlen-llama-cpp-python)
Community notes

Community notes