ggml: a dependency-free C tensor library for machine learning
Tensor library for machine learning
At a glance
- What is it?
- ggml is the C/C++ tensor library underneath llama.cpp and whisper.cpp. It targets portability and quantized inference, not training, and the documentation is thinner than its adoption suggests.
- Who is it for?
- Adopt ggml if you are embedding quantized inference in a C or C++ program and want no runtime dependencies, a CMake build and kernels for x86, ARM and RISC-V. Do not adopt it if you need autograd, a training loop, a Python-first API or a documented support policy; ggml is a library, not a framework, and the README points contributors at llama.cpp for core changes.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What ggml is for, and who should pick it up
ggml is a tensor library written in plain C and C++. The README states its goal directly: a simple, portable and efficient tensor library for machine learning with minimal setup. The phrase that matters is minimal setup. There is no Python runtime, no package manager and, per the README, no dependencies at all. You clone the repository, run CMake and get a library you can link into a C or C++ program.
The audience is narrow and specific. If you are writing an inference engine, a desktop application that ships a model inside the binary, or an embedded service that cannot pull in a Python stack, ggml is aimed at you. The README lists x86, ARM, RISC-V, LoongArch, PowerPC, s390x and WebAssembly as targets, which is a wider platform list than most tensor runtimes publish. That breadth is the product. If your workload runs comfortably in PyTorch on a single Linux box, ggml solves a problem you do not have.
The clearest signal of who this is for is what it does not do. The README describes quantization, SIMD kernels and zero runtime allocations. It never mentions autograd, optimizers, training loops or a dataset abstraction. ggml is an inference-oriented tensor library. Treat any expectation of training as a misunderstanding of the scope.
How ggml works: graphs, no runtime allocation
The design constraint the README leads with is zero memory allocations during runtime. That single sentence explains most of the architecture. You build a computation graph up front, and the library executes it without calling the allocator along the way. For latency-sensitive or embedded work this removes a class of unpredictable pauses, and it means memory planning happens before the first token or pixel is processed.
The README lists broad backend support: CPU, GPU, NPU and browser. The examples directory shows what that looks like in practice. There are directories for gpt-2, gpt-j, mnist, sam, yolo, magika and perf-metal, plus a python example. The perf-metal example is the tell that backends are first-class rather than an afterthought bolted onto a CPU path.
Quantization is the other half of the mechanism. The README states support for 2- to 8-bit integer quantization plus MXFP4 and NVFP4 microscaling formats. Those are weight formats, and they exist so a model fits in memory that would otherwise be too small. The GGUF file format is documented in docs/gguf.md, and it is the container that carries quantized weights into the runtime. If you are choosing between ggml and a general tensor framework, the honest summary is that ggml trades ecosystem breadth for control over memory, binary size and numeric format.
Building ggml from source and running the simple example
There is no package to install. The README's quick start is the only supported path it documents: clone, configure with CMake, build in Release. The -j 8 flag sets parallelism and can be changed to match your machine.
git clone https://github.com/ggml-org/ggml
cd ggml
mkdir build && cd build
cmake ..
cmake --build . --config Release -j 8After the build completes you should have the library and the example binaries under the build directory. The README points at examples/simple for a minimal, fully commented example covering matrix multiplication. That is the right first target: it exercises the graph construction and execution path without dragging in a tokenizer or a model file. The README does not give a separate command for building only that target, so use the build output of the commands above and look for the simple example there.
Read examples/simple before writing your own graph. The repository also ships a pkg-config template at ggml.pc.in, which is how downstream C and C++ projects are expected to find the installed library rather than hardcoding include and link paths. The examples/python directory exists, but the README does not present Python bindings as the primary interface, and requirements.txt is dominated by PyTorch, TensorFlow and Keras entries used by the example scripts rather than by the core library.
Where ggml is the wrong tool
The most concrete limitation is the contribution path. The README asks that changes to the core ggml library, including the CMake build system, be opened as pull requests in llama.cpp rather than in this repository, on the grounds that doing so makes the PR more visible and better tested. Read that carefully. The canonical development surface for the library is a different repository. If you plan to patch ggml and maintain a fork, you are working against the grain of the project's own process.
Documentation is the second gap. The README links three things: docs/gguf.md, a Hugging Face blog post introducing ggml, and a wiki page of tips and tricks hosted on llama.cpp. There is no API reference in the README. A tensor library without a reference manual means reading headers in include/ and src/ to learn the surface. That is normal for C projects of this style, but it is a real cost you should price in before starting.
The third limitation is scope. There is no autograd, no optimizer and no training story anywhere in the README. If your problem is fine-tuning, ggml is not the tool, and no amount of quantization support changes that. Finally, the README does not document rollback, version compatibility guarantees or a deprecation policy across the v0.2x releases, so pinning a version and reading the release notes before upgrading is the only safe posture.
ggml against ONNX Runtime and PyTorch
The nearest alternative for portable inference is ONNX Runtime. The difference in approach is where the work happens. ONNX Runtime consumes a serialized graph in the ONNX format and provides a stable, versioned runtime API with prebuilt binaries for many platforms. ggml expects you to construct the computation graph in C or C++, link a library and manage the memory plan yourself. ONNX Runtime gives you a format boundary and a supported ABI; ggml gives you direct control and a much smaller dependency surface.
Against PyTorch the split is simpler. PyTorch is a training framework with autograd, a large Python ecosystem and eager execution. ggml is a C tensor library that the README describes as having no dependencies and zero runtime allocations, aimed at shipping inference. If you need to experiment, train and iterate, PyTorch wins on every axis that matters. If you need a binary that runs on RISC-V or inside a browser without a Python interpreter, PyTorch is not in the conversation.
The GGUF format is where ggml's ecosystem advantage lives. docs/gguf.md documents a container designed for quantized weights, and the related searches around GGML versus GGUF suggest the naming confuses people. The distinction is worth stating plainly: ggml is the library, GGUF is the file format it reads. Choosing ggml usually means choosing GGUF as your weight distribution format, and that decision is harder to reverse than the library choice.
Maintenance, releases and the MIT licence
The repository is not archived, and the last push was on 2026-09-09. Recent releases are v0.23.0 on 2026-09-04, v0.22.0 on 2026-08-25 and v0.21.0 on 2026-08-21. Three releases in roughly two weeks is a fast cadence, and the version numbers are still in the 0.x range. Fast 0.x releases mean the API can move; they also mean you should read release notes before bumping a pinned version in production.
The licence is MIT, stated in the README badge and present as a LICENSE file at the repository root. MIT is permissive: it allows commercial use, modification and redistribution with the copyright notice and licence text retained. That is a summary of the licence text, not legal advice. If you ship ggml inside a product, the obligation to verify is yours, and the file to read is LICENSE at the root of the repository.
Upgrade cost is the part the README does not answer. It does not describe a stability policy, a deprecation process or a compatibility matrix between library versions and GGUF files. The practical consequence is that upgrading ggml is an integration task, not a version bump: you rebuild, you re-run the examples, and you check that your graph construction still compiles against the current headers.
Editorial conclusion
Adopt ggml if you are embedding quantized inference in a C or C++ program and want no runtime dependencies, a CMake build and kernels for x86, ARM and RISC-V. Do not adopt it if you need autograd, a training loop, a Python-first API or a documented support policy; ggml is a library, not a framework, and the README points contributors at llama.cpp for core changes. Before committing, build the examples/simple target and read docs/gguf.md to confirm the GGUF format matches how you plan to ship weights.
Frequently asked questions
What language is whisper.cpp written in?
The README does not state the language of whisper.cpp. It describes ggml itself as a plain C/C++ implementation without any dependencies, and whisper.cpp is built on ggml, but the README makes no claim about whisper.cpp's own source language.
Can I download Llama.cpp for Windows?
The README does not describe downloading llama.cpp or any prebuilt Windows package. It only documents building ggml from source with git clone and CMake, and it lists Windows among the platforms the library targets.
What are llama.cpp models?
The README does not define llama.cpp models. It documents ggml as a tensor library and points to docs/gguf.md for the GGUF file format, which is the container ggml uses for quantized weights.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ggml-org-ggml)