# How to install whisper.cpp and choose a model that fits your memory

> whisper.cpp runs OpenAI's Whisper speech recognition model from C/C++ with no library dependencies, and the practical decisions are all in the model download and the input format, not in the framework choice. The RAM table is the part worth reading twice.

**ggml-org/whisper.cpp** — whisper.cpp is a dependency-free C/C++ port of OpenAI's Whisper speech-recognition model, optimized for Apple Silicon and supporting CPU-only inference on many platforms.

- Repository: https://github.com/ggml-org/whisper.cpp
- Stars: 54,014 · Forks: 6,201
- Language: C++
- License: MIT
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/ggml-org-whisper-cpp

## The model is a separate download in ggml format, and memory is the constraint

No weights ship with the code. `sh ./models/download-ggml-model.sh base.en` fetches one Whisper model converted to the `ggml` format, and the conversion itself belongs to another project. The README does price the result, and the price is memory, not disk. `tiny` is 75 MiB on disk and about 273 MB in memory, `base` is 142 MiB and about 388 MB, `small` is 466 MiB and about 852 MB, `medium` is 1.5 GiB and about 2.1 GB, and `large` is 2.9 GiB with about 3.9 GB resident. Resident memory runs several times the file size in every row, so a download that looks harmless on disk is the wrong number to plan against. The Makefile also offers `large-v1`, `large-v2`, `large-v3` and `large-v3-turbo` targets, and the table prices none of them. If you are sizing hardware, treat those four as unknown rather than assuming the `large` row covers them.

## whisper-cli reads 16-bit WAV only, so ffmpeg sits in front of the tool

The whole quick start is five commands, and none of them install a runtime dependency into your project.

```bash
git clone https://github.com/ggml-org/whisper.cpp.git
cd whisper.cpp
sh ./models/download-ggml-model.sh base.en
cmake -B build
cmake --build build -j --config Release
./build/bin/whisper-cli -f samples/jfk.wav
```

That produces `./build/bin/whisper-cli`, and the flag to read its own options is `-h`. What the CLI will not do is read an mp3 or a flac. The README is explicit that the example runs only with 16-bit WAV files and asks you to convert the input first, with a resample to 16 kHz mono:

```bash
ffmpeg -i input.mp3 -ar 16000 -ac 1 -c:a pcm_s16le output.wav
```

The claim of no dependencies holds for the library, not for your pipeline. A product that hands users mp3 and m4a will need ffmpeg or the transcode path in `examples/ffmpeg-transcode.cpp` in front of the binary, and that is a dependency decision the repository makes for you by leaving the file unhandled.

## A Makefile target named after a model also builds the project and runs the samples

`make base.en` is a demo, not a weights fetch. The target downloads the model in `ggml` format, configures and builds with cmake, then runs inference on every `.wav` file in `samples`, which is why the README offers it as a quick demo rather than a setup step. If all you want is the file on disk, calling `sh ./models/download-ggml-model.sh` with the model name directly is the narrower command. The same applies to `make -j samples`: that target pulls several audio files, including material hosted on Wikimedia, on archive.org and on OpenAI's own CDN, and converts them to 16-bit WAV with ffmpeg. Both targets therefore assume a working network path, a build toolchain and, for the samples target, ffmpeg on the machine. In a locked down build environment where outbound fetches are restricted, these convenience targets are the first thing to replace with your own asset pipeline.

## Integer quantization is a second binary and a second file, with no prices to compare

Quantization is not a flag on the download. A separate `quantize` binary reads a model and writes a new one, and you then point the CLI at the result by path:

```bash
./build/bin/quantize models/ggml-base.en.bin models/ggml-base.en-q5_0.bin q5_0
./build/bin/whisper-cli -m models/ggml-base.en-q5_0.bin ./samples/gb0.wav
```

The method name is an argument, so the family of quantized formats is as wide as whatever `quantize` accepts, and the file you get back is named by whatever you pass as the second argument. Here is where the documentation thins out. The README says quantized models need less memory and disk and can process more efficiently depending on the hardware, then stops: there is no per-method table of output size, no accuracy comparison and no speed figure for `q5_0` against the full precision file. Choosing a quantization level is therefore a judgement call you will have to measure yourself on your own audio. The upside is real, since halving the resident footprint is the only lever the repository offers against the memory table, but the repository does not hand you the numbers to pick a point on that curve.

## Core ML moves only the encoder, and older macOS is flagged for hallucination

On Apple Silicon the encoder can run on the Apple Neural Engine through Core ML, which the README puts at more than x3 faster than CPU-only execution. Two things narrow that. The acceleration is scoped to the encoder inference, so the rest of the pipeline stays where it was, and getting it to work is a build step rather than a switch: the Core ML model has to be generated first, with Python dependencies installed and the Xcode command line tools present.

```bash
pip install ane_transformers
pip install openai-whisper
pip install coremltools
```

Python 3.11 is the recommended version, and the README recommends macOS Sonoma or newer for the reason that older systems might experience transcription hallucination. That caveat is the one to carry into production. Hallucination in a speech recognizer means output text that was never spoken, and here it is tied to an operating system version rather than to a flag, so the same binary and the same model behave differently on two Macs. If the transcript feeds anything downstream, test the machine you will actually ship on. `cmake -B build -DGGML_BLAS=1` is the equivalent knob for POWER, where the README reports faster than realtime transcription on POWER9 and POWER10 with a BLAS package installed.

## The model sits in one header and one source file, which is also what you vendor

The architecture is unusually compressible, and that is the reason integration is easy. The README states that the entire high-level implementation of the model is contained in `whisper.h` and `whisper.cpp`, with the rest of the code belonging to the `ggml` machine learning library. The public surface is a C-style API, not a C++ one, which means the binding you write is a header include and a set of function calls, and it is why there is a Java binding, a Node addon under `examples/addon.node`, an Objective-C example for iOS, an Android example and a WebAssembly build under `examples/whisper.wasm`. Each of those is a separate integration path, so the supported platform list in the README is not a claim that one binary runs everywhere. The alternative worth naming is the Python original this repository ports: same model, different runtime. What changes here is where the inference runs and how you call it, and if your requirement is alignment or word level timestamps, this repository does not present that work as its own.

## Two tag schemes shipped on the same day, so a version pin is not one value

The default branch is `master`, the licence is MIT, and the last push was on 2026-09-28. The tags are where a consumer has to look twice: v1.9.4 and b5130 both landed on 2026-09-11, and b5127 on 2026-09-10. A semantic version and a bare build number ship in the same window, so `latest` and a pinned tag can point at artefacts that were produced minutes apart by different naming conventions. If you vendor this, pin the exact tag you tested and record the model file alongside it, because a change in the runtime and a change in the `ggml` format model fail in different ways. Build configuration is spread as well: `CMakePresets.json` at the root, a `ci/` directory, a `.devops/` directory and a separate `README_sycl.md` for the SYCL build. Finding the flags for a given accelerator means searching those files rather than reading the main README, which covers a subset of the hardware paths.

## Conclusion

Use whisper.cpp when transcription has to run on the device, in a process you can embed through the C-style API in include/whisper.h, on a CPU or an accelerator the project already supports. Skip it if you need the Python ecosystem around Whisper, or if a machine with roughly 4 GB of free memory is not available, because the largest model in the README table asks for about 3.9 GB before your application takes any. Before you commit, check three things: the model file you intend to ship in ggml format, the fact that whisper-cli refuses anything but 16-bit WAV, and the macOS version if you plan to use Core ML, where the README warns about transcription hallucination on systems older than Sonoma.

## FAQ

### What is Whisper in C++?

It is the `whisper.cpp` port of OpenAI's Whisper automatic speech recognition model, written in plain C/C++ without dependencies. The high-level model implementation is contained in `whisper.h` and `whisper.cpp`, and the rest belongs to the `ggml` library.

### How to use Whisper cpp?

Clone the repository, download a model in `ggml` format with `sh ./models/download-ggml-model.sh base.en`, build with `cmake -B build` and `cmake --build build -j --config Release`, then run `./build/bin/whisper-cli -f samples/jfk.wav`. Detailed options are listed by `./build/bin/whisper-cli -h`.

### Can Whisper.cpp perform real-time transcription?

The README reports faster than realtime transcription on POWER9 and POWER10 on Linux with the `GGML_BLAS` build option and a BLAS package installed, and more than x3 faster than CPU-only execution for the encoder on Apple Silicon through Core ML. It gives no real time figures for other platforms.

### how to install whisper.cpp

The install is a cmake build rather than a package: `cmake -B build` followed by `cmake --build build -j --config Release`, which produces the example binaries under `./build/bin/`. A model download is a separate step, and `make base.en` runs the download, the build and the sample transcription in one target.

### what is whisper.cpp used for

Offline, on-device speech recognition, including an offline voice assistant built on the `examples/command` example, and embedding through the C-style API on iOS, Android, Java, WebAssembly, Windows, Raspberry Pi and Docker.

## Sources

- [Official README](https://github.com/ggml-org/whisper.cpp#readme)
- [Project repository](https://github.com/ggml-org/whisper.cpp)
- [Release notes](https://github.com/ggml-org/whisper.cpp/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ggml-org-whisper-cpp
