# Const-me/Whisper: DirectCompute Whisper inference for Windows

> A Windows-only port of whisper.cpp that runs OpenAI's Whisper ASR on Direct3D 11 compute shaders. It ships a GUI, a native C++ DLL, a C# wrapper and PowerShell bindings, and it trades portability for speed on the GPU you already own.

**Const-me/Whisper** — High-performance GPGPU inference of OpenAI's Whisper automatic speech recognition (ASR) model

- Repository: https://github.com/Const-me/Whisper
- Stars: 10,671 · Forks: 969
- Language: C++
- License: MPL-2.0
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/const-me-whisper

## What Const-me/Whisper solves, and for whom

OpenAI's Whisper is a Python model. Running it in a Windows desktop application normally means shipping a Python interpreter, PyTorch, and a CUDA or CPU backend. The README measures that stack at 9.63 gigabytes of runtime dependencies. Const-me/Whisper is a Windows port of whisper.cpp, which is itself a C++ port of OpenAI's Whisper, and it replaces all of that with a single native DLL and a Direct3D 11 compute-shader backend. The README puts the resulting Whisper.dll at 431 kilobytes.

The audience is narrow and clearly stated: developers building 64-bit Windows software who want speech recognition inside their own process. The supported platform is 64-bit Windows only, the library targets Windows 8.1 or newer (the author has only tested Windows 10), and it needs a Direct3D 11.0 capable GPU plus AVX1 and F16C on the CPU side. If you are writing a C# application, there is an idiomatic wrapper on NuGet, and version 1.10 added scripting support for PowerShell 5.1. If you are writing C++, you get a COM-style API. If you just want to transcribe a file, there is a GUI.

The project is not archived, and the last push was on 2026-05-24. That is a repository with recent activity, but the newest release listed is 1.12.0 from 2023-07-22, so the tagged release line is older than the branch.

## How the DirectCompute backend actually runs the model

The backend is DirectCompute, which the README glosses as compute shaders in Direct3D 11. That choice is the whole design. Instead of a vendor SDK such as CUDA, the model's matrix multiplications and other operations are compiled to HLSL shaders and dispatched on whatever GPU the machine has. The README calls this vendor-agnostic, and the practical consequence is that an AMD integrated APU and an Nvidia discrete card go through the same code path.

The shaders are not shipped as source. The build instructions include a CompressShaders C# project under Tools; running it prints a line such as Compressed 46 compute shaders, 123.5 kb -> 18.0 kb. That compressed form is what the DLL carries.

Precision is mixed F16 and F32. The README justifies this by pointing out that Windows has required support for R16_FLOAT buffers since Direct3D feature level 10.0, so the half-precision path is safe on essentially any GPU from that era onward. Audio handling is delegated to Media Foundation, which covers most audio and video formats and most Windows capture devices, with two documented exceptions: Ogg Vorbis is not supported, and professional capture devices that expose only ASIO are not supported.

For live capture there is voice activity detection, implemented from a 2009 paper by Mohammad Moattar and Mahdi Homayoonpoor. There is also a built-in profiler that measures the execution time of individual compute shaders, which is the main tool you have if a particular GPU turns out to be slow.

## Installing WhisperDesktop and transcribing a first file

The fastest path is the pre-built GUI. The README says to download WhisperDesktop.zip from the Releases section of the repository, unpack the ZIP, and run WhisperDesktop.exe. No installer, no runtime to configure.

On the first screen the application asks you to download a model. The README recommends ggml-medium.bin, which it lists at 1.42GB, because that is the model the author has mostly tested with.

```bash
# Not a shell command: download WhisperDesktop.zip from the repository's Releases section,
# unpack it, then run the executable from the unpacked folder.
WhisperDesktop.exe
```

After the model loads, the next screen transcribes an audio file. A third screen captures live audio from a microphone and either transcribes or translates it. Nothing in the README documents a command-line equivalent for the GUI workflow.

For developers who want the library rather than the application, the build path is Visual Studio. Clone the repository, open WhisperCpp.sln in Visual Studio 2022 (the author uses the freeware community edition, version 17.4.4), switch to the Release configuration, then build and run the CompressShaders project under Tools before building anything else.

```bash
# Inside Visual Studio 2022, after opening WhisperCpp.sln:
# 1. Set the configuration to Release
# 2. Right-click the CompressShaders project under Tools, choose "Set as startup project"
# 3. From the menu, choose Debug / Start Without Debugging
# Expected console output looks like:
# Compressed 46 compute shaders, 123.5 kb -> 18.0 kb
# 4. Build the Whisper project for the native DLL, or WhisperNet for the C# wrapper
```

One redistribution detail matters if you ship this. If your host application is built with Visual C++ 2022 or newer and you redistribute the Visual C++ runtime as a merge module or vc_redist.x64.exe, the README says to open the Whisper project properties, go to C/C++, Code Generation, and switch Runtime Library from Multi-threaded (/MT) to Multi-threaded DLL (/MD) before rebuilding. The binary gets smaller.

## Where the performance story is weaker than it looks

The headline benchmark in the README is a GeForce 1080Ti transcribing 3 minutes 24 seconds of speech with the medium model in 19 seconds, against 45 seconds for PyTorch with CUDA. That is a real number from the author's machine, and it is also the only configuration that gets a strong result.

The author is candid that the GPU selection at hand was limited. Optimization targets were the 1080Ti, the Radeon Vega 8 inside a Ryzen 7 5700G, and the Radeon Vega 7 inside a Ryzen 5 5600U. The 1080Ti reaches relative speed 5.8 on the large model and 10.6 on the medium model. The Ryzen 5 5600U APU reaches about 2.2 on the medium model, which the README describes as not great but still much faster than realtime. An Intel HD Graphics 4000 from 2012 managed 0.14 relative speed on the medium model and 0.44 on the small model, which is far slower than realtime.

The README goes further and says it is not sure performance is ideal on discrete AMD GPUs or integrated Intel GPUs, and that those might need different builds of the most expensive compute shaders, named as mulMatTiled.hlsl and mulMatByRowTiled.hlsl. So the vendor-agnostic claim is about correctness and coverage, not about equal tuning. If your target hardware is not one of the three cards above, treat the published numbers as an upper bound you have to verify yourself.

The platform constraint is the other hard edge. There is no Linux or macOS build, and the repository is explicit that 64-bit Windows is the only supported platform.

## How it compares with whisper.cpp

Whisper.cpp, by Georgi Gerganov, is the upstream project this one ports. The difference is the compute backend. Whisper.cpp is built around CPU inference with optional backends, and it runs on Linux, macOS and Windows. Const-me/Whisper commits to Direct3D 11 compute shaders and to Windows alone, and buys GPU acceleration on any D3D 11.0 class adapter as a result.

The practical split follows from that. If you need a cross-platform binary, a Linux server process, or a CPU-only deployment, whisper.cpp is the right starting point and this project is the wrong one. If you are already inside a Windows application and want the model to run on the user's existing GPU without a CUDA dependency, the DirectCompute route is the reason this fork exists at all. The repository keeps the whisper.cpp lineage visible in its naming: the solution file is WhisperCpp.sln.

A second comparison point is the Python original. PyTorch with CUDA is the reference implementation and the one the README benchmarks against, but it is also the one that brings the 9.63 gigabyte dependency tree. The trade here is clear: you give up cross-platform reach and the Python ecosystem, and you get a small native binary.

## Licence and the cost of keeping up

The repository is licensed under MPL-2.0. That is a file-level copyleft licence: modifications to files already covered by the licence stay under it, while larger works that combine this library with other code can be distributed under other terms. This is a summary of how the licence is usually described, not legal advice, and anyone redistributing a product around Whisper.dll should read the LICENSE file in the repository and get their own counsel.

On maintenance, the facts are limited. The repository is not archived, and the last push was on 2026-05-24, so the branch has moved recently. The newest release in the list is 1.12.0 from 2023-07-22, which means there is no tagged release covering whatever landed on master since then. A team adopting this should decide whether they are pinning to a release or tracking the branch, because those are different support stories.

Upgrade cost has one specific wrinkle that the README makes visible. The build is not a single command. You need Visual Studio 2022, and you must run the CompressShaders tool before building the library, because the DLL depends on the compressed shader output. Any CI pipeline has to reproduce that step. The README also notes that the repository carries a lot of development-only code: alternative model implementations, FP64 versions of some compute shaders, debug tracing and a trace comparison tool. Those are disabled by preprocessor macros or constexpr flags, and the author says he hopes it is fine to keep them there. For a reader auditing the codebase, that is a real amount of surface area to sort through.

## Conclusion

Adopt Const-me/Whisper if you are shipping a 64-bit Windows application and want Whisper inference inside your own process without dragging in PyTorch, CUDA or a Python runtime. The README puts the DLL at 431 kilobytes against 9.63 gigabytes of runtime dependencies for the PyTorch path, and the COM-style API plus the WhisperNet NuGet package make it embeddable from C# or from PowerShell 5.1. Do not adopt it if you need Linux, macOS, a headless server, or a GPU older than Direct3D 11.0, and do not expect the maintainer to have tuned it for your card: the README names only a 1080Ti, a Radeon Vega 8 and a Radeon Vega 7 as targets, and it explicitly says discrete AMD and integrated Intel GPUs have not been optimized. Before committing, verify three things on your own hardware: that your GPU and CPU meet the D3D 11.0, AVX1 and F16C requirements, that the medium model transcribes your audio faster than realtime (the 2012 Intel HD 4000 managed only 0.14 relative speed, which is far slower than realtime), and that your audio format is not Ogg Vorbis, which Media Foundation does not handle here.

## FAQ

### How do I install Const-me/Whisper on Windows?

Download WhisperDesktop.zip from the Releases section of the repository, unpack the ZIP, and run WhisperDesktop.exe. The first screen asks you to download a model, and the README recommends ggml-medium.bin at 1.42GB. Building the library yourself instead requires Visual Studio 2022 and running the CompressShaders tool before building the Whisper project.

### How do I use Const-me/Whisper to transcribe audio?

The GUI has a screen that transcribes an audio file and a separate screen that captures and transcribes or translates live audio from a microphone. Audio handling goes through Media Foundation, so most audio and video formats work, with Ogg Vorbis as a documented exception.

### Is OpenAI's Whisper free, and does that apply to Const-me/Whisper?

The README does not discuss pricing for OpenAI's Whisper. It does state that this repository is licensed under MPL-2.0, so the code here is distributed under that licence rather than as a paid product.

## Sources

- [Const-me/Whisper on GitHub](https://github.com/Const-me/Whisper)
- [Issues](https://github.com/Const-me/Whisper/issues)
- [License: MPL-2.0](https://github.com/Const-me/Whisper/blob/master/LICENSE)
- [README](https://github.com/Const-me/Whisper/blob/master/README.md)
- [Releases](https://github.com/Const-me/Whisper/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/const-me-whisper
