CapsWriter-Offline: Offline Speech Input for Windows, Reviewed
PC 端语音输入工具,离线识别,高准确率、低延迟,支持热词、LLM润色。按住CapsLock或鼠标侧键X2说话,松开自动上屏。
At a glance
- What is it?
- CapsWriter-Offline is a Windows-only, fully offline voice input tool that types while you hold CapsLock. It is fast and configurable, but the platform support is narrow and the LLM features need a separate local model.
- Who is it for?
- Adopt CapsWriter-Offline if you work on Windows 10 or 11, want dictation that never leaves the machine, and are willing to edit config_server.py and config_client.py by hand. Do not adopt it if you need macOS, Linux or mobile input, or if you expect a polished GUI installer that manages models for you.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 16 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What CapsWriter-Offline Is For
The README states the core promise in one line: hold CapsLock, speak, release, and the text appears. The project targets Windows 10 and 11 (64-bit) users who type a lot and want dictation that never touches a network. It is built for people on air-gapped or restricted machines, and the README notes that a USB drive is enough to carry it between computers.
The design goal is stated as low latency and high accuracy with heavy customization. That combination is why the project ships a C/S split: a server process runs the ASR model, a client process captures the keyboard and microphone. The README says an old Windows 7 machine cannot run the server models, but can still run the client for input. That is a real architectural concession, not a marketing line.
The feature list is broader than dictation. The README lists file transcription (drop an audio or video file on the client exe to get .srt, .txt and .json output), inverse text normalization that turns spoken numbers into digits, hot-word replacement, regex replacement, LLM roles, a tray menu, diary archiving and saved recordings. If you only need push-to-talk typing, most of that is optional weight.
How the Client and Server Split Works
The repository layout shows start_server.py and start_client.py at the top level, with a core/ directory and separate requirements-server.txt and requirements-client.txt files. The server owns the ASR engine and the model files under models/. The client owns the keyboard hook, the microphone capture and the tray icon.
The README describes the trigger as CapsLock or mouse side button X2, with a walkie-talkie mode and a click-to-record mode. The client sends audio to the server, the server returns text, and the client types it. Because the server holds the model, the client can stay light, which is the stated reason a weak machine can still act as an input terminal.
The README's model table makes the trade-off explicit. Paraformer and SenseVoice-Small are ONNX-only and fast but rated three stars for accuracy. Fun-ASR-Nano is ONNX plus GGUF and rated four stars. Qwen3-ASR is ONNX plus GGUF and rated five stars for accuracy but three stars for speed. The README gives a 20-second audio transcription reference: Paraformer 0.6s on CPU, SenseVoice-Small 0.6s on CPU and 0.15s on GPU, Fun-ASR-Nano 2.0s on CPU and 0.5s on GPU, Qwen3-ASR-1.7B 4.0s on CPU and 1.0s on GPU. Those are the project's own figures, not independent measurements.
The pyproject.toml pins requires-python to >=3.14 and lists sherpa-onnx, onnxruntime-directml, gguf, sounddevice, keyboard, pynput, pystray and pywin32 among the dependencies. The Windows-specific packages are a strong signal that porting is not a small job.
Installing CapsWriter-Offline and Typing Your First Sentence
The README's quick start does not begin with pip. It begins with the VC++ runtime, which the documentation lists as a prerequisite, and ffmpeg, which is only needed for file transcription and must be on PATH.
After that, download the software package from the latest release and the model archive from the models release, then extract the model into the matching folder under models/. The README is explicit that the model goes into the folder for the corresponding engine.
With the model in place, the two executables are started by double-clicking. The README says both minimize to the tray automatically.
start_server.exe
start_client.exeOnce both are running, hold CapsLock or mouse button X2 and speak. The README says the result is typed on release, with trailing commas and periods removed by default.
Configuration lives in two files at the repository root. The README says all settings are there and can be edited directly.
# config_server.py
# config_client.pyHot words go into hot.txt, and regex or simple equality rules go into hot-rule.txt. The README describes hot-word matching as phonetic fuzzy matching with a similarity threshold: above the threshold, the replacement is forced. That is worth knowing before you fill the file, because a low threshold will rewrite words you did not intend to change.
Where CapsWriter-Offline Falls Short
Platform support is the first limitation, and the README does not soften it. It states that only Windows 10 and 11 (64-bit) are guaranteed to work. Linux has no test environment and no packaging, so compatibility is not assured. macOS is blocked because the underlying keyboard library has dropped macOS support and the system restrictions are too many.
This also answers the Android and iOS search interest directly: there is no mobile build in the repository. The top-level entries contain no Android or iOS project files, and the README lists no mobile client.
The second limitation is operational. The README's own FAQ asks why pressing the key does nothing, and the answer is that the client's console window must still be running. To type into an application running with administrator rights, the client must also run as administrator. That is a Windows input-hook constraint, and it means the tool is not invisible to the operating system.
The third is accuracy on poor audio. The FAQ says that if the recognition result is empty, you should check the recording files under the year/month/assets folders, listen to them, and consider a desktop USB microphone. In other words, the model cannot compensate for a bad microphone, and the project's answer is hardware, not software.
The LLM roles are a fourth consideration. The README says a role is triggered when the start of a recognition result matches a role name, and that the result is then handled by that role. The role needs an LLM backend; the dependency list includes openai, ollama and httpx, which implies a local or remote model service is involved. The README does not document what happens when that backend is unreachable, so offline dictation and LLM polish are not the same guarantee.
How It Compares to Other Offline Dictation Tools
The README itself points to LazyTyper and 闪电说 as alternatives, noting that both have offline engines, support Windows, Linux and macOS, and have graphical interfaces. That is an honest comparison from the author, and it draws the boundary clearly: if you need cross-platform support or a GUI, those are the better fit.
The difference in approach is architectural. CapsWriter-Offline is a two-process system with a Python server that loads ONNX and GGUF models through sherpa-onnx, configured by editing Python files. The alternatives named in the README are presented as having polished graphical pages. Choosing CapsWriter-Offline means choosing a config-file workflow and a Windows-only input hook in exchange for control over the engine, the hot-word rules and the replacement pipeline.
For file transcription specifically, the README positions the client exe as a drop target that produces .srt, .txt and .json. A general-purpose transcription tool with a GUI will usually be easier for one-off files; CapsWriter-Offline's advantage is that the same engine and the same hot-word rules apply to both live dictation and batch files.
Maintenance, Licensing and Upgrade Cost
The repository is not archived. The last push was on 2026-08-26. The most recent release listed is v2.6 on 2026-05-30, described as a large set of detail improvements; v2.5 on 2026-05-09 added the Qwen3-ASR-1.7B model, and v2.4 on 2026-02-05 added DirectML GPU acceleration for the Fun-ASR-Nano encoder. The release cadence over that period is roughly every few months, with model support being the main theme.
Upgrading is not a package-manager operation. The README's install path is downloading a release archive and extracting model files into models/. That means an upgrade involves replacing the executables and, when the release notes mention a new model, downloading that model separately. The pyproject.toml lists pyinstaller under a build dependency group, which matches the repository's build.spec, build-client.spec and zip_release.py files: releases are built, not pip-installed.
The licence is MIT, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are included. Model files are downloaded separately from the software release, so their terms are a separate question from the MIT licence on the code. That is a distinction to check before shipping the tool inside a product; it is not legal advice.
The dependency list is large, including numba, pydub, rapidfuzz, nagisa and soynlp. For a source install rather than the packaged exe, that is a meaningful environment to maintain.
Editorial conclusion
Adopt CapsWriter-Offline if you work on Windows 10 or 11, want dictation that never leaves the machine, and are willing to edit config_server.py and config_client.py by hand. Do not adopt it if you need macOS, Linux or mobile input, or if you expect a polished GUI installer that manages models for you. Before committing, verify that VC++ redistributables are installed, that ffmpeg is on PATH if you want file transcription, and that the model folders match the engine names the config expects.
Frequently asked questions
Is there a CapsWriter-Offline APK for Android?
No. The README states that only Windows 10 and 11 (64-bit) are guaranteed to work, and it lists no Android client. The repository's top-level files contain no mobile project.
Can CapsWriter-Offline run on macOS or Linux?
The README says Linux has no test environment and no packaging, so compatibility is not guaranteed. For macOS it states that the underlying keyboard library has dropped macOS support and the system restrictions are too many, so it is not supported for now.
Why does nothing happen when I press CapsLock in CapsWriter-Offline?
The README's FAQ says to confirm that the start_client.exe console window is still running. It also notes that to type into a program running with administrator rights, the client must be run as administrator too.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/haujetzhao-capswriter-offline)