hear: a macOS speech recognition CLI for audio files and the microphone
Command line interface for the built-in speech recognition and transcription capabilities in macOS.
At a glance
- What is it?
- hear wraps the speech recognition macOS already ships with and exposes it as a command line tool for transcribing audio files and live input. It is small, BSD licensed, and limited by Apple's on-device locale support.
- Who is it for?
- Adopt hear if you are on macOS 13 or later, want transcription to stay on-device, and are comfortable with the locale limits Apple imposes on that mode. Do not adopt it if you need Linux or Windows, long unattended recordings, or a stable machine-readable output format, since the README documents none of those.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 31 days ago.
- What is it written in?
- Mainly Objective-C, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What hear does that macOS does not expose by itself
macOS has shipped speech recognition since 10.15 Catalina, but Apple provides no command line front end for it. The README frames hear as the counterpart to say, the long-standing speech synthesis tool: where say turns text into audio, hear turns audio into text. That gap is the whole reason the project exists. Anyone scripting transcription on a Mac otherwise has to write against the Speech framework directly.
The audience is narrow and concrete. People who want to transcribe a recording from a shell script, generate subtitles, or pipe microphone input into another tool. The README lists subtitles as a topic, and the file transcription example redirects output into a text file, which is the shape a subtitle or transcript pipeline wants. It is not a service, not a cross-platform library, and not a wrapper around a cloud API. It is a single binary that calls the recognition engine already present on the machine.
On-device versus server recognition, and why -d matters
The most consequential flag in hear is -d. The README states that without it, data may be sent to Apple servers, and that those servers have a hard limit of about 500 characters before quitting. With -d, recognition runs on the device and no audio leaves the machine. For anything longer than a sentence or two, or for audio you would rather not upload, -d is effectively mandatory.
The cost is coverage. The README says plainly that only locales supported by Siri work with the -d flag. That is a real constraint, not a footnote. If your language is not among them, on-device transcription is unavailable to you in this tool, and the server path caps you at roughly 500 characters per run. There is no third option documented.
This also means hear inherits whatever Apple ships. Model quality, language coverage and accuracy are Apple's, not the project's. The repository is a thin Objective-C layer over system frameworks, which is why the archive is around 50 KB. You are not installing a speech model. You are installing a way to call the one already on your Mac.
Installing hear from the release archive
The README points to a zip download for version 0.8, described as ARM/Intel 64-bit and requiring macOS 13 or later, signed with a Developer ID and notarized by Apple. After expanding the archive, the install script places the binary and the man page. The README says the defaults are /usr/local/bin for the binary and /usr/local/share/man/man1/ for the man page.
sudo bash install.shRunning that from inside the expanded directory should leave you with a hear binary on your PATH and a man page reachable through man hear. Because the script writes into /usr/local, it needs sudo, which is stated in the README's own example.
There is also a Homebrew route, and it carries a heavier prerequisite: the README says a full installation of Xcode is required to build hear via Homebrew. That is not the same as the command line tools. The tap and install commands are given as follows.
brew tap sveinbjornt/hear https://github.com/sveinbjornt/hear
brew install sveinbjornt/hear/hearIf you only want the binary, the zip is the lighter path. If you already have Xcode and prefer package management, Homebrew compiles it. If you are building from a clone instead, the README gives one command from the repository root, requiring Xcode command line build tools, and notes the resulting binary lands in products/.
make build_unsignedA note the README includes under troubleshooting: if running the binary produces an abort signal, opening it once from the Finder via right-click and Open should prompt for the permissions it needs and route it through your terminal client. That is a permissions problem, not a broken build.
A first transcription, from microphone and from a file
With hear installed, running it with no arguments starts recognition on the microphone or the default audio input device. The README's usage section shows the bare command for that case.
hearThere is a single-line output mode, which the README documents as -m. It is useful when you want the result on one line rather than as accumulating output.
hear -mFile transcription combines -d with -i and a path. The README's example redirects stdout into a text file, which tells you the transcript goes to standard output and nothing else is written for you.
hear -d -i /path/to/someone_speaking.wav > transcribed_text.txtAfter that command you should have a plain text file containing the recognized speech. The README says all formats supported by CoreAudio should work, listing WAV, MP3, AIFF, AAC, CAF and ALAC as examples. The word should is the README's, and it is worth taking literally: format support is inherited from the system audio stack, not enumerated and tested by the project.
The repository also ships a test script, run from the repository root, which the README presents as the way to test the built tool. The Makefile comments that more extensive tests than these cannot be run because of missing permissions from macOS, which is an honest admission about how much automated coverage is possible here.
bash test/test.shWhere hear stops being the right tool
The platform boundary is absolute. hear is a macOS tool built against macOS frameworks, and nothing in the repository suggests a port. On Linux or Windows it is not an option at all.
The 500 character ceiling on the server path is the sharpest functional limit. The README states it as a property of Apple's servers as of 2026, not as something hear can work around. So a long interview recording on the non-on-device path will not complete. You need -d, and -d needs a Siri-supported locale.
Output is another gap. The README shows transcripts going to stdout as text. There is no documented JSON mode, no timestamped segment output, no speaker diarization, and no confidence scores. If your pipeline needs word-level timing for subtitle synchronization, the README does not describe a way to get it. The subtitles topic suggests people use hear for that purpose, but the README does not document the mechanism.
Finally, there is no documented rollback or uninstall procedure. The install script copies files into /usr/local, and the README does not describe removing them. That is a small thing to plan around, but it is undocumented rather than solved.
hear compared with whisper.cpp and other local transcription tools
The obvious alternative for local transcription is an implementation of OpenAI's Whisper models, such as whisper.cpp. The difference in approach is fundamental rather than cosmetic. hear calls the speech recognition engine that is already part of macOS. It ships no model, and the download is about 50 KB. Whisper-based tools ship or download a model, which is orders of magnitude larger, and run inference themselves.
That trade-off cuts both ways. hear gets you transcription with no model management and no GPU or compute planning, because Apple's framework handles it. In exchange you are bound to Apple's locale list, Apple's accuracy, and Apple's decisions about what the on-device engine can do. A Whisper-based tool is not bound to a locale list in the same way and tends to expose more output structure, at the cost of a real model file and the compute to run it.
If your audio is in a language outside the Siri-supported set, or you need structured output, the Whisper route is the one that can actually deliver it. If you want a tiny binary that transcribes a WAV on a Mac with no setup beyond an install script, hear is the smaller answer.
Maintenance, licensing and what to check before depending on it
The repository is not archived, and the last push was on 2026-08-30. Release 0.8 arrived on 2026-04-29, preceded by 0.7 in 2025-11-09 and 0.6 in 2025-05-29. That is a steady cadence of roughly two releases a year, and the recent push date indicates the project is still being touched. The version string in the Makefile is 0.8, matching the latest release.
The licence is BSD-3-Clause, held by Sveinbjorn Thordarson with a copyright range of 2022 to 2026. The three clauses are the standard ones: keep the copyright notice in source redistributions, reproduce it in binary redistributions, and do not use the copyright holder's name or contributors' names to endorse derived products without permission. The licence text also carries the usual all-caps disclaimer of warranty. If you redistribute hear inside a product, the notice obligation is the part to look at, and that is a question for your own counsel rather than something this article can settle.
Upgrade cost is low by construction. The binary is around 50 KB and depends on system frameworks, so a new release is a download and a reinstall rather than a migration. The README does not document a versioned upgrade path or a rollback, so treat each install as a replacement. The one thing that can shift under you without a hear release is Apple's recognition engine itself, since hear calls into it. A macOS update can change behavior with no change to this project's version number.
Editorial conclusion
Adopt hear if you are on macOS 13 or later, want transcription to stay on-device, and are comfortable with the locale limits Apple imposes on that mode. Do not adopt it if you need Linux or Windows, long unattended recordings, or a stable machine-readable output format, since the README documents none of those. Before relying on it, verify that your target locale works with the -d flag and test the exact audio formats you plan to feed it, because the README states only that CoreAudio-supported formats should work.
Frequently asked questions
Does hear send my audio to Apple's servers?
By default it may, because the README states that without the -d flag data may be sent to Apple servers. Passing -d restricts recognition to on-device capabilities so the audio stays local.
Which languages does hear support for on-device transcription?
The README says only locales supported by Siri work with the -d flag. It does not list them, so you have to check whether your target locale is among them.
What audio formats can hear transcribe?
The README states that all formats supported by CoreAudio should work, giving WAV, MP3, AIFF, AAC, CAF and ALAC as examples. That support comes from the system audio stack rather than from hear itself.
Why does hear abort with a signal when I run it?
The README's troubleshooting section attributes this to permissions. It suggests opening the binary once from the Finder with right-click and Open, which prompts for the required permissions and launches it in your terminal client.
Can I install hear without Xcode?
Yes. The README offers a zip download for version 0.8 that installs via install.sh. The Homebrew route is the one that requires a full Xcode installation, according to the README.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/sveinbjornt-hear)