Model or dataset
matthartman/ghost-pepper avatar
matthartman/ghost-pepper

Ghost Pepper: on-device speech-to-text and meeting transcription for macOS

100% private on-device voice models for speech-to-text and meeting transcription on macOS

3,192 stars188 forksSwiftLicense varies

At a glance

What is it?
Ghost Pepper is a free, MIT-licensed menu bar app that runs Whisper, Parakeet, Qwen3-ASR and Nemotron speech models locally on Apple Silicon. It is a good fit if you dictate into other apps all day and refuse to send audio to a cloud API. It is not a fit if you are on Intel Macs, on older macOS, or need a real-time streaming transcript.
Who is it for?
Adopt Ghost Pepper if you are on an M1-or-newer Mac running macOS 14.0 or later, you dictate into other apps rather than inside a browser tab, and you accept that the Accessibility permission is what makes the paste work. Do not adopt it if you are on Intel hardware, if you need a live streaming transcript while someone is still talking, or if your organization blocks Accessibility grants and has no MDM profile to pre-approve them.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 65 days ago.
What is it written in?
Mainly Swift, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Ghost Pepper solves, and who it is actually for

Most dictation tools on macOS work by shipping your audio to a server. That is fine for a grocery list and not fine for a client call, a medical note, or anything covered by an internal data-handling policy. Ghost Pepper takes the opposite position: the README states that no cloud APIs are used and no data leaves the machine, and the privacy audit table lists speech-to-text, cleanup, recording, meeting transcription, summary generation, OCR and file storage as local. The only cloud features named are optional and disabled by default: Zo AI chat, Trello integration, and Granola meeting import, each requiring your own API keys.

The intended user is a Mac owner who talks to their computer more than they type. The core interaction is a hold-to-talk hotkey: hold Control, speak, release, and the transcribed text is pasted into whatever text field has focus. That is a keyboard-replacement workflow, not a note-taking app. The second workflow is meeting transcription, which records a call and writes a markdown file containing notes, transcript and a locally generated summary. If neither of those describes your day, the app has little to offer you beyond novelty.

The project is explicit about its audience in the header: macOS 14.0+, Apple Silicon (M1+), free and open source. That is a narrower platform claim than most Mac utilities make, and it is worth taking literally.

How the model stack is wired together

Ghost Pepper is a Swift application that delegates inference to four separate open-source projects rather than implementing its own runtime. WhisperKit handles the Whisper family. FluidAudio handles Parakeet v3. MLX Audio handles the two Nemotron streaming models. LLM.swift runs the cleanup models, which are Qwen 3.5 in 0.8B, 2B and 4B sizes. Weights arrive from Hugging Face and are cached locally after the first download.

The interesting design decision is that speech recognition and text cleanup are two separate models with separate trade-offs. Whisper small.en is the default recognizer at roughly 466 MB; the default cleanup model is Qwen 3.5 0.8B at roughly 535 MB. Cleanup is what strips filler words and resolves self-corrections, and the README documents that the cleanup prompt is editable. That matters more than it sounds: if the cleanup model rewrites your words in a way you dislike, the prompt is the knob, not the model.

The model table is also where the real cost lives. Whisper tiny.en is about 75 MB and English-only. Whisper small.en is the accuracy default. Parakeet v3 covers 25 languages at roughly 1.4 GB. Qwen3-ASR 0.6B covers 50+ languages at roughly 900 MB but requires macOS 15 or later. The two Nemotron streaming models, at roughly 633 MB and 721 MB, are the low-latency options for English and multilingual respectively. Meeting transcription is described as chunked, and recordings use AVAudioEngine plus ScreenCaptureKit without streaming, which is consistent with a batch pipeline rather than a live one.

Installing Ghost Pepper and running your first dictation

The README points to a signed DMG rather than a package manager. Download it from the releases page, open the disk image, drag the app to Applications, then launch it. The first launch asks for Microphone and Accessibility permissions. Microphone is self-explanatory. Accessibility is what lets the app register a global hotkey and paste text by simulated keystrokes, so without it the hold-to-talk flow does not produce output in other apps.

On macOS Sequoia the README warns that Gatekeeper may show an "Apple could not verify" message the first time. The documented fix is to open System Settings, go to Privacy & Security, scroll to the Ghost Pepper entry and click Open Anyway, then Confirm in the popup. The README states this is a one-time step.

If you prefer to build it yourself, the instructions are short. Clone the repository, open the Xcode project, and run.

bash
git clone https://github.com/matthartman/ghost-pepper
cd ghost-pepper
open GhostPepper.xcodeproj

Then build and run with Cmd+R in Xcode. Expect the first launch to spend time downloading model weights, since the README says models download automatically and are cached locally.

Once permissions are granted, the usage pattern is a single gesture. Hold Control, speak, release. The README's own summary of the flow is that you "release to transcribe and paste into any text field." If nothing appears, check Accessibility first, then check which speech model is selected, because the default is English-only and a multilingual recording will produce poor results until you switch models in Settings. Launch at login is on by default on first run and can be turned off in Settings.

The Accessibility permission is the real deployment constraint

Granting Accessibility is not a minor onboarding step. It gives an application the ability to synthesize keystrokes into other applications, and on managed Macs it normally requires admin rights. Ghost Pepper's README acknowledges this directly and provides an MDM path: a Privacy Preferences Policy Control payload with the bundle identifier of the signed app, the signing team's Apple Developer team ID, and the Accessibility permission key `com.apple.security.accessibility`. Jamf, Kandji and Mosaic are named as examples of tools that can push such a profile.

That is a thoughtful inclusion, but it also tells you who will struggle. If you work on a locked-down corporate machine and your IT department will not pre-approve Accessibility for an unsigned-by-policy third-party app, the paste-into-any-field feature is dead on arrival. You could still use the meeting transcription side, which relies on Microphone and ScreenCaptureKit rather than keystroke simulation, but you would be using half the product.

The second constraint is that the privacy claims rest on a self-audit. The README describes the audit as "verified by AI code review" and invites you to run the audit prompt against the repository yourself. That is a reasonable posture for an open-source project, and the invitation to reproduce it is more than most apps offer. It is still not an independent third-party audit, and the README does not claim it is.

Where Ghost Pepper is the wrong tool

The hardware floor rules out a large group. Apple Silicon M1 or newer only, and macOS 14.0 or later. Intel Macs are not supported, and the README does not document a fallback path for them. If you are still on an Intel machine, this project is not for you regardless of how appealing the local-inference pitch is.

The model-specific requirement is easy to miss. Qwen3-ASR 0.6B, described as the highest multilingual quality option, requires macOS 15 or later. On macOS 14 you can still use Whisper small multilingual or Parakeet v3, but not the model the README ranks best for multilingual work.

Latency is the other honest limitation. The cleanup stage is an LLM pass, and the README's own speed column gives roughly 1 to 2 seconds for the default Qwen 3.5 0.8B, 4 to 5 seconds for the 2B, and 5 to 7 seconds for the 4B. That is a delay between releasing the Control key and seeing text appear. For dictating a paragraph it is tolerable. For a live captioning use case, or for anything where you need the words on screen while the speaker is still talking, the batch pipeline is the wrong architecture. The two Nemotron streaming models are the closest thing to a low-latency option, but the README frames them as streaming speech models, not as a live captioning product.

Storage is a quieter issue. The default combination of Whisper small.en and Qwen 3.5 0.8B is roughly 1 GB before you record anything. Switching to Qwen3-ASR plus a 4B cleanup model pushes past 3 GB of cached weights.

How it compares to cloud dictation and to MacWhisper-style tools

The obvious alternative is a cloud dictation service. The difference is not accuracy, it is the threat model. A cloud service transcribes on someone else's hardware and returns text; Ghost Pepper transcribes on your machine and never opens a socket for audio. The trade is that you own the model downloads, the disk usage, the permission grant and the model selection. Cloud services hide all four of those decisions from you, which is exactly why they are easier and exactly why they are unsuitable for some recordings.

The closer comparison is to other local Whisper front ends for macOS, which typically present a file-transcription interface: you drop in an audio file, wait, and get a transcript. Ghost Pepper's centre of gravity is different. It is a menu bar app whose primary verb is dictating into the focused text field, with meeting transcription as the second workflow that produces a markdown file with notes, transcript and a summary. If your job is transcribing existing audio files in bulk, a file-oriented tool will fit better than a hold-to-talk app. If your job is writing emails, tickets and messages by voice, the hotkey model is the point.

The other structural difference is the model menu. Ghost Pepper exposes Whisper, Parakeet, Qwen3-ASR and Nemotron behind one interface, with separate cleanup models, and lets you edit the cleanup prompt. That flexibility is real, but it also means you are the one choosing between an English-only 75 MB model and a 1.4 GB multilingual one. Tools that ship a single fixed model make that decision for you.

Maintenance, licensing and what an upgrade actually costs

The repository is not archived, and the last push was on 2026-07-27, which is recent enough that the project is not dormant. The release history shows v2.4.4 on 2026-07-27, v2.4.2 on 2026-07-18 and v2.4.1 on 2026-07-18, so the cadence in that window was tight. There is an appcast.xml at the repository root, which is consistent with Sparkle-based in-app updates; Sparkle is listed in the acknowledgments.

The licence is MIT. The README states it and the repository carries a LICENSE file. Practically, that means you can read, modify and redistribute the source, and there is no copyleft obligation on your own code. It does not mean the models are MIT: the weights come from Hugging Face and are produced by separate projects (WhisperKit, FluidAudio, MLX Audio, LLM.swift), each of which carries its own licence. If you plan to redistribute a build with models bundled, check those upstream licences separately. This is a description of what the repository states, not legal advice.

The upgrade cost is mostly disk and bandwidth, not engineering. Model weights are cached locally after download, and switching speech or cleanup models means fetching new weights from Hugging Face. There is no documented rollback procedure in the README, and no documented migration path for the on-disk markdown transcripts if you change storage layout. For a single-user Mac app that is a small concern. For a team deploying via MDM, it means version pinning is your responsibility.

Editorial conclusion

Adopt Ghost Pepper if you are on an M1-or-newer Mac running macOS 14.0 or later, you dictate into other apps rather than inside a browser tab, and you accept that the Accessibility permission is what makes the paste work. Do not adopt it if you are on Intel hardware, if you need a live streaming transcript while someone is still talking, or if your organization blocks Accessibility grants and has no MDM profile to pre-approve them. Before installing, open PRIVACY_AUDIT.md and read the file-level results rather than the summary table, check that your macOS version supports the model you want (Qwen3-ASR 0.6B requires macOS 15+), and budget the disk space: the default Whisper small.en plus the default Qwen 3.5 0.8B cleanup model is roughly 1 GB of cached weights, and the multilingual options push that past 2 GB.

Frequently asked questions

Does Ghost Pepper send my audio to the cloud?

The README states that no cloud APIs are used and no data leaves your machine, and the privacy audit table lists speech-to-text, text cleanup, audio recording and meeting transcription as local. The only cloud features named are Zo AI chat, Trello integration and Granola meeting import, all disabled by default and requiring your own API keys. Model downloads from Hugging Face are described as a one-time step.

Which Macs and macOS versions does Ghost Pepper support?

The README lists macOS 14.0+ and Apple Silicon (M1+) as the requirements, so Intel Macs are not supported. One model, Qwen3-ASR 0.6B, additionally requires macOS 15 or later. On macOS 14 you can still use the Whisper multilingual model or Parakeet v3.

Which speech model should I pick in Ghost Pepper?

The README marks Whisper small.en as the default at roughly 466 MB and describes it as the best accuracy option for English. Parakeet v3 covers 25 languages at roughly 1.4 GB, and Qwen3-ASR 0.6B covers 50+ languages at roughly 900 MB but needs macOS 15 or later. Whisper tiny.en is the smallest at roughly 75 MB and is English only.

Why does Ghost Pepper need Accessibility permission?

The permissions table gives two reasons: Microphone to record your voice, and Accessibility for the global hotkey and pasting via simulated keystrokes. On managed devices the README notes that granting Accessibility normally needs admin access, and documents an MDM Privacy Preferences Policy Control payload for pre-approval.

Official sources

  1. Issues
  2. matthartman/ghost-pepper on GitHub
  3. README
  4. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/matthartman-ghost-pepper.svg)](https://hysenlabs.com/projects/matthartman-ghost-pepper)