SayIt inserts polished dictation at your cursor, with seven local GGUF models
Open-source voice typing for Windows — a Wispr Flow / Superwhisper alternative. Press a shortcut, speak, and AI-polished text lands at your cursor. Local models, your own API keys, or a self-hosted backend.
At a glance
- What is it?
- A Windows voice typing app built as a Tauri and React client over a FastAPI backend, offering local GGUF recognition, direct cloud API keys, or a self-hosted server. The interesting part is that a selection changes what your speech means, and that the local mode's privacy promise depends on whether AI cleanup is on.
- Who is it for?
- Adopt SayIt if you dictate into editors and terminals all day on Windows and want a choice of who transcribes you, since local GGUF models, direct cloud keys and a self-hosted server are three genuinely different trust positions. Skip it if you are on macOS or Linux, because the installer and the client toolchain are Windows shaped, and read the licence before you plan to ship it inside a commercial product, since AGPL-3.0 covers the network use case.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Three modes, and the local one only stays local while cleanup is off
The mode table is where the trust story is decided, and one cell in it carries more weight than the others. Local mode keeps speech recognition on the PC, and the privacy claim is stated conditionally: with AI cleanup off, nothing leaves the device. Switch cleanup on and you are sending text to whichever provider you configured, which is a different arrangement from the one the mode name suggests.
Cloud API mode is the middle option and the easiest to reason about, because your PC talks directly to the ASR and AI providers you configure, with no SayIt server in the path at all. Server mode is for teams and managed deployments, where audio goes to a backend you control or to the public trial server. The public trial server is what the quick start points at, which is a sensible default for a first run and a poor default for anything you are typing a password-adjacent secret into.
The app shows which mode is active and where audio and text are processed, rather than leaving you to remember what you configured.
Seven GGUF models locally, and a long list of providers in the cloud
Local recognition is not a single bundled model. Seven GGUF models ship, and the voice engine page downloads and switches between them, using detected GPUs automatically. Parakeet Unified EN is described as fastest and most accurate for English, SenseVoice Small and Fun-ASR Nano cover the small-model end, Nemotron 3.5 ASR is the one named for breadth at 32 languages, and the remaining three are Qwen3-ASR sizes, which is how you trade memory for accuracy.
Cloud recognition names Doubao, Qwen, OpenAI, Gemini, OpenRouter, Xiaomi MiMo and Groq Whisper, plus any OpenAI-compatible service at an address you supply. AI cleanup runs on a partly different list: DeepSeek, Qwen, Doubao, Zhipu GLM, MiMo, Groq and Ollama, again with an address for anything OpenAI-compatible. Ollama on the cleanup side is the interesting entry, since a local model there is a second way to keep text on the machine without giving up the cleanup pass.
Each key field has a link to the provider's own console beside it, which saves the round of hunting for an API key page.
Selecting text turns your speech into an instruction instead of a transcript
Dictation is one mode. Selection is another, and it is the more interesting mechanism. With text selected first, what you say becomes an editing instruction rather than content: translate it, tighten it, rewrite it, or ask a question about it, and the result replaces the selection directly. That turns a speech tool into a batch of common editing operations that would otherwise need a separate prompt and a copy and paste.
Context-aware writing is a separate switch and it ships off by default. When on, it reads the text around your cursor so new dictation matches the surrounding tone and terminology. Off by default is the right call, because the feature works by looking at what else is on your screen.
One guard is worth naming explicitly: password fields are skipped. An app that reads around the cursor and inserts text at the focus would otherwise be a bad thing to alt-tab away from during a login, and skipping those fields is the specific mitigation for it.
A fixed cost dominates short clips, and the numbers come from one AWS instance
The performance reference is specific enough to be useful and narrow enough to be read carefully. It measures Qwen3-ASR-1.7B running under vLLM on an AWS EC2 `g5.xlarge`, which is an A10G with 24 GB of VRAM, and reports both latency and real-time factor.
| Audio length | ASR latency | RTF | | --- | --- | --- | | 30 seconds | ~0.8 s | 0.025 | | 1 minute | ~1.6 s | 0.026 | | 2 minutes | ~2.1 s | 0.017 | | 3 minutes | ~2.5 s | 0.014 | | 5 minutes | ~3.0 s | 0.010 |
The shape of the table is the finding. Latency grows from 0.8 seconds to 3.0 seconds while audio grows from 30 seconds to 5 minutes, a tenfold increase in input for under four times the wait, so roughly half a second is fixed overhead and the marginal cost is well under one percent of the audio duration. In practice that means a one-line dictation pays the whole fixed cost, and a paragraph barely pays anything extra.
What the table does not tell you is how any of this behaves on a laptop GPU, which is the configuration most people would run Local mode on, and it says so by being about a rented server.
Self-hosting is a compose file, and the default server model wants 16 GB
The backend is FastAPI with WebSocket streaming, Qwen3-ASR for recognition, and an optional OpenAI-compatible model for cleanup. Docker Compose is the recommended deployment path, and the setup is four commands:
git clone https://github.com/crosswk/SayIt.git
cd SayIt/server
cp config.example.yaml config.yaml
cp .env.example .env
# Add your provider and deployment settings to .env/config.yaml
docker compose up -d --buildBoth a YAML config and an env file are copied, so there are two places to put settings and the guide in `server/README.md` is where the split is explained. That same guide is stated to cover configuration, deployment, security and API details, and for a service that receives audio and returns text over a WebSocket, the security section is the one to read before exposing it beyond your own network.
The hardware floor is specific: GPU speech recognition requires an NVIDIA GPU, and 16 GB or more of VRAM is recommended for the default server model. Server mode is therefore not a cheap way to add local recognition to a small box.
Building the client needs Rust, CMake and the Vulkan SDK, and about 20 minutes
The desktop client is Tauri plus React, and development builds run through the Tauri CLI:
cd client
npm install
npm run tauri devThe toolchain behind that is Node.js 18 or newer, Rust 1.75 or newer, CMake 3.20 or newer, and the Vulkan SDK. The cost is stated plainly: the first native build compiles the C++ speech engine and takes around 20 minutes, and later builds use the cache. Twenty minutes is a one-time price for a native speech component, and it is worth knowing before you decide a build has hung.
There is a Windows-specific trap too. On a non-English Windows installation you should set `CL=/utf-8` before building, so MSVC reads the UTF-8 source files correctly. That is the kind of detail that turns into an hour of misread strings if nobody wrote it down.
For the server alone the loop is shorter, a virtual environment, `pip install -r backend/requirements.txt`, then `uvicorn app.main:app --port 8000`, needing Python 3.10 or newer and an NVIDIA GPU with CUDA for GPU inference.
AGPL-3.0, a 0.2.x line, and a community that runs on WeChat
Two facts about the project rather than the software are worth a reader's attention.
The first is the licence. AGPL-3.0 is the copyleft that reaches past distribution to network use, and for an application people might be tempted to wrap inside another product, that is the clause to read before you build on it. It is a deliberate choice for a project positioned against commercial alternatives, and it is the price of the code being available.
The second is where the conversation happens. Release announcements go out through a WeChat official account and a user group, and the README says plainly that both are Chinese-language channels and points English speakers at a GitHub issue instead. Anyone outside that ecosystem is working without the informal support network, which is a real difference between this and an equivalent project run in English.
The repository layout is four directories, `client/` for the Tauri and React desktop app, `server/` for the FastAPI backend, gateway, web demo and deployment files, `docs/` for user guides, and `dev-docs/` for internal notes. Releases are still on the 0.2 line, v0.2.0 on 2026-09-09, v0.2.1 on 2026-09-21 and v0.2.2 on 2026-09-24, with the last push to the repository on 2026-09-30.
Editorial conclusion
Adopt SayIt if you dictate into editors and terminals all day on Windows and want a choice of who transcribes you, since local GGUF models, direct cloud keys and a self-hosted server are three genuinely different trust positions. Skip it if you are on macOS or Linux, because the installer and the client toolchain are Windows shaped, and read the licence before you plan to ship it inside a commercial product, since AGPL-3.0 covers the network use case. Verify first that the mode you pick matches your privacy claim: in Local mode the guarantee is that nothing leaves the device with AI cleanup off, and turning cleanup on changes that. Licence is AGPL-3.0, the current release is v0.2.2 from 2026-09-24, and the last push was 2026-09-30.
Frequently asked questions
What is SayIt?
It is an open-source voice typing app for Windows. You press a shortcut, speak, and the transcribed and optionally cleaned-up text is inserted wherever your cursor is, in editors, chat apps, browsers and other Windows software.
Does SayIt work without sending my audio anywhere?
In Local mode, yes, with a condition: speech recognition stays on the PC and nothing leaves the device while AI cleanup is off. Turning cleanup on sends text to the provider you configured. The app shows which mode is active and where audio and text are processed.
Which speech recognition models does SayIt offer locally?
Seven GGUF models with GPU acceleration where available: Parakeet Unified EN, described as fastest and most accurate for English, SenseVoice Small, Fun-ASR Nano, Nemotron 3.5 ASR for 32 languages, and three Qwen3-ASR sizes. Detected GPUs are used automatically.
How do I self-host the SayIt backend?
Clone the repository, copy `config.example.yaml` to `config.yaml` and `.env.example` to `.env` from `server/`, add your provider settings, then run `docker compose up -d --build`. The backend is FastAPI with WebSocket streaming and Qwen3-ASR, and GPU recognition needs an NVIDIA GPU with 16 GB or more of VRAM recommended.
Can SayIt edit text I select?
Yes. Select text first and your speech becomes an editing instruction, so you can ask it to translate, tighten, rewrite or answer a question about the selection, and it replaces the selection directly. Context-aware writing, which reads text around the cursor to match tone, is off by default, and password fields are skipped.
What licence is SayIt released under?
AGPL-3.0. The current release is v0.2.2 from 2026-09-24, and the last push to the repository was 2026-09-30.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/crosswk-sayit)