Dragonfly: Python Speech Recognition Framework for Dragon, Kaldi, and WSR
Speech recognition framework allowing powerful Python-based scripting and extension of Dragon NaturallySpeaking (DNS), Windows Speech Recognition (WSR), Kaldi and CMU Pocket Sphinx
At a glance
- What is it?
- Dragonfly is a Python framework for writing custom voice commands that work across Dragon NaturallySpeaking, Windows Speech Recognition, Kaldi, and CMU Pocket Sphinx. It treats speech grammars as first-class Python objects, which lets developers define, load, and modify command sets in ordinary Python code without modifying any underlying speech engine.
- Who is it for?
- Dragonfly suits Python developers who need voice control over their desktop, want to script commands for Dragon NaturallySpeaking or Kaldi without learning a proprietary macro language, or need cross-engine portability in a single grammar definition. It is not a speech recognition engine itself: you still need Dragon, WSR, Kaldi, or CMU Pocket Sphinx installed and configured.
- Can I use it commercially?
- Yes, with conditions. LGPL-3.0 is a weak copyleft licence: you can use it inside commercial and closed-source software, but if you distribute changes to its own files, you must publish those changes under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 54 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Dragonfly Solves and Who It Is For
Writing custom voice commands for commercial speech recognition software like Dragon NaturallySpeaking traditionally requires using Dragon's own scripting tools or third-party utilities tightly coupled to that one engine. Dragonfly replaces that by providing a Python API where grammars, rules, and actions are Python objects. The same grammar code can run against Dragon, Windows Speech Recognition, Kaldi, or CMU Pocket Sphinx by switching the backend.
This is a fork of the original t4ngo/dragonfly project and is published on PyPI as dragonfly2. Dragonfly is designed for three broad use cases: general programming by voice in any programming language, speech-enabling GUI applications using context-specific command sets, and dictating prose. The action framework handles keystroke simulation and text input, and it works on Windows, macOS, and Linux (X11 only for Linux).
The target audience is developers and power users who are comfortable with Python and want to extend or automate their desktop through voice, rather than users looking for a turn-key dictation product.
Supported Speech Engines and Platform Constraints
Dragonfly supports four speech recognition backends. Dragon (the product line from Nuance) covers all versions up to the current version 16, including Home and Professional Individual editions; other editions may work but are not explicitly listed as tested. Windows Speech Recognition is included with Windows Vista, Windows 7 and later, and is freely available for Windows XP.
Kaldi is the open-source backend. It is released under the AGPL and runs on multiple platforms, making it the primary option for Linux and macOS users who do not have a Dragon license. CMU Pocket Sphinx is a second open-source option.
The action framework is cross-platform for Windows, macOS, and Linux, but on Linux only X11 sessions are fully supported. The README does not describe Wayland support. This is a meaningful constraint for developers on modern Linux distributions that default to Wayland.
Installing Dragonfly and Writing a First Command
Dragonfly installs as a Python package from PyPI:
pip install dragonfly2The simplest grammar defines a single spoken command and a callback. A CompoundRule combines the spoken form with the recognition handler:
from dragonfly import Grammar, CompoundRule
class ExampleRule(CompoundRule):
spec = "do something computer"
def _process_recognition(self, node, extras):
print("Voice command spoken.")
grammar = Grammar("example grammar")
grammar.add_rule(ExampleRule())
grammar.load()Save this in a command module file, load it via the Natlink user directory or a module loader, and say the phrase "do something computer". When Dragon recognizes it, the callback prints its message to the Natlink messages window. With other engines it goes to the console.
For Linux deployments using Kaldi or X11 features, the README notes that the wmctrl, xdotool, and xsel system packages must also be installed.
MappingRule: Defining Multiple Context-Specific Commands
The more common pattern in real voice grammars is MappingRule, which maps multiple spoken phrases to actions in a single class. An AppContext limits the grammar to a specific window. The README shows a Notepad example:
from dragonfly import (Grammar, AppContext, MappingRule, Dictation, Key, Text)
class NotepadRule(MappingRule):
mapping = {
"save [file]": Key("c-s"),
"save [file] as": Key("a-f, a/20"),
"save [file] as <text>": Key("a-f, a/20") + Text("%(text)s"),
"find <text>": Key("c-f/20") + Text("%(text)s\n"),
}
extras = [
Dictation("text")
]
context = AppContext(executable="notepad")
grammar = Grammar("Notepad example", context=context)
grammar.add_rule(NotepadRule())
grammar.load()The optional text element in the spec captures arbitrary dictated text, which Key and Text actions then use. The grammar is active only when a window matching the Notepad executable is in the foreground. This pattern scales: you define one MappingRule per application or mode, each with its own context, and Dragonfly loads and unloads the active ones as focus changes.
Architecture: Grammars, Actions, and the Module Loader
Dragonfly separates the grammar layer from the action layer. A Grammar holds Rules; a Rule defines a spoken form with optional dynamic elements like Dictation or Choice; an Action performs the keyboard or text operation when a rule matches.
The action sub-package is cross-platform and independent of which speech engine is in use. This means keystrokes and text actions work on macOS or Linux even when the backend speech engine is different.
The framework relies on a module loader or the Natlink user directory to discover and activate command modules at startup. The documentation at dragonfly.readthedocs.org covers the loader setup in detail, but it is not reproduced in the repository itself. This is a design point worth noting: the README does not contain a self-contained tutorial for getting the loader running. New users need to follow the online documentation to connect everything.
Limitations: No Engine, X11 Only on Linux, and Grammar Loading Complexity
Dragonfly is a framework, not a speech recognition engine. It provides no acoustic model and no audio capture. Every supported backend except Kaldi requires a separately installed and licensed speech engine. Dragon NaturallySpeaking requires a paid license. Windows Speech Recognition requires Windows. Only the Kaldi and CMU Pocket Sphinx backends are fully free and cross-platform, and both require their own setup before Dragonfly can use them.
On Linux, the action framework works only under X11. If you run a Wayland session and need keystroke simulation or text insertion to work, Dragonfly is not a direct fit.
Loading grammars requires a module loader or integration with Dragon's Natlink layer. The overhead for someone who is not already a Natlink or Dragon user is non-trivial. The README points to online documentation rather than providing a step-by-step setup guide in the repository.
Dragonfly Compared to an End-to-End Speech Control Framework
A direct alternative for users who want voice automation without a commercial engine is Coqui STT (the continuation of Mozilla DeepSpeech). Coqui provides an end-to-end neural speech recognition model that runs locally and is entirely open-source. It does not require a commercial license or Windows.
The difference in approach is fundamental. Dragonfly is grammar-based: you define exactly which phrases the engine should recognize, and the engine matches against those. This gives high accuracy on the commands you define but requires explicit authoring of every command. Coqui uses a neural model that transcribes arbitrary speech, which gives broader coverage but can introduce recognition errors on unusual vocabulary or commands that are not in its training distribution.
Dragonfly works best when you have a defined set of commands and want precise, low-latency recognition. Coqui or similar end-to-end systems work better when you need open-domain transcription or are not willing to enumerate every valid utterance.
Editorial conclusion
Dragonfly suits Python developers who need voice control over their desktop, want to script commands for Dragon NaturallySpeaking or Kaldi without learning a proprietary macro language, or need cross-engine portability in a single grammar definition. It is not a speech recognition engine itself: you still need Dragon, WSR, Kaldi, or CMU Pocket Sphinx installed and configured. On Linux, only X11 is fully supported; Wayland support is absent. The last push was on 2026-08-07, confirming active maintenance.
Frequently asked questions
Does Dragonfly work without Dragon NaturallySpeaking?
Yes. Dragonfly supports Kaldi and CMU Pocket Sphinx as open-source backends that do not require a Dragon license. The Kaldi backend is multi-platform and is the primary option for Linux and macOS users.
What is the PyPI package name for Dragonfly?
The package is named dragonfly2 on PyPI and is installed with pip install dragonfly2. For functionality, the README refers to it simply as Dragonfly.
Does Dragonfly support Wayland on Linux?
The README states that on Linux, Dragonfly is only fully functional on X11. Wayland support is not described.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/dictation-toolbox-dragonfly)