Kaldi Active Grammar switches whole grammars on at the start of each utterance
Python Kaldi speech recognition with grammars that can be set active/inactive dynamically at decode-time
At a glance
- What is it?
- A Python package for Kaldi and Dragonfly that makes each command grammar independently activatable per utterance, so fewer phrases compete for recognition, and that ships as platform-specific wheels bundling binaries from a private Kaldi fork.
- Who is it for?
- Adopt it if you run voice command and control through Dragonfly on a 64-bit machine and can accept an AGPL-3.0 dependency on a Kaldi fork. Do not adopt it for plain dictation work, where a general speech stack is simpler, or if your team needs source distributions or a model other than a left-biphone nnet3 chain model.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 29 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A Kaldi decoding graph is static for the whole session
The problem this package addresses is structural in Kaldi itself. Decoding graphs are monolithic, they need expensive up-front off-line compilation, and they stay static while decoding runs. Add a grammar and you recompile; every grammar in the graph can be recognised at any moment.
Kaldi's newer grammar framework relaxes that in three ways at once. Multiple independent grammars, with nonterminals, can be compiled separately and stitched together dynamically at decode time. What it does not offer is a way to switch them off, so all the grammars remain active and recognisable for the whole session.
That is the gap kaldi-active-grammar fills. Each grammar or rule can be marked active or inactive independently, dynamically, on a per-utterance basis, set at the beginning of each utterance. Dragonfly can then activate only the grammars that suit the environment the user is actually in.
Fewer live grammars is an accuracy change, not only a speed one
It would be easy to present per-utterance switching as a performance feature. The stated benefit is accuracy, and the mechanism is straightforward: every grammar left active is a set of phrases competing for the same acoustic evidence, so activating fewer of them means fewer possible recognitions.
For a command and control system this matters more than for general dictation. A window manager has a different vocabulary from a code editor, and a browser from a terminal. When a user switches context, the grammars that do not belong there are still in the graph if nothing switches them off, and a mis-recognition can land in the wrong command.
The per-utterance boundary is what makes this usable. State is set once at the start of an utterance rather than being threaded through the decoder, so an application can change its listening profile between utterances without touching the compiled artefacts.
The dictation grammar is compiled once and shared by every command
The second structural saving is about compilation cost rather than recognition. The dictation grammar can be shared between all the command grammars, which means it compiles quickly without every command grammar having to embed large-vocabulary dictation directly.
That is what makes the dynamic stitching affordable. If free dictation had to be duplicated inside every command grammar, the compile step would grow with each one and the architecture would be no better than the monolithic graph it replaces.
For anyone who only wants plain dictation, there is a separate interface that takes either a specified HCLG.fst file or the pre-trained dictation model included with the package. The documentation is blunt that this is the boring use case, and treats it as supported rather than as an afterthought, which is convenient when you are evaluating whether the package is worth its native dependencies at all.
Wheels only, because the native library has to match the platform
Distribution policy is the first thing to understand about this package. The Python package includes the binaries needed for decoding on Windows, Linux and macOS, and they are built from the author's own Kaldi fork, which is stated to be intended for use by kaldi-active-grammar directly rather than as a stand-alone library.
Because of that, only platform-specific wheels are produced and supported. Source distributions are intentionally unsupported, on the grounds that a usable package has to contain the matching native Kaldi library. The build files make the same point in code: the build system requires setuptools, wheel, scikit-build, cmake and ninja, and setup.py forces the wheel to be platform-specific because of manually-loaded native libraries.
Two environment variables exist for people building outside that path. KALDIAG_BUILD_SKIP_NATIVE skips the native build when the required libraries have already been placed in the package, and KALDIAG_SETUP_RAW switches to plain setuptools instead of scikit-build. If you build from source, you are building native code, not installing a Python library.
Only left-biphone nnet3 chain models, and yours must be converted
Model support is narrow and stated up front: only Kaldi left-biphone models work, specifically nnet3 chain models, and only with specific modifications.
A compatible general English nnet3 chain model trained on roughly 3000 hours of open audio ships under the project releases, with model information and comparisons in docs/models.md and improved models under development. The install path is to download and unzip it, then pass the directory path to the kaldi-active-grammar constructor.
Using your own model is allowed with a caveat the documentation does not soften: standard Kaldi models must be converted to be usable. Conversion can be performed automatically, and that automation is not fully implemented yet. That sentence is the first thing to test if your acoustic model is not the shipped one, because the whole per-utterance mechanism sits on top of a model you may not be able to bring in at all.
The resource floor is honest as well: Python 3.6 or newer on 64-bit only, Windows, Linux or macOS, and roughly 1GB or more of disk for the model plus temporary storage and cache, and roughly 1GB or more of RAM for model and grammars, both scaling with grammar complexity.
On Windows, the install is a zip you do not have to assemble
For Windows users the project publishes self-contained portable packages under its releases, described as batteries-included because they carry Python, the libraries and the model together.
There are three: kaldi-dragonfly-winpython bundles kaldi-active-grammar with dragonfly2, kaldi-dragonfly-winpython-dev is the more recent development version of the same pairing, and kaldi-caster-winpython-dev adds caster on top. Each one unzips and runs, which removes the wheel-per-platform step, the native library matching and the model download from the critical path.
Linux has a different answer. A community Docker image, kmdouglass/caster-kaldi, runs KaldiAG with Dragonfly and Caster inside a container while using the host's microphone, which is the practical way to get audio into a container without a virtual device.
The surrounding ecosystem is versioned rather than bundled. A Dragonfly backend compatible with this package was merged as of Dragonfly v0.15.0, and Caster support arrived with KaldiAG v0.6.0 alongside Dragonfly v0.16.1. Since v0.2 the project develops itself with itself, and a plain-dictation example sits alongside full_example.py, mimic.py, mix_dictation.py and audio.py in the examples directory.
AGPL-3.0, a pinned native revision, and a release every two years
The licence is AGPL-3.0, which matters more here than in a pure Python package because the package ships modified Kaldi binaries rather than merely importing a library.
The tree shows how seriously the native side is pinned. A kaldi-native-revision.txt records which Kaldi revision the binaries came from, alongside a CMakeLists.txt, a Justfile, BUILDING.md and separate requirement files for building, editable installs and testing.
Testing is split by marker rather than run wholesale. pytest defaults to excluding the prolonged and stress markers, and each is opt-in: prolonged covers long-term decoder tests at 1000 utterances per framework, source_build covers source-tree build tooling that an installed wheel does not contain, and stress covers a long-term harness in tests/stress. xfail_strict is on, so an unexpected pass fails the suite.
Release timing has been uneven. Version 3.0.0 and 3.1.0 landed within weeks of each other in late 2021, then 3.2.0 waited until 2025-11-02, and the last push was 2026-09-03 on the master branch. The visible README also cuts off mid-sentence inside its own installation steps, so the exact pip and binary steps are not in this copy.
Editorial conclusion
Adopt it if you run voice command and control through Dragonfly on a 64-bit machine and can accept an AGPL-3.0 dependency on a Kaldi fork. Do not adopt it for plain dictation work, where a general speech stack is simpler, or if your team needs source distributions or a model other than a left-biphone nnet3 chain model. Before anything else, check whether the pre-trained English model in the project releases converts cleanly, since converting a standard Kaldi model is the step the documentation marks as unfinished.
Frequently asked questions
What does kaldi-active-grammar add to Kaldi's own grammar framework?
Kaldi's grammar framework compiles grammars separately and stitches them together at decode time, but all of them stay active. kaldi-active-grammar lets each grammar or rule be marked active or inactive independently and dynamically, per utterance, so only the grammars for the current environment compete for recognition.
Which Kaldi models does kaldi-active-grammar support?
Only left-biphone models, specifically nnet3 chain models with specific modifications. A general English model trained on about 3000 hours of open audio ships under the project releases, and any standard Kaldi model has to be converted first, with the automatic conversion not fully implemented.
Can I run kaldi-active-grammar on Windows without installing Python first?
Yes. The project releases include self-contained portable packages such as kaldi-dragonfly-winpython and its development variant, plus a caster variant, each bundling Python, the libraries and the model so you can unzip and run.
Why does kaldi-active-grammar publish no source distribution?
Only platform-specific wheels are produced, because a usable package must bundle the matching native Kaldi library built from the author's fork. Source distributions are intentionally unsupported, and setup.py forces the wheel to be platform-specific for the same reason.
How much memory and disk does kaldi-active-grammar need?
About 1GB or more of disk for the model plus temporary storage and cache, and about 1GB or more of RAM for the model and the grammars, both scaling with grammar complexity. Python 3.6 or newer is required and 64-bit is mandatory.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/daanzu-kaldi-active-grammar)