echogarden-project/echogarden: six commands, sixteen engines, three licence statements
Cross-platform speech toolset, used from the command-line or as a Node.js library. Includes a variety of engines for speech synthesis, speech recognition, forced alignment, speech translation, voice isolation, language detection and more.
At a glance
- What is it?
- A TypeScript speech toolset that synthesises, transcribes, aligns, translates, denoises and separates voice, with no Python or Docker and engines delivered as TypeScript, WebAssembly or ONNX. The practical traps are an install that needs an extra flag on npm 12 and a licence position the repository states three different ways.
- Who is it for?
- This fits someone who needs word-level timings on long recordings without provisioning Python or a container, and who is willing to pick one engine per task. Read the licence position yourself before shipping it inside a product, verify the postinstall flag before your first install, and treat the linked release notes as history rather than a changelog for the current line.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 23 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
On npm 12 the install silently loses its native binaries
A dedicated section exists for one reason: npm v12 or later disables postinstall scripts by default. The fix is to add a flag naming the two packages that need to run install hooks, `--allow-scripts=onnxruntime-node,wtf_wikipedia`, to both the install and the update command. The reason it matters is the sentence that follows: the postinstall script used by the onnxruntime-node package downloads binaries that are crucial on some platforms, so without the flag the install may simply fail. Nothing in the one-line install command signals this, so a fresh install on a current npm produces either a broken runtime or a confusing error depending on platform. The plain install itself is otherwise unremarkable:
npm install -g echogarden@latestUpdating is separate, and flagged as unreliable at major version boundaries:
npm update -g echogardenThe documentation is honest that this may not reach the very latest major version, and points at a version checker as the reliable route.
Three licence statements, two licence files, and a third name that does not exist
The licensing here needs reading carefully rather than skimming. The repository metadata carries no licence value at all. The package manifest declares `MIT AND GPL-3.0`. The feature list says the project is fully open-source under GPL v3. The licence section says all source code is MIT, and adds that the package as a whole can be used under MIT when GPL licensed libraries such as eSpeak-NG are not loaded, which is a conditional rather than a single grant. The tree holds both `LICENSE.MIT.md` and `LICENSE.GPLv3.md`, so both texts are present. None of visible text reconciles them, and the conditional means whether you are MIT or GPL depends on which engines you load at runtime. There is also a packaging bug: the published file list includes `LICENSE.md`, a name that appears nowhere in the tree, so the tarball npm serves does not carry either of the two licence files that do exist.
The release notes page stops at 1.0.0 and there is no changelog
The documentation index links a release notes page and describes its scope as releases up to 1.0.0. The package version is 3.4.0. Between those two points sit three minor releases in under a month: v3.2.0 on 2026-08-10, v3.3.0 on 2026-08-30 and v3.4.0 on 2026-09-08. The top-level listing contains no changelog file of any kind, so for a package that has crossed three major versions since the last written notes, there is no single place in the tree to find what changed. Version consistency elsewhere is good, since the manifest version matches the newest tag and the last push date. The practical effect is that upgrading means reading the diff or asking in issues rather than reading notes, which is a real cost on a toolset whose engine list and options both shift between versions.
The manual test script runs a Node flag that has since been withdrawn
Two test entry points are defined. The default runs vitest. The second, named test-manual, invokes the built output directly and opens with `node --experimental-wasi-unstable-preview1 --no-warnings --trace-uncaught`. That flag is the interesting part: it is a preview-gated switch for WebAssembly system interface support, and preview flags of that shape have been removed from Node as the feature stabilised. The manifest still declares an engine floor of node >=18, and the installation text recommends Node v18 or later with v22 or later recommended. So the documented floor and the documented recommendation straddle the release where such preview flags disappear, and the script that exercises the WASI-dependent parts is the one carrying the flag. Anyone who follows the recommendation and then tries the manual suite on a current runtime is the person most exposed to it.
Six sample commands cover a surface that is much larger
The interface sample is six lines:
echogarden speak "Hello World!"
echogarden speak-file story.txt --engine=kokoro
echogarden transcribe speech.mp3
echogarden translate-speech speech.webm subtitles.srt
echogarden align speech.opus transcript.txt
echogarden isolate speech.wavThose six cover synthesis from text and from file, transcription, speech translation to subtitles, alignment against an existing transcript, and voice isolation. Behind them sits a much wider set. Synthesis offers Kokoro and VITS as offline models plus sixteen further offline and online engines, with cloud services from Google, Microsoft, Amazon, OpenAI and ElevenLabs among them. Recognition uses a custom TypeScript and ONNX port of the Whisper architecture, plus whisper.cpp, plus other engines and four cloud providers. Language detection uses Whisper or Silero for audio and TinyLD or FastText for text. Voice activity detection offers WebRTC VAD, Silero VAD, an RNNoise-based variant and a built-in adaptive gate. Denoising uses RNNoise and NSNet2, and source separation uses the MDX-NET architecture.
Alignment and translation each carry their own language count
These are separate operations with separate reach, and the counts differ. Forced alignment uses several variants of dynamic time warping, named as DTW and DTW-RA, with support for multi-pass hierarchical processing, or alternatively guided decoding using Whisper recognition models, and it claims support for 100+ languages. Speech translation is narrower: it takes speech in any of the 98 languages Whisper supports and produces English, with timing described as near word-level rather than word-level. Two further alignment modes sit alongside: one that synchronises audio in one language against a provided English transcript using the Whisper engine, and one that synchronises audio against a translation into other languages when both a transcript and its translation are supplied. Text-to-text translation is present too, but only through a cloud-based Google Translate engine, which is the one listed feature in the set that cannot run offline.
Claiming no system dependencies is achieved through a family of WASM packages
The headline claim is that the toolset does not require Python, Docker or other system-level dependencies, and that it does not lean on essential platform-specific binaries, because engines are written in pure TypeScript, ported through WebAssembly, or imported via the ONNX runtime. The dependency list shows how that is carried. Alongside the AWS SDK clients for Polly and Transcribe streaming sit a run of first-party scoped packages, including ones for audio input and output, FastText compiled to WebAssembly, Flite through WASI, the FreeVoice activity detector, ICU segmentation, and the FFT library the frequency-domain work needs. Each is a pinned minor range rather than a loose version. The one place a real native binary can still appear is the ONNX runtime package whose postinstall fetches platform binaries, which is exactly the case the npm 12 flag exists to cover. Hardware acceleration is handled separately, with a dedicated guide for enabling the CUDA ONNX execution provider.
Pronunciation correction is the one part that is not a generic pass-through
Most entries in the feature list describe plumbing. The exception is the set of enhancements attached to the Kokoro, VITS and eSpeak-NG synthesis engines, which go beyond reading text aloud. Those engines get text normalisation so that dates and currency amounts are pronounced idiomatically rather than read digit by digit, English heteronym disambiguation backed by a simple rule-based model, a set of pronunciation corrections, and acceptance of user-provided pronunciation lexicons for words the rules get wrong. That last one is the escape hatch, and it matters because a rule-based disambiguation model will mis-handle names and loanwords. Delivery of models is handled by an internal package system that downloads and installs voices and models on demand, so the pronunciation additions arrive with an engine rather than as a separate configuration step.
Editorial conclusion
This fits someone who needs word-level timings on long recordings without provisioning Python or a container, and who is willing to pick one engine per task. Read the licence position yourself before shipping it inside a product, verify the postinstall flag before your first install, and treat the linked release notes as history rather than a changelog for the current line.
Frequently asked questions
How do I install echogarden?
Install Node.js v18 or later, with v22 or later recommended, then run `npm install -g echogarden@latest`. On npm v12 or later you also need `--allow-scripts=onnxruntime-node,wtf_wikipedia` because postinstall scripts are disabled by default and the ONNX runtime package downloads required binaries through one.
What licence applies to echogarden?
The repository states it several ways without reconciling them. The package manifest declares `MIT AND GPL-3.0`, the feature list calls the project GPL v3, and the licence section says source code is MIT while the package as a whole may be used under MIT when GPL licensed libraries such as eSpeak-NG are not loaded. Both an MIT and a GPLv3 licence file are present in the tree.
Which speech engines does echogarden support?
Synthesis uses Kokoro and VITS offline models plus sixteen other offline and online engines including Google, Microsoft, Amazon, OpenAI and ElevenLabs. Recognition uses a TypeScript and ONNX port of Whisper, whisper.cpp and other engines. Denoising uses RNNoise and NSNet2, voice isolation uses MDX-NET, and language detection uses Whisper, Silero, TinyLD or FastText.
What languages does echogarden forced alignment and translation cover?
Forced alignment uses DTW and DTW-RA variants or guided decoding with Whisper models and claims 100+ languages. Speech translation covers the 98 languages Whisper supports and outputs English, with timing described as near word-level rather than word-level.
Does echogarden need Python or Docker?
No. Engines are written in pure TypeScript, ported via WebAssembly, or imported through the ONNX runtime, and the supported platforms are Windows, macOS and Linux on both x64 and ARM64. The exception is the ONNX runtime package, whose postinstall downloads platform binaries.
How do I find out what changed between echogarden versions?
The linked release notes page covers releases up to 1.0.0 while the package is at 3.4.0, and there is no changelog file at the top level of the repository. Releases v3.2.0, v3.3.0 and v3.4.0 landed between 2026-08-10 and 2026-09-08.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/echogarden-project-echogarden)