Irene: a Russian offline voice assistant whose documentation describes version 13.0 while the newest release tag is 8.1 from 2023
Ирина - русский голосовой ассистент для работы оффлайн. Поддерживает скиллы через плагины.
At a glance
- What is it?
- Irene is a Python voice assistant that runs speech recognition and speech output locally, with a plugin system, a settings manager, a web API image and a beta mode where a language model picks plugins by tool calling. It is a long lived project with several speech stacks side by side, four different requirements files, and a container that copies a directory the repository does not contain.
- Who is it for?
- Irene is worth trying if you speak Russian, you want recognition and speech output to work without a network round trip, and you are comfortable running one of several scripts depending on which engine you want. Before you build on it, settle four things.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 72 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The documentation describes version 13.0 while the newest release is 8.1 from 2023
Read the version claims in one file and they do not agree with the release list. The installation section describes a settings manager available since version 9.0, AI plugins in beta since version 12.0, and a dedicated section for version 13.0 introducing a streaming recogniser model released in September 2025. The only published release is 8.1, dated 2023-03-27, and the README calls the default branch master rather than main. That gap has a practical consequence, because the repository offers two Windows install paths and the second one is explicitly marked as outdated and tells you to download a release, noting that Python and Git are bundled in it and nothing needs installing. That path installs the 2023 build. The first path, which points at a separate installer repository and a ZIP download, and which is the one the document calls the fastest, is the only route that tracks the current code. Both paths point at the same follow-up document for fine tuning, so a reader who takes the deprecated route and then reads the 13.0 section is working from two different builds of the same program.
The container copies a temp directory that is not in the repository
The Dockerfile builds in two stages, downloading a speech model with a pinned checksum style filename and then assembling a Python 3.10 slim image. Copying the application in is where it breaks:
COPY temp irene/tempThere is no temp directory in the top level listing. What is there is a long set of named directories, among them options_docker, plugins_inactive, plugins_vasi and docker_plugins, plus the media, model, utils and lingua_franca directories the same stage copies successfully. Either the directory is produced before the build, in which case the Dockerfile depends on an undocumented preparation step, or a build from a fresh clone has nothing to copy at that line. The entrypoint makes the second reading worse rather than better, because it installs dependencies when the container starts rather than when the image is built:
ENTRYPOINT ["/bin/sh", "-c", "pip install -r requirements.txt && python runva_webapi.py"]So every start of that container reaches out to a package index and resolves the dependency file again, which means a running instance's behaviour depends on what the index serves that day rather than on what was baked into the image.
The image downloads the older speech model the documentation has moved past
The first build stage fetches one specific model archive:
RUN curl https://alphacephei.com/vosk/models/vosk-model-small-ru-0.22.zip -o ./c611af587fcbdacc16bc7a1c6148916c-vosk-model-small-ru-0.22.zipThat is version 0.22, and the same file elsewhere names it as the old model. The 13.0 section sets out a word error rate comparison on a Sova devices benchmark, where the small 0.22 model scores 29.89, the newer 0.54 line scores 11.6 with a note that the project ships 0.56 and may do slightly better, and Whisper Large V3 scores 15.9. The documentation therefore argues for a newer streaming model and the container ships the older one, which is the single largest gap between what the documentation recommends and what the packaged artifact contains. The same section also offers two ways to cut memory use, downloading an int8 quantisation of the full streaming model or the smaller streaming model instead, at the stated cost of lower quality.
A certificate and key pair sit in the repository root and are copied into the image
Two files at the top level are named localhost.crt and localhost.key, and the container build copies both into the application directory as part of the same stage that copies the web API script and its JSON configuration. The image then exposes port 5003. That combination is a conventional way to give a local web API HTTPS without a real certificate, and a self signed pair for localhost has limited value on its own. It is still worth naming, for three reasons. The private key is in version control, so every clone carries it and the file will outlive the container that needed it. Nothing in the visible documentation explains whether the key is regenerated on first run or shared by every installation, and a shared key is only harmless while the listener stays on a local interface. And the port is exposed by the image, so whether that interface is reachable beyond the machine depends entirely on how the container is run rather than on anything the project controls. Replace the pair per deployment and bind the service to loopback.
The quick start launches a different engine than the version 13 documentation recommends
The installation steps tell you to run one script from the root folder, and by default it starts the offline vosk recogniser for microphone input with the pyttsx3 engine for output. The 13.0 section recommends something else. That path needs two packages installed by hand, sherpa-onnx and its binary build for CPU, with a link to more options for GPU, and then a different script:
pip install sherpa-onnx sherpa-onnx-binNeither package appears in the requirements file, so following the quick start and then the 13.0 section means installing something the installer did not install. The section is also explicit about what that path cannot do yet: it is in beta, it works with only the first microphone on the local device, and the over the network variant is not supported. The repository carries several more entry points for other recognisers and for the web API, so the practical picture is one project with multiple speech stacks and a quick start that points at the conservative one.
Four requirements files, and the base one mixes pinned packages from 2023 with unpinned ones
Dependencies are spread across four files, and the documentation distinguishes them: the main file for a quick install and a fixed file for pinned, checked versions when you want more stability. A third file exists for the container and a fourth for the packaged executable runner. The main file itself is a mix:
wikipedia-api~=0.5.4
pyttsx3~=2.90
vosk~=0.3.45
vosk-tts==0.3.52
termcolor~=1.1.0
sounddevice
soundfile
websockets==15.0.1
pyautogui
requests
numpy
audioplayer
python-dateutil~=2.6.0
gradio~=3.28.3
fsspec==2023.1.0
elevenlabs==1.0.3Six audio and runtime packages carry no version at all, so a fresh install resolves whatever those projects publish now. Two others are pinned exactly at 2023 releases, including a cloud text to speech SDK for a commercial voice service that the visible documentation never mentions, since the documented speech output is local. A desktop automation library that can move the mouse and read the screen is also in the base list with no explanation anywhere in the documentation. On Linux and macOS the documentation separately asks you to install system packages for the audio player before pip runs, which is a reminder that this is desktop software rather than a service.
The language model plugins bill every call by default and the sample configuration stops mid string
The AI plugins feature lets you call a function by speaking loosely, and the mechanism is tool calling: a language model is asked which plugin to invoke and with which arguments. That requires an endpoint in OpenAI compatible form that supports tool calls, and the default the documentation names is a small hosted model reached through another project by the same author, with each call billed and described as costing a few kopecks, plus a free trial credit. The alternative is a local model in LM Studio or Ollama that supports tools, described as slower but functional. Two details are worth carrying into a decision. The configuration key for plugin types takes a comma separated list, so classic and AI plugins can run together, and the documentation states that classic plugins always take priority, which means the language model path is a fallback rather than a replacement for the fixed command grammar. And the sample configuration block ends mid value, on the base URL line with its quote never closed, so the last field of the excerpt is incomplete.
Three different Python version statements and no single documented command
The opening line says Python 3.5 or newer is required, with a parenthetical that the dependency may be lower but any version 3 will do. The installation section gives a different range, describing an installed Python of roughly 3.7 to 3.11 as what you need. The container pins a third, building on a Python 3.10 slim base image. None of these are contradictory in a way that will stop you, but they mean the supported floor is not stated once, and the 3.5 claim is older than several of the pinned dependencies in the same project. Around that sit the other gaps between a reader and a running assistant. The settings folder is not in the repository; it appears after the first launch, which is where you edit behaviour. Support is arranged as five separate installation documents for Windows in two forms, Linux, macOS and debugging, plus a Telegram group for questions and an issue tracker for bugs, and a plugin catalogue with active, inactive and separate plugin directories alongside an installer script for them.
Editorial conclusion
Irene is worth trying if you speak Russian, you want recognition and speech output to work without a network round trip, and you are comfortable running one of several scripts depending on which engine you want. Before you build on it, settle four things. Which version you are actually running, because the release list stops at 8.1 from March 2023 while the documentation describes 13.0 features, and the install path that downloads a release is already marked deprecated. Which engine you want, because the quick start launches the older recogniser while the newer streaming path needs a separate package installed by hand and is documented as beta with a single microphone and no network mode. Which container you use, because the one in the repository copies a directory that is not in the repository and downloads the older speech model the documentation has moved past. And what you do about the certificate and key pair sitting in the repository root, since both are copied into the image that serves a web API. If you only need offline commands, install from the repository and ignore the container. If you want the language model plugins, budget for it separately: the default route bills every call to a hosted model, and the local alternative is described as slower.
Frequently asked questions
How do I start the Irene voice assistant on Linux or macOS?
Install the dependencies with pip using the requirements file, or the fixed file for pinned versions, install the system packages the audio player needs first, then run runva_vosk.py from the root folder. It starts the offline vosk recogniser and the pyttsx3 output engine, and the options folder with the settings is created after the first launch.
Does the Irene voice assistant work without an internet connection?
Yes for the default configuration, which uses a local offline recogniser and a local speech engine. The exception is the AI plugins feature, which needs an OpenAI compatible model endpoint with tool calling support, by default a hosted small model billed per call, or a local model in LM Studio or Ollama that is described as slower.
What does the Vosk streaming mode in Irene require?
Two packages installed by hand, sherpa-onnx and its binary build for CPU, after which you run runva_vosk_sherpa.py. The documentation marks it as beta and says it works with only the first microphone on the local device, with the over the network variant unsupported.
What is the newest published release of the Irene voice assistant?
Version 8.1, dated 2023-03-27. The documentation itself describes features for versions 9.0, 12.0 and 13.0, including a streaming recogniser model from September 2025, and the Windows path that downloads a release is marked as outdated.
How do I install Irene on Windows?
The documented fastest path is a separate installer repository, where you download the code as a ZIP and follow its instructions, then run start-settings-manager.bat to configure plugins and see the extra commands. The older path downloads a release that bundles Python and Git, and is marked as outdated.
Which speech recognition models does Irene compare?
On the Sova devices benchmark quoted in the documentation, the small 0.22 model scores 29.89 word error rate, the 0.54 line scores 11.6 with a note that the project ships 0.56, and Whisper Large V3 scores 15.9. Lower is better.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/janvarev-irene-voice-assistant)