ARIA v2: Windows system audio turned into live subtitles with local Whisper, Sherpa-ONNX and Vosk
ARIA - AI Realtime Intelligent Audio | Universal real-time AI subtitles for Windows
At a glance
- What is it?
- ARIA is a Windows desktop app that captures system audio or a microphone and renders live subtitles in a movable overlay, with separate translation. It ships as a prebuilt package with an embedded Python runtime, and the Full build carries Whisper Large-v3, Sherpa-ONNX, Vosk and NLLB offline.
- Who is it for?
- ARIA fits Windows users who want subtitles for audio they cannot re-route: browser playback, calls, games, and video players, especially where the source has no caption track. Skip it if you are on Linux or macOS, if you need a pip-installable library rather than a desktop app, or if a 7.6 GB download is unacceptable.
- Can I use it commercially?
- Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository last received commits 53 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What ARIA solves, and for whom
Most captioning tools sit inside one application. A video player has its own subtitle track, a meeting client has its own live captions, a browser extension handles one site. The gap is everything else: audio that plays on the machine with no caption path at all. ARIA takes Windows system audio as its input, runs it through a speech recognition engine, and draws the text in an overlay window that floats above whatever is playing. The README lists video players, browsers, games and calls as the intended sources.
The audience is narrow but clear. It is a Windows-only desktop application, and pyproject.toml classifies it as "Operating System :: Microsoft :: Windows" with an intended audience of "End Users/Desktop". There is no server component, no API to call from another program, and no Linux or macOS build. If you need captions inside a Linux pipeline, this is not the project. If you are on Windows 11 and want captions over a stream that has none, the Lite build exists precisely for that case.
Three recognition engines behind one overlay
ARIA does not run a single model. The recognition mode table maps each mode to a different engine and a different trade-off. Precise uses Faster-Whisper and the README recommends it for recorded video and talks, where accuracy matters more than latency; the Full package ships Whisper Large-v3 at roughly 3 GB for this. Realtime uses Sherpa-ONNX or Vosk and targets conversations and streams where fast updates matter; the bundled models are a 500 MB Sherpa-ONNX bilingual model for Chinese and English and a 1 GB Vosk model for Japanese. Live Captions delegates to the Windows 11 Live Captions feature, which means no CUDA and no NVIDIA GPU.
The data flow is therefore: capture (system audio or microphone, via PyAudioWPatch), recognition (one of the three engines), then rendering into two independent overlays, one for the source text and one for the translation. The README states that the subtitle and translation overlays can be moved and resized independently, and that settings persist between runs. Translation can go to Google Translate, Bing Translator, Youdao, or the local NLLB model bundled in the Full package.
The language matrix is where the mode choice bites. Precise supports Chinese, English, Japanese, Korean, Spanish, French and more. Realtime supports Chinese and English through Sherpa-ONNX and Japanese through Vosk, with dashes for Korean and for Spanish and French. Live Captions covers everything listed, but the README notes its availability is determined by the Windows feature, not by ARIA. So the mode that runs everywhere has the least control, and the mode with the broadest language list wants an NVIDIA GPU.
Installing ARIA from the prebuilt packages
There is no pip install path documented for end users, despite the pyproject.toml entry point named aria. The README directs you to two prebuilt archives hosted on Google Drive and Baidu. The Lite build is about 600 MB and requires Windows 11, version 22H2 or later, because it depends on the Live Captions feature. The Full build is about 7.6 GB and runs on Windows 10 or 11, with an NVIDIA GPU recommended for Precise mode. The size difference is the offline models.
After extracting, the launcher differs by build. The README gives these two names:
ARIA.bat # Full package launcher
run_lite.bat # Lite-package launcherThe package includes its own Python runtime, so you do not install Python separately. The repository layout reflects that: python/ holds the embedded runtime, src/ holds the application source, models/ holds the offline models, and the Lite build lives in its own ARIA-v2-lite/ directory with its own python/ and src/ and its own run_lite.bat.
For the Full build the getting-started steps are: start ARIA.bat, select a recognition mode, configure translation if you want it, then press Start. For Lite you start run_lite.bat and you are limited to Live Captions mode by design. If you would rather run from source, pyproject.toml declares requires-python >=3.10 and lists the runtime dependencies, including faster-whisper, sherpa-onnx, vosk, PyQt6, PyAudioWPatch, and the Live Captions mode dependencies pywinauto, uiautomation, comtypes and pyautogui. The console entry point is aria = "realtime_subtitles.ui:run_app".
pip install -e .
ariaThat path is inferred from pyproject.toml, not from the README, which does not document a source install.
Where ARIA gets in the way
The Lite and Full split is the first real constraint. Lite is not a smaller version of the same tool; it is Live Captions mode only, and the README says so directly: "Lite is intentionally limited to Live Captions mode." If you download the 600 MB build expecting to switch to Whisper later, you cannot. You re-download 7.6 GB.
Precise mode's GPU recommendation is the second. The README says it works best with an NVIDIA GPU. A machine without CUDA is pushed toward Live Captions or Realtime, and Realtime drops Korean, Spanish and French from its language coverage. The mode that needs no GPU is also the one whose language support ARIA does not control.
Translation has a dependency the project does not hide. The README states that online translation services "can be rate-limited or change without notice" and recommends the local NLLB model when offline operation and predictable availability matter more. That is a fair warning, but it also means the convenient setup (online providers) is the fragile one. The local alternative adds 1.2 GB to an already large download.
Finally, the metadata and the release history disagree in a way worth noting. pyproject.toml declares version 1.0.0 and a classifier of "Development Status :: 3 - Alpha", while the most recent release listed is v2.0.0-lite from 2026-02-08. Treat the package metadata as stale rather than as the current state. The last push to the repository was on 2026-08-09.
ARIA against building your own Faster-Whisper loop
The obvious alternative is not another subtitle app. It is writing the pipeline yourself: capture system audio with PyAudioWPatch, feed it to faster-whisper, and draw a Qt window on top. ARIA's dependencies are exactly that stack, so the comparison is about what the packaging adds rather than about the recognition quality.
What you would have to rebuild: the three-mode switch between Faster-Whisper, Sherpa-ONNX and Vosk; the Live Captions path, which uses pywinauto, uiautomation, comtypes and pyautogui to drive a Windows feature through its UI rather than through an audio API; two independently movable overlays; the four translation backends including local NLLB; and the session and transcript logging with a selectable log timezone. The Live Captions mode in particular is not something you would write by choice, since it automates the Windows feature instead of calling a model.
What you keep by writing it yourself: control over the model, the ability to run headless, and no 7.6 GB archive. If your target is a single language and a single source, a few hundred lines of faster-whisper plus a Qt overlay is a smaller thing to own. ARIA's value is the mode matrix and the packaging, not the recognition engine, which is a dependency in both cases.
Licence and the cost of staying current
ARIA is GPL-3.0, stated in the README and in pyproject.toml as license = { text = "GPL-3.0" }. For an end user running the prebuilt package this changes nothing about daily use. For anyone who wants to embed ARIA in a product, the copyleft terms apply to the combined work, and the bundled components are separate projects with their own terms: Faster Whisper, Sherpa-ONNX, Vosk, PyQt6 and PyAudioWPatch are all acknowledged in the README. PyQt6 in particular is dual-licensed by its vendor, so a commercial redistribution needs its own review. This is a description of what the files say, not legal advice.
Upgrade cost is dominated by the download, not by a package manager. Because distribution is through Google Drive and Baidu archives rather than PyPI, there is no dependency resolver to run and no version pinning to manage, but there is also no incremental update. A new Full release means fetching roughly 7.6 GB again, models included. The models themselves are the reason: Whisper Large-v3 at 3 GB, Vosk Japanese at 1 GB, NLLB at 1.2 GB, Sherpa-ONNX at 500 MB. If you only ever use one mode, you are still carrying the others.
The repository has a CHANGELOG.md at the top level, so release notes exist, but the README does not document a rollback procedure or an in-place upgrade path. The README is also silent on how to verify the integrity of the downloaded archives.
Editorial conclusion
ARIA fits Windows users who want subtitles for audio they cannot re-route: browser playback, calls, games, and video players, especially where the source has no caption track. Skip it if you are on Linux or macOS, if you need a pip-installable library rather than a desktop app, or if a 7.6 GB download is unacceptable. Before committing, verify three things in the Full package: that your GPU is good enough for Precise mode, that your target language is covered in the Realtime column of the language table (Korean and Spanish are marked with a dash there), and that the online translation provider you pick is not the one you depend on for offline work, since the README warns those services can be rate-limited or change without notice.
Frequently asked questions
How do I get ARIA?
Download one of the two prebuilt archives linked in the README, from Google Drive or Baidu. Lite is about 600 MB and needs Windows 11 version 22H2 or later; Full is about 7.6 GB and runs on Windows 10 or 11. Extract the package, then start ARIA.bat in the Full package or run_lite.bat in the Lite package.
Does ARIA need a separate Python installation?
No. The README states the package includes its own Python runtime, and the repository layout shows a python/ directory in both the Full package and ARIA-v2-lite/. A separate Python installation is not required.
Which languages does ARIA support in Realtime mode?
Realtime mode supports Chinese and English through Sherpa-ONNX and Japanese through Vosk. The language table marks Korean, Spanish and French with a dash for Realtime, while Precise mode supports the broadest set of languages.
Can ARIA run without an NVIDIA GPU?
Yes, but not in Precise mode, which the README says works best with an NVIDIA GPU. Live Captions mode uses the Windows 11 Live Captions feature and needs no NVIDIA GPU or CUDA setup, and Realtime mode uses Sherpa-ONNX or Vosk.
Does ARIA translate subtitles offline?
The Full package includes a local NLLB translation model of about 1.2 GB, which the README recommends when offline operation and predictable availability matter. The other options are online services: Google Translate, Bing Translator and Youdao, which the README warns can be rate-limited or change without notice.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/sayksii-aria)