Sokuji: Real-Time Two-Way Speech Translation for Bilingual Meetings
Real-time two-way speech translation for bilingual meetings — auto-detects the spoken language and translates both directions, cloud or fully offline on-device. Desktop (Windows · macOS · Linux) + browser extension (Chrome · Edge) for Zoom, Meet, Teams & any app.
At a glance
- What is it?
- Sokuji is a TypeScript desktop app and browser extension from Kizuna AI Lab that detects the spoken language and translates both directions of a conversation, using cloud providers or fully on-device inference. It is AGPL-3.0, and the last push to the repository was on 2026-09-10.
- Who is it for?
- Adopt Sokuji when you need one install that covers both directions of a bilingual meeting and you want the option to keep audio on-device. Skip it if your platform is not Windows, macOS, Linux, Chrome or Edge, or if you cannot accept AGPL-3.0.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Sokuji solves for bilingual meetings
A bilingual meeting has an awkward asymmetry. One person speaks, the other waits for a translation, then answers, and the translation runs only one way unless someone switches a setting mid-sentence. Sokuji targets that specific shape of conversation. The README describes the goal as "two people, two languages, one conversation" and states that the app auto-detects which language is being spoken and translates it into the other, in both directions, in real time. You set Language A and Language B, capture system audio and microphone together, and the two streams are handled in a single session.
The intended user is not a professional interpreter. It is an engineer or product person on a call with a counterpart who speaks a different language, where the alternative is a second tool, a shared document, or a human interpreter. The README also notes a live subtitle mode: you can share your screen so the other side reads along. That is a different use case from replacing the audio, and it matters because subtitles degrade more gracefully than synthesized speech when the connection is poor.
The project is built by Kizuna AI Lab, and the name is explained in the README: "Kizuna" means bond in Japanese, and Sokuji (即時) is the word for immediate. The repository is TypeScript, licensed AGPL-3.0, and the last push was on 2026-09-10, with v0.40.3 released on 2026-09-07.
How the translation path is wired
The pipeline is a straight line: audio in, speech recognition, translation, speech synthesis, audio out. The README's diagram shows the input as your voice in any language, then a branch point where you choose cloud or local, then the translated voice, then the target application: Zoom, Teams, Meet, Discord, or any app.
On the cloud side, the README lists nine providers: OpenAI, Google Gemini, Palabra.ai, Kizuna AI, Doubao AST 2.0, Soniox, Zoom AI, an OpenAI-compatible endpoint, and Local Inference. The two-way meeting feature is attributed in the README to "Soniox two-way mode", with a claim of 60+ languages and 3,600+ language pairs. The OpenAI-compatible entry is the interesting one for anyone running their own gateway, because it means the translation stage does not have to be a vendor you have already integrated.
On the local side, the README states that ASR, translation and TTS run on-device through WASM and WebGPU, with no API key and no internet. It lists 44 ASR models, 75 translation models, and 137 TTS models across 53 languages, with one-click download and IndexedDB caching. The package.json build configuration shows the WASM payloads that ship with the desktop app: sherpa-onnx-asr, sherpa-onnx-asr-stream, sherpa-onnx-tts, ort, vad, piper-plus and gtcrn. That list is the concrete evidence that local inference is bundled rather than fetched, and it also tells you where the installer size comes from.
Installing Sokuji and running a first session
There are two distribution channels, and they are not interchangeable. The desktop app works with any application that accepts microphone input, which the README lists as Zoom, Teams, Discord, Slack, games and OBS. The browser extension works inside web meeting platforms: Google Meet, Teams, Zoom, Yandex Telemost, Discord, Slack, Gather.town, Whereby and Jitsi Meet.
The desktop builds are on the Releases page. Windows uses `Sokuji-x.y.z.Setup.exe`; macOS has separate Apple Silicon and Intel packages, `Sokuji-x.y.z-arm64.pkg` and `Sokuji-x.y.z-x64.pkg`; Linux ships `.deb` packages for Debian and Ubuntu on x64 and arm64. Note that the README's install table lists the `.deb` files while the Electron build configuration in package.json also targets AppImage for both architectures. If you are on a distribution that does not take `.deb`, check the Releases page rather than assuming.
The extension installs from the Chrome Web Store or Microsoft Edge Add-ons. If you prefer not to use the stores, the README documents a developer-mode path:
# 1. Download sokuji-extension.zip from the Releases page
# 2. Extract it
# 3. Open chrome://extensions/ and enable Developer mode
# 4. Click "Load unpacked" and select the extracted folderBuilding from source is a standard Node workflow. The README gives these four lines, and the two npm scripts are the ones you will actually use:
git clone https://github.com/kizuna-ai-lab/sokuji.git
cd sokuji && npm install
npm run electron:dev # Development
npm run electron:build # ProductionFor a first real session, the sequence the README describes is: open Sokuji, set Language A and Language B, then capture system audio and microphone together. The system-audio capture is what lets the app hear the other participant, so if you only grant microphone access you get one direction. Then join your meeting and select Sokuji as the audio device. If you want the other side to read along, enable the subtitle mode and share your screen.
Where Sokuji is the wrong tool
The platform matrix is the first hard boundary. Desktop covers Windows, macOS and Linux; the extension covers Chrome and Edge, with Brave marked as coming soon in the README. Firefox and Safari are not listed. If your organisation standardises on Firefox, the extension route is closed and you are left with the desktop app.
The second boundary is the two-way feature itself. The README attributes it to Soniox two-way mode. That means the headline capability is tied to a specific cloud provider, not to the local inference path. The README presents Local Inference as the privacy option and Soniox as the two-way option, and it does not state that local models provide the same automatic bidirectional detection. Anyone whose reason for choosing Sokuji is offline operation should verify that specific behaviour before assuming the two features compose.
Third, the desktop app captures system audio. On Linux that generally means a loopback device, and the repository contains a top-level directory named `BlackHole`, which is the macOS virtual audio driver used for the same purpose. This is not a zero-configuration capture path; you are routing audio, and routing audio is where setups break. The README does not document a fallback when system audio capture is unavailable.
Finally, the licence. AGPL-3.0 is a strong copyleft licence, and the README does not discuss what it means for embedding Sokuji in another product. That question belongs with your legal team, not with this article.
How Sokuji differs from captioning tools
The nearest alternative most teams already have is a live captioning layer rather than a translation layer. Zoom, Google Meet and Teams all offer captions, and some offer translated captions. The difference in approach is what happens to the audio. Captioning leaves the original voice untouched and adds text on top; Sokuji's primary mode replaces or accompanies the audio with synthesized speech in the target language, and its subtitle mode is an addition rather than the whole product.
That distinction decides which one you want. If the goal is accessibility for a listener who reads the language but cannot hear well, captions are the right tool and Sokuji is overkill. If the goal is a participant who does not speak the language at all, captions force them to read at speaking speed while also formulating a reply, which is exactly the load Sokuji's audio path removes.
The second difference is provider independence. Platform captions are bound to the platform. Sokuji's provider list includes an OpenAI-compatible endpoint, so the translation stage can point at infrastructure you already run, and the local path removes the network entirely. The trade-off is that you now own the configuration: model selection, provider credentials, and the audio routing that platform captions handle for you.
Maintenance, releases and the AGPL-3.0 licence
The release cadence is visible and recent. v0.40.1 landed on 2026-09-05, v0.40.2 on 2026-09-05, v0.40.3 on 2026-09-07, the last push to the repository was on 2026-09-10, and package.json declares version 0.41.0 with a separate `sidecarVersion` of 0.3.0. The repository is not archived. The presence of a `CHANGELOG.md`, a `CLAUDE.md`, a `CONTEXT.md`, an `evals/` directory and a `benchmark/` directory suggests the project is being developed with its own tooling rather than as a side project.
Upgrade cost is the part to think about before adopting. The desktop app bundles WASM runtimes for ASR, TTS, VAD and the ONNX runtime, so each release is a substantial download, and local model files are cached in IndexedDB rather than shipped inside the installer. The `.env.example` file also documents a build-time feature-gating system for Kizuna-managed providers: a master gate `VITE_ENABLE_KIZUNA_AI` plus one gate per provider. The comment in that file states that adding a gate is not enough on its own, because it must also be forwarded in `extension/vite.config.ts` and in every build step of `.github/workflows/build.yml`, and that `featureGateForwarding.consistency.test.ts` fails when either is missed. If you build Sokuji yourself, that is the mechanism you will trip over when enabling a provider.
On licensing: the repository is AGPL-3.0, and the README does not address commercial embedding, hosted derivatives, or what the network clause means for a modified build. This article is not legal advice; the licence text in `LICENSE` and your own counsel are the sources that matter.
Editorial conclusion
Adopt Sokuji when you need one install that covers both directions of a bilingual meeting and you want the option to keep audio on-device. Skip it if your platform is not Windows, macOS, Linux, Chrome or Edge, or if you cannot accept AGPL-3.0. Verify your meeting platform and the actual model download size before committing.
Frequently asked questions
Does Sokuji work offline?
The README states that Local Inference runs ASR, translation and TTS on-device through WASM and WebGPU, with no API key and no internet. Cloud providers are the alternative path and require connectivity.
Is Sokuji available as a Chrome extension?
Yes. The browser extension is published on the Chrome Web Store and Microsoft Edge Add-ons, and the README also documents loading it unpacked from `sokuji-extension.zip` in developer mode. The extension works inside web meeting platforms such as Google Meet, Teams, Zoom, Discord and Jitsi Meet.
Which platforms does the Sokuji desktop app support?
The README lists Windows, macOS and Linux. Releases include `Sokuji-x.y.z.Setup.exe` for Windows, `.pkg` packages for Apple Silicon and Intel macOS, and `.deb` packages for Debian and Ubuntu on x64 and arm64.
Can Sokuji translate both directions at once?
The README describes a two-way mode where you set Language A and Language B, capture system audio and microphone together, and the app auto-detects the spoken language and translates it into the other. The README attributes this mode to Soniox two-way mode.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/kizuna-ai-lab-sokuji)