# Voxbento: simultaneous interpretation for conferences, in a browser

> A FastAPI portal where conference interpreters watch the stage in a Jitsi iframe and push translated audio out over WHIP and WHEP, with Python deliberately kept out of the audio path, plus automatic transcription and translation running in the background.

**fossasia/voxbento** — Open Source AI powered Interpretation Platform https://voxbento.com

- Repository: https://github.com/fossasia/voxbento
- Website: https://voxbento.com
- Stars: 1,502 · Forks: 64
- Language: Python
- License: Apache-2.0
- Published: 2026-10-06 · Updated: 2026-10-06 · Language: en
- Canonical page: https://hysenlabs.com/projects/fossasia-voxbento

## Python is not in the audio path, and that is the whole design

The architecture diagram in the README is the most useful thing in the repository, and its most important line is a negative statement: Python is never in the audio path. An interpreter's browser captures microphone audio and posts it with WHIP to MediaMTX on port 8889. MediaMTX terminates WebRTC, remuxes, and serves it back out over WHEP to attendees, with the README claiming sub-second latency. The Python process never touches a sample.

That decision has obvious benefits and one obvious cost, and it is worth being precise about both. The benefit is that translation, transcription and any latency in the model APIs cannot stutter the audio an audience is hearing. The cost is that the platform needs a media server, and the README's prerequisites table lists MediaMTX 1.x as a dependency alongside Python 3.13+, uv and Docker with Docker Compose.

Coordination takes the other path. Who is active in a booth, relay handoff and chat all run over a WebSocket at `/ws/booth/{booth_id}` to a FastAPI portal on port 8000, which also handles state, JWT and REST. So the split is clean: media goes peer to media server, control messages go to Python.

## The audio delay knob is the feature that makes it usable

Rooms can optionally set a listener-side audio synchronization delay for WHEP playback. The default is 0 ms, which leaves the existing low-latency HTML audio path unchanged, and organizers can set values such as 1000, 2000, 5000 or 8000 ms.

The scenario is specific and common. A conference livestream, caption feed or embedded player may itself be delayed by a few seconds behind the room. Translated audio that arrives at genuine low latency is then worse than useless, because the interpreter is describing something the audience has not heard yet. Letting the organizer match the audio to the video is the difference between a usable stream and an unusable one.

The implementation detail matters here and the README is explicit about it. The delay is applied in the listener browser only. MediaMTX, WHIP, WHEP and RTP packets are not changed. That means it costs nothing on the server, it cannot degrade the interpreter's own monitoring feed, and each attendee can be at a different offset.

One consequence is that the booth and the audience experience are decoupled. A booth that has no delay configured is still correct for attendees watching raw stage video, which is why the default stays at 0.

## Transcription and translation run beside the booth, not through it

The FastAPI portal fans out to two background processes. Transcription runs ffmpeg into Deepgram, OpenAI or a local model. Translation runs through Groq, Anthropic or Gemini. Neither is in the audio path; both produce text.

That is the sensible split for a live interpretation product, because captioning and translation output can tolerate latency in a way that spoken audio cannot. It also means the platform is usable at two very different levels of investment: a volunteer who speaks two languages and wants to broadcast needs only a browser and a headset, while an event that wants automatic captioning in twenty-two languages needs API keys, and the cost model changes completely.

There is also a floor audio bot. A headless Chromium instance joins the Jitsi meeting, pulls audio through ffmpeg into RTSP, hands it to MediaMTX, and feeds transcription and translation. That is how the system gets the stage audio when no human interpreter is covering a particular pair of languages, and it is why the Jitsi stack is a hard dependency rather than an optional extra.

The `pyproject.toml` confirms the local-model story is real rather than aspirational. Dependencies include `faster-whisper`, `ctranslate2`, `transformers` and `numpy`, which is a full local speech stack, plus `tenacity` for retry logic on the API calls.

## Bringing it up with one command and three required secrets

Option 1 in the README starts every service, portal, MediaMTX and Jitsi, in containers.

```bash
git clone https://github.com/fossasia/voxbento.git
cd voxbento
cp .env.example .env
echo "ADMIN_PASSWORD=$(openssl rand -hex 16)" >> .env
echo "API_KEY_ENCRYPTION_KEY=$(openssl rand -hex 32)" >> .env
echo 'DOCKER_HOST_ADDRESS=192.168.1.x' >> .env
docker compose up --build
```

Then you open `http://localhost:8000`. Three environment variables are mandatory and each has a reason worth knowing. `ADMIN_PASSWORD` is your admin credential. `API_KEY_ENCRYPTION_KEY` encrypts third-party API keys in the database and must be 32 characters or longer. `DOCKER_HOST_ADDRESS` must be your machine's LAN IP, because Jitsi's JVB and MediaMTX WebRTC ICE both need candidates that are actually reachable; the README gives the macOS and Linux one-liners for finding it.

Key rotation is supported without breaking existing rows. Provide a comma-separated list of keys and Voxbento encrypts new tokens with the first one while attempting decryption with all of them. That is a small design decision that will matter to you on the day a key leaks.

The compose file defaults the database to SQLite on a Docker volume, with a comment giving the PostgreSQL URL to override it for production. Alembic handles migrations, and the container runs `alembic upgrade head` before starting uvicorn.

## A Python version conflict you will hit before anything else

Three documents disagree about the Python version, and reconciling them is the first task. The README prerequisites table says Python 3.13+. `pyproject.toml` pins `requires-python = ">=3.13,<3.14"`, which excludes 3.14. The `Dockerfile` is built `FROM python:3.14-slim`, which is 3.14, the one version the project metadata refuses.

The Docker path may well work anyway, since `uv sync` can install into an interpreter the metadata does not describe, and nothing in the README mentions the conflict. But a pip install from source on 3.14 will refuse, and a native developer following the README's 3.13+ claim on 3.14 will get an error with no obvious cause. If you run this outside Docker, pin 3.13.

A second naming detail suggests this is a project that outgrew its packaging. The `pyproject.toml` still calls the package `eventyay-interpretation-portal` with the description Lightweight collaborative interpretation booth for Eventyay, while the repository, the homepage and the environment defaults all say Voxbento. The Docker compose file still defaults the Jitsi room to `eventyay-stage-room`. This is harmless at runtime but it tells you the packaging metadata has not been renamed to match the product.

Native development setup is documented in `CONTRIBUTING.md`, and the full API reference, environment variables and configuration live at docs.voxbento.com rather than in this repository.

## Security defaults, and where this stops being a small deployment

Several decisions here are more careful than the project size suggests. The `.env.example` marks the encryption key as CRITICAL and documents the rotation semantics. JWT settings are separate from the secret key, with `JWT_EXPIRY_SECONDS` defaulting to 86400 and an empty `JWT_SECRET` falling back to `SECRET_KEY`. Passwords use bcrypt, and API keys are encrypted at rest with the `cryptography` package rather than stored as plaintext.

One default is worth flagging rather than repeating. In `.env.example`, `SECRET_KEY=change-me` is the shipped value, so a deployment that copies the example without editing it is running with a known key. The compose file does the same thing, defaulting to `change-me`, and the Jitsi auth passwords are both `changeme`. That is normal for an example file and a real risk if you treat the file as configuration rather than a template.

The size of the deployment is the other constraint. A working install is a FastAPI portal, MediaMTX, a Jitsi stack with its own auth, optional ffmpeg for the bot, and a headless Chromium for floor audio. That is several containers and a Caddyfile in the tree for TLS termination, with `JITSI_BASE_URL=https://jitsi.voxbento.com` already filled in pointing at production. Behind that sits a database and, optionally, keys for Deepgram, OpenAI, Groq, Anthropic or Gemini.

The repository publishes no releases at all, so version tracking means reading commit history rather than tags. The last push was on 2026-09-20 and it is not archived.

## Conclusion

Voxbento solves a problem most conference tooling ignores, which is that a spoken translation delivered over a conference video feed is unusable if it arrives half a second out of step with the speaker on screen. Routing the interpreter audio through MediaMTX with WHIP ingest and WHEP playback, and adding a listener-side delay of 1000 to 8000 ms when the video lags, is the right shape for that. It is the wrong tool for a two-language event with an audience of two, given a Docker stack, a self-hosted Jitsi and ffmpeg. The last push was on 2026-09-20 and the repository is not archived, but it publishes no releases at all. Read the architecture diagram in the README first, then reconcile the Python version question between `pyproject.toml` and the `Dockerfile` before you deploy.

## FAQ

### What is Voxbento?

A real-time interpretation platform for live events, described as an open source AI powered interpretation platform. Interpreters monitor the stage video through a Jitsi iframe in the browser and broadcast translated audio to attendees with low latency, which suits conferences that need spoken translation rather than captions alone.

### How does Voxbento stream interpreter audio to attendees?

The interpreter browser posts microphone audio to MediaMTX with WHIP, MediaMTX terminates WebRTC and remuxes, and attendees play it back over WHEP, which the README describes as sub-second latency. Python is never in the audio path, so API latency in transcription or translation cannot affect the audio the audience hears.

### What does Voxbento need before it will start?

Python 3.13 or newer, uv as the package manager, MediaMTX 1.x as the WebRTC and HLS audio server, and Docker with Docker Compose for the Jitsi stack. Three environment variables are mandatory: `ADMIN_PASSWORD`, a `API_KEY_ENCRYPTION_KEY` of at least 32 characters, and `DOCKER_HOST_ADDRESS` set to your machine's LAN IP so WebRTC ICE candidates are reachable.

### Can Voxbento delay translated audio to match a lagging video feed?

Yes, and that is what the listener-side synchronization setting is for. Rooms can set 1000, 2000, 5000 or 8000 ms, with a default of 0 ms which leaves the low-latency path unchanged. The delay applies in the listener browser only, so MediaMTX, WHIP, WHEP and RTP packets are untouched.

### Does Voxbento support automatic captioning and translation?

Yes, as background processes beside the audio path rather than inside it. Transcription runs ffmpeg into Deepgram, OpenAI or a local model, and translation runs through Groq, Anthropic or Gemini. The `pyproject.toml` also depends on `faster-whisper`, `ctranslate2` and `transformers`, so local speech models are supported as an alternative to the hosted APIs.

### What license is Voxbento released under?

Apache-2.0, stated in the repository metadata with a `LICENSE` file in the tree. The licence carries no reciprocity requirement, though the Apache-2.0 notice obligations still apply to anyone redistributing it.

## Sources

- [fossasia/voxbento on GitHub](https://github.com/fossasia/voxbento)
- [Issues](https://github.com/fossasia/voxbento/issues)
- [License: Apache-2.0](https://github.com/fossasia/voxbento/blob/main/LICENSE)
- [Project website](https://voxbento.com)
- [README](https://github.com/fossasia/voxbento/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/fossasia-voxbento
