# joinly lets an agent join your meeting, and its server has no authentication by design

> Middleware that joins a browser-based call on your behalf, exposes meeting tools through a model context protocol server, and streams the conversation to whichever speech and language providers you choose. The interesting part is the boundary it draws: the server accepts no authentication and takes configuration from the client, so the documentation tells you to bind it to loopback and treat it as single-client.

**joinly-ai/joinly** — Make your meetings accessible to AI Agents

- Repository: https://github.com/joinly-ai/joinly
- Website: https://joinly.ai
- Stars: 567 · Forks: 95
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/joinly-ai-joinly

## No authentication, client-supplied configuration, localhost only

The server mode is documented with a warning that is worth reading before anything else. The model context protocol server has no authentication and accepts configuration supplied by the client, so it is meant to run locally with a single trusted client, and the documentation says to bind it to loopback and not expose the port to a network. The launch line reflects that:

```bash
docker run -p 127.0.0.1:8000:8000 ghcr.io/joinly-ai/joinly:latest
```

Two consequences follow and neither is hidden. Any process that can reach that port can drive the session, and any client can change the server's configuration at runtime, including which model and which speech providers are used. That is a deliberate trade for a tool whose whole purpose is to be driven by an agent you control, and it is the reason the server mode exists separately from the bundled client mode rather than being a flag on the same process.

## The image is 2.3 GB because it carries a browser and models

Self-hosting here means pulling an image of roughly 2.3 gigabytes, and the documentation says why: it packages a browser and models. That is the honest cost of the design, since joining a meeting means driving a real browser rather than calling a meeting API, which is also how it reaches any platform that works over a browser. The dependency list confirms the shape of it. A browser automation library is pinned to an exact version, which is the single exact pin in an otherwise floating list. Local speech is present as well, with a faster-whisper package for transcription and an ONNX runtime paired with a Kokoro package for synthesis, alongside a voice activity detection wheel. So the image can run speech entirely on your own hardware, and the browser it drags along is the price of the meeting integration rather than of the model.

## Client mode and server mode cannot run at the same time

The image starts as a server by default, and there is a second way to use it. Passing the client flag starts a bundled conversational agent client directly instead, which is the quickstart path: the container is given a meeting URL and joins the call. The documentation is explicit about the trade. In client mode no server is started, so no other client can connect to it, and therefore none of the third-party model context protocol servers can be attached either. In server mode the container exposes the server on a port, you run a separate client package from your own machine with a tool runner, and that client can load additional MCP servers from a JSON configuration. The same configuration file format covers a local server launched as a subprocess and a remote server reached over HTTP with OAuth, and each entry becomes tools the agent can call inside the meeting.

## Every speech provider ships as a mandatory dependency, not an extra

The pitch is bring your own providers: whisper or Deepgram for transcription, Kokoro, ElevenLabs, or Deepgram for synthesis, and any language model including a local one. The configuration side lives up to that, with command line flags selecting speech to speech and text to speech providers separately, and the environment example carrying keys for Deepgram, ElevenLabs, and a third synthesis service. The packaging side does not. Both cloud speech SDKs are unconditional dependencies rather than optional extras, as are the local speech packages, so a deployment that uses one provider still installs the clients for the others, and the only declared optional extra is a GPU build of the ONNX runtime. The single optional dependency group in the project is for hardware acceleration. Everything else, including the browser and both speech stacks, is in the base image.

## Two variable names for the same model setting

The environment example covers five language providers, and they are not configured the same way. The OpenAI, Anthropic, and Azure blocks set a model variable and a provider variable under one naming pair. The Google block uses a different pair, a model-name key and a provider key without the intervening segment. Anyone copying the file and switching providers has to notice which block they are in, because the wrong key is silently ignored and the service falls back to whatever default it had. The file also assigns the model variable once per provider with a different value each time, so an unmodified copy resolves to the last block rather than to a configuration you chose, and the accompanying instruction to delete the placeholders for providers you are not using is doing real work. Azure additionally needs an endpoint and an API version, which the other providers do not.

## A misspelled key, and a local model warning

Two details in the configuration are easy to trip on. The key for the third synthesis service is spelled with the letters transposed, so a search of the codebase for the provider name will not find it and a search for the key name will. And the local model path carries three separate caveats in the example file: it requires the model to be pulled first, it only works through the example client or another script outside the container rather than through the bundled client flag, and small models often fail to call the tools correctly, with the example setting using a one-and-a-half-billion-parameter model. That last caveat is the important one. An agent in a meeting is only useful if it can call tools, and tool calling is exactly what small local models are worst at, so the offline configuration is the one most likely to underperform.

## Interruption handling is the feature, and it is proprietary logic

The capability that separates this from a meeting recorder is conversational flow: built-in logic for handling interruptions and multiple speakers, so the conversation does not stall when two people talk at once or when the agent is cut off mid-answer. The agent responds in real time by voice or chat while remaining in the call, and it can execute tasks rather than only talk, which the demos show by having it read current news from the web when asked what the product is and then create an issue in a demonstration repository, and separately by connecting to a hosted workspace through MCP and editing a page live during the conversation. That combination is the argument for the browser-based approach: nothing about meeting platforms exposes a clean tool-calling surface, so the agent has to be present the way a person is.

## Nine months between releases, and a repository built around its own tooling

The release cadence has a long gap in it. Version 0.5.2 shipped in November 2025, 0.5.3 in December, and then 0.5.4 came in September 2026, roughly nine months later, and the last commit matches that release date rather than running ahead of it. So the current state is a single recent tag with nothing after it. The tree itself is organised as a workspace with three members: the server package, a client package, and a shared package, and the server depends on the latter two from the workspace, which means a published install pulls them by version range rather than vendoring them. Alongside sit a development container definition, a Python version file, a lock file, a pre-commit configuration, and a directory and instructions file for a coding agent, so the project is set up to be worked on by one.

## Conclusion

Use it if you want an agent present in a meeting rather than a transcript afterwards, since live participation with interruption handling is the whole capability and it works across Meet, Zoom, and Teams. Two constraints to plan around. The server has no authentication and accepts client-supplied configuration, so it stays on loopback with one trusted client, never on a shared network. And the self-hosted claim comes with a 2.3 GB image that bundles a browser and models, which is a different commitment from a small container.

## FAQ

### What is joinly and what does it do?

It is open-source connector middleware that lets AI agents join and participate in video calls. An MCP server exposes meeting tools and resources so an agent can execute tasks and respond by voice or chat in real time during the meeting, across Google Meet, Zoom, Microsoft Teams, or any platform reachable over a browser.

### How do I run joinly?

Through Docker. Create a folder, put an environment file with your language provider key in it, pull the image from the container registry, then run the container with the environment file and a meeting URL using the client flag. The image is about 2.3 GB because it packages a browser and models.

### Is the joinly MCP server safe to expose on a network?

No, and the documentation says so in a warning. The server has no authentication and accepts configuration supplied by the client, so it is meant for local use with a single trusted client, bound to loopback, with the port not exposed.

### Can I use Ollama with joinly?

Yes, with caveats recorded in the environment example: you need to pull the model first, it only works through the example client or another script outside the container rather than through the bundled client flag, and small models often fail to call the required tools correctly.

### How do I give the joinly agent more tools?

Run the server mode, which publishes port 8000 on loopback, then start the client package from your own machine with a JSON configuration listing additional MCP servers under an mcpServers key. Each entry becomes tools available in the meeting, and the format covers both a local server launched as a subprocess and a remote server with OAuth.

## Sources

- [joinly-ai/joinly on GitHub](https://github.com/joinly-ai/joinly)
- [License: MIT](https://github.com/joinly-ai/joinly/blob/main/LICENSE)
- [Project website](https://joinly.ai)
- [README](https://github.com/joinly-ai/joinly/blob/main/README.md)
- [Releases](https://github.com/joinly-ai/joinly/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/joinly-ai-joinly
