Vexa: self-hosted meeting bots that stream speaker-attributed transcripts
Open-source meeting transcription API for Google Meet, Microsoft Teams & Zoom. Auto-join bots, real-time WebSocket transcripts, MCP server for AI agents. Self-host or use hosted SaaS.
At a glance
- What is it?
- Vexa is an Apache-2.0 stack that sends a bot into Google Meet, Microsoft Teams and Zoom calls and streams transcripts over WebSocket or HTTP. This review covers the Docker Compose install, the API, the agent plane, and where the project is still thin.
- Who is it for?
- Adopt Vexa if you want transcripts of Meet, Teams and Zoom calls to stay on infrastructure you control, or if you need a bot that joins a call and emits a live stream rather than a file after the fact. Do not adopt it if you need Jitsi in production today, since the README says live validation is still pending, or if you want a hosted agent plane, which the README states is self-hosted only.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Vexa solves, and who it is for
Most meeting AI products begin after a transcript exists. The vendor's bot joins the call, the audio leaves your network, and what comes back is a summary you rent access to. Vexa starts one step earlier. A real bot joins Google Meet, Microsoft Teams or Zoom as a participant, and the stack streams a speaker-attributed transcript while the meeting is still running. That bot fleet is the part the README calls "the genuinely hard part," and it is the reason the project exists.
The audience follows from that. If you are building a product that needs live meeting text, or you run a team that cannot send call audio to a third party, Vexa is aimed at you. The README is explicit that the transcription API is a complete standalone product: send a bot, read the stream, and ignore the agent side entirely. The agent plane, which compiles meetings into a Markdown knowledge base in a git repo and lets sandboxed coding agents read and write it, is a second product built on the same runtime.
The licence is Apache-2.0, and the README describes the deployment as air-gap-ready. That combination is the pitch: portable files, your own models, your own network.
One gateway, two domains, and a container per workload
The architecture the README describes is a single API gateway routing to two domains. Meetings handles capture. Agents works the resulting knowledge. Both sit on one runtime, which is the component that spawns every bot and every agent in its own sandboxed container.
The README states that a bot and an agent are the same runtime.v1 workload: isolated, ephemeral, reaped on idle. That is the design decision worth noticing. Rather than writing two orchestration layers, the project reuses the machinery that already spawns meeting bots to run coding agents. The README claims this machinery is proven by thousands of meeting bots, a figure that appears only in the README and cannot be verified from the repository layout.
Data flow for the capture path is straightforward. A client posts to /bots with a platform and a meeting code. The runtime starts a container. The bot joins the call, transcribes, and the transcript is readable over HTTP at /transcripts/{platform}/{native_meeting_id} or streamed live. The README mentions draft-then-confirmed segments, meaning interim text is replaced as the transcriber settles. On the agent path, replies stream as Server-Sent Events, with message-delta frames carrying text and commit frames marking anything written into the workspace.
One constraint is visible in the repository rather than the README. The top level holds contracts.seal.json, schema.seal.json and architecture.seal.json alongside a large set of gate scripts in package.json (gate:isolation, gate:contract-version, gate:schema, among many others). The project enforces its internal contracts mechanically. That is a real cost for contributors, and it is also why the runtime abstraction has stayed stable enough to serve both domains.
Installing Vexa with Docker Compose and sending a first bot
The README gives two paths. The hosted service needs no install: sign in, copy an API key, and post to the cloud endpoint. New accounts get $5 of bot credit, which the README says is about 16 hours at $0.30/hr. The self-hosted path is what gets you the agent plane, and it is the one worth walking through.
Prerequisites named in the README are make, Docker engine v26 or later, and a transcription credential. The default behaviour matters here: POST /bots requires speech-to-text and answers 503 when it is missing. You can get a token from the Vexa account page, or self-host the GPU transcription unit for a fully air-gapped setup. Capture-only is an explicit opt-out, either per spawn or per deployment.
git clone https://github.com/Vexa-ai/vexa.git && cd vexa
make allThe README states that make all pulls published, release-validated images rather than building, so a modest machine is enough. It seeds .env, pulls the images including the bot, and prints your API key and URLs. Contributors who want to build from the checkout use make dev instead, which the README says wants 8 vCPUs and 16 GB RAM. There is also make lite for a single-container image.
When it finishes you should see a block like this, with your own key in place of the placeholder:
Terminal UI : http://localhost:13000
API gateway : http://localhost:18056
API key : vxa_...The README calls the Terminal UI at port 13000 the fast path, and that is fair: you are already signed in to a self-host account, and you can paste a meeting URL to send a bot without writing any curl. If you prefer the API, the README's example exports the key and base URL, then posts a bot:
export API_KEY=vxa_...
export API_BASE=http://localhost:18056
curl -X POST "$API_BASE/bots" \
-H "X-API-Key: $API_KEY" -H "Content-Type: application/json" \
-d '{"platform":"google_meet","native_meeting_id":"abc-defg-hij","bot_name":"Vexa"}'The platform field accepts google_meet, teams, zoom or jitsi, and native_meeting_id is the code from the join URL. To read the transcript back, the README uses GET /transcripts/google_meet/abc-defg-hij with the same header. Expect a bot to appear as a participant in the call, and text to arrive in draft-then-confirmed segments.
Where Vexa is the wrong tool
Jitsi is the clearest boundary. The README says join and capture are offline-proven but live validation is pending, and it links issue #883. Treat Jitsi as unsupported for production until that changes. Meet, Teams and Zoom are the platforms the project actually stands behind.
The transcription dependency is the second boundary. Because POST /bots returns 503 without speech-to-text configured, a fresh self-hosted install that skips the credential step will look broken rather than degraded. You either point at the hosted token or run the GPU transcription unit yourself. Capture-only mode exists, but it is an opt-out you have to set deliberately, and it gives you no transcript.
Authenticated meetings are a third area the README does not settle. The Makefile has a login target that provisions an authenticated-bot session, signing in once and persisting it to userdata storage, with docs at /authenticated-bots. What is not described is how that session behaves when a platform changes its sign-in flow, or how it scales across many accounts. If your meetings require authentication, verify that path before you plan around it.
Finally, the agent plane is not in the hosted service. The README states that sandboxed knowledge agents are self-hosted only. If you wanted hosted transcription plus hosted agents from one vendor, Vexa does not offer that combination.
How Vexa differs from Recall.ai and from meeting summarizers
Recall.ai is the closest comparison in kind: a meeting bot API that joins calls and returns transcripts. The difference is where the stack runs. Recall.ai is a hosted service; you send it a meeting and it handles the bot fleet. Vexa ships the bot, the gateway and the runtime as an Apache-2.0 repository you can run on your own Docker or Kubernetes, with an air-gapped configuration the README describes. That is the trade: you take on the deployment and the transcription credential, and in exchange the audio and the resulting files stay inside your network.
The second comparison is against meeting summarizers such as the note-taker category generally. Those products start after the transcript exists and end at a summary document inside their cloud. Vexa's agent plane takes a different position: meetings compile into Markdown files in a git repo, and sandboxed agents read and write that repo like developers, in isolated ephemeral containers with no egress. The README frames this as knowledge as code, which is a reasonable description of the artefact even if the framing is marketing. The practical difference is that a summary is a document you read, while a Markdown repo is something you can diff, grep and feed to another tool.
Neither comparison makes Vexa strictly better. A hosted API removes an entire operational surface. A polished summarizer will have a better interface than a Markdown repo for most non-technical users.
Maintenance, releases and what the licence asks of you
The repository is not archived, and the last push was on 2026-09-07, so the project is being worked on now. Release cadence is fast: v0.12.26 on 2026-08-31, then v0.12.27 and v0.13.1 both on 2026-09-07. The v0.13.1 release is labelled "the minutes line (alpha)," which tells you the newest surface is not settled. Upgrading across those versions is where the cost sits.
The repository carries contracts.seal.json, schema.seal.json and architecture.seal.json, plus a long list of gate scripts in package.json covering isolation, exports, graph, schema, contract-version, config-contract, db-schema and db-budget. Those gates imply that internal contracts are versioned and checked, which should make upgrades more predictable than the version numbers alone suggest. They also mean a fork that changes an interface will have to deal with the seals.
On licensing, the root LICENSE is Apache-2.0. The repository also contains THIRD_PARTY_LICENSES.md, license-exceptions.json, image-licenses.json, a licenses/ directory, CLA/, CONTRIBUTOR_RIGHTS.md and a security/ directory. The presence of image-licenses.json is worth noting if you plan to redistribute container images, because the images may carry terms separate from the source licence. Apache-2.0 includes a patent grant and requires attribution and notice retention, but nothing here is legal advice; read LICENSE and THIRD_PARTY_LICENSES.md yourself before shipping a derivative.
Editorial conclusion
Adopt Vexa if you want transcripts of Meet, Teams and Zoom calls to stay on infrastructure you control, or if you need a bot that joins a call and emits a live stream rather than a file after the fact. Do not adopt it if you need Jitsi in production today, since the README says live validation is still pending, or if you want a hosted agent plane, which the README states is self-hosted only. Before committing, verify that your transcription path works end to end, because POST /bots returns 503 when speech-to-text is absent, and confirm that the bot can authenticate into whatever meeting platform you actually use.
Frequently asked questions
How do I install Vexa and run it self-hosted?
Clone the repository and run make all. The README states that this pulls the published, release-validated images, seeds .env, and prints your API key along with the Terminal UI at http://localhost:13000 and the API gateway at http://localhost:18056. You need make and Docker engine v26 or later.
Does Vexa work with Google Meet, Microsoft Teams and Zoom?
The README lists google_meet, teams, zoom and jitsi as the platform values accepted by POST /bots. Jitsi is described as join and capture offline-proven with live validation pending, so Meet, Teams and Zoom are the platforms the project stands behind.
Why does POST /bots return 503 on my Vexa install?
The README states that POST /bots requires speech-to-text by default and answers 503 when it is missing. Supply a token from the Vexa account page or self-host the GPU transcription unit, or set transcribe_enabled to false on the spawn to run capture-only.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/vexa-ai-vexa)