Self-hosted service
murtaza-nasir/speakr avatar
murtaza-nasir/speakr

Speakr: a self-hosted transcription and note-taking stack you run yourself

Speakr is a personal, self-hosted web application designed for transcribing audio recordings

4,047 stars341 forksPythonAGPL-3.0

At a glance

What is it?
Speakr is a Flask and Docker application that turns recordings into searchable notes, with WhisperX, OpenAI or AssemblyAI behind it. It is built for people who want the transcripts to stay on their own server, and it costs you a deployment to keep alive.
Who is it for?
Adopt Speakr if you already run Docker and want transcripts, diarization and cross-recording search on hardware you control, and if you accept tracking an alpha release line. Do not adopt it if you need a hosted service with no server to maintain, or if you cannot supply an ASR backend, since the app is a front end to one.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Speakr solves, and who ends up running it

A meeting recording is only useful if you can find the sentence again three months later. Speakr's answer is a single self-hosted web application that takes audio in, transcribes it through an engine you choose, and keeps the resulting transcript, summary, tags and notes in your own database. The README frames the audience directly: "privacy-conscious groups and individuals", running "entirely on your own infrastructure". That is a narrower group than it sounds. You need a machine that stays up, a place to put the audio, and an ASR backend reachable from the app. In exchange, nobody else holds the recordings. The feature list is unusually broad for a project this size: capture from microphone or system audio, a watched import folder, diarization, voice profiles, per-recording chat, and library-wide semantic search. The breadth is the point and also the risk, because each of those pieces depends on a backend that Speakr does not ship.

The pipeline: capture, transcribe, understand, organize

Speakr is a Flask application (the requirements pin flask 3.1.0, flask-sqlalchemy, flask-login and flask-wtf) with server-rendered templates and Vue loaded from a CDN at runtime. The repository's package.json says so explicitly: Vue components "are still rendered server-side by Flask and Vue is loaded from CDN at runtime", and the npm package exists only to host Vitest for pure-helper modules under static/js/modules/utils. So the architecture is a conventional server-rendered app with pockets of client-side behaviour, not a separate single-page frontend.

The transcription step is deliberately pluggable. The README lists self-hosted WhisperX, OpenAI, Mistral / Voxtral, AssemblyAI, OpenASR, Alibaba FunASR, or a custom ASR webservice, and states that "the right connector is auto-detected from your configuration". That design has a consequence worth stating plainly: several advertised features are backend-conditional. Speaker diarization works through WhisperX or OpenAI's diarizing models. Voice profiles, which recognize the same person across recordings via embeddings, require the WhisperX backend. Custom vocabulary and hotwords are, in the README's wording, "most effective with the WhisperX backend". If you point Speakr at a plain transcription endpoint, you get transcripts and lose the speaker layer.

Above the transcript sit the understanding features: automatic summaries with prompts customizable per recording, tag or folder; event and action-item extraction; per-recording chat whose answers cite timestamp chips that jump playback; and Inquire Mode, a semantic search across the whole library. Inquire Mode has an opt-in beta agentic mode that iteratively searches and reads recordings, showing each step and citing numbered sources. The README is specific about the permission boundary here: transcripts are always readable, summaries and private notes only if you allow them.

Installing Speakr from the Docker image and running a first transcription

The README points at Docker Hub under the image learnedmachine/speakr, and the repository carries a Dockerfile that builds in stages: a Python 3.11-slim builder installs requirements.txt against constraints.txt, optionally adds requirements-embeddings.txt when LIGHTWEIGHT is not 0, and downloads vendor JS, CSS and font assets through scripts/download_offline_deps.py. A second stage fetches static FFmpeg binaries from BtbN/FFmpeg-Builds. The Dockerfile comments explain why that mirror replaced the older johnvansickle builds: the latter is described as frozen at 7.0.2 and therefore shipping the MagicYUV decoder flaw CVE-2026-8461, fixed upstream in 8.1.2. The build pins a dated release rather than a rolling tag and verifies each architecture tarball's SHA-256.

The Dockerfile installs the Python dependencies with pip against the constraints file, which is the step that decides what the image contains:

dockerfile
COPY requirements.txt requirements-embeddings.txt constraints.txt ./
RUN pip install --no-cache-dir --prefix=/install -c constraints.txt -r requirements.txt

The same stage downloads the vendor assets, so an offline or air-gapped build still gets its JavaScript, CSS and fonts:

dockerfile
COPY scripts/download_offline_deps.py scripts/
RUN pip install --no-cache-dir requests && \
    PRODUCTION=${PRODUCTION} python scripts/download_offline_deps.py

Once the container is running, open the app in a browser and complete the first-run setup, then add your ASR backend in the settings so the connector can be auto-detected from your configuration.

For a first real use, drag an existing audio file onto the interface rather than recording live. The README describes drag and drop as one of the input paths, alongside microphone and system or browser-tab capture. Once the file is processed you should see a transcript with synced playback: clicking a line jumps to that moment, and the chat-style bubble view is available as an alternative layout.

The LIGHTWEIGHT build argument controls whether the embeddings requirements are installed at all. Setting it to 1 skips requirements-embeddings.txt, which is the file that would carry the semantic-search dependencies. Expect the library-wide Inquire Mode to be the part that suffers. The README does not document a rollback path for a container built this way.

Where Speakr gets in your way

The release line is the first thing to notice. The newest tag in the repository is v0.10.6-alpha, published on 2026-09-19, and the two before it, v0.10.5-alpha and v0.10.4-alpha, landed on 2026-09-01 and 2026-08-31. An alpha version number on every release is an honest signal about what you are installing. The project is not archived and the last push was on 2026-09-19, so work is happening, but you are tracking a fast-moving alpha rather than a frozen product.

Dependencies are the second constraint. requirements.txt pins numpy 1.24.3, scikit-learn 1.3.0 and scipy below 1.15, and holds authlib below 1.7.0 with a comment explaining that authlib 1.7.x mis-handles the client_id claim placement for some OpenID Connect providers, producing `unsupported_header: Unsupported {'client_id'} in header` on the SSO callback. That pin protects you from a known upstream bug and also holds you back from the fixed line until it lands. It is the kind of constraint you inherit when you self-host.

The third limitation is structural rather than a bug. Speakr is a front end to a transcription engine, and the interesting features sit on top of a specific one. Choose WhisperX and you take on hosting WhisperX with its GPU or CPU cost. Choose a simpler endpoint and diarization, voice profiles and hotword biasing degrade or disappear. The README never claims otherwise, but the feature list reads as though all of it is always available, and it is not.

Finally, a single-user setup is not where the design pays off. Groups, group-scoped tags that auto-share recordings, granular view/edit/reshare permissions, OIDC single sign-on and per-user LLM and transcription budgets are all aimed at teams. Running Speakr for one person means running a multi-user application and using a fraction of it.

Speakr against Scriberr and the hosted transcription apps

The closest comparison in the search data is Scriberr, another self-hosted transcription tool. The practical difference is scope. Speakr's README describes a system built around a library: folders, bulk operations, smart tags that carry their own AI prompt and ASR settings and stack when combined, retention policies with per-recording protection from cleanup, automated export, groups, public links and a REST API v1 with Swagger UI. That is a knowledge base with a transcription step inside it. If what you want is a file in and a transcript out, that apparatus is overhead you will maintain and mostly not use.

The other comparison is a hosted service. A cloud transcription product removes the server, the Docker image, the FFmpeg binaries and the authlib pin from your life, and it will usually be cheaper than the electricity and attention a self-hosted stack costs. What it does not give you is the README's central promise, that recordings stay on your infrastructure. Speakr's answer to that promise is concrete: OIDC single sign-on against Keycloak, Azure AD, Google, Auth0 or Pocket ID, admin-controlled public links, HMAC-signed and SSRF-guarded webhooks, and usage budgets per user. Those are controls you would otherwise buy, not build.

Licence, upgrade cost and what you are signing up for

Speakr is licensed AGPL-3.0, per the README badge and the LICENSE file in the repository root. The practical implication for a company is that AGPL reaches network use: if you modify Speakr and let users interact with it over a network, the licence's source-availability terms are the ones to read, and the repository carries a CLA.md, which means contributions are handled under a contributor agreement. This is not legal advice; if Speakr ends up inside a product you sell, that is a question for someone qualified to answer it.

Upgrade cost is mostly dependency management. The pinned authlib ceiling means an SSO-dependent deployment is waiting on an upstream fix before it can move. The FFmpeg stage pins a dated BtbN release and verifies SHA-256, and the Dockerfile notes that BtbN deletes autobuild assets after roughly two weeks, so a rebuild of an old commit can fail when the pinned tarball is gone. The release cadence, three alpha tags in under three weeks, means you should decide in advance whether you upgrade on every tag or hold a known-good image. The README does not document a migration or rollback procedure, so back up the volume that holds your data before you pull a new tag.

Editorial conclusion

Adopt Speakr if you already run Docker and want transcripts, diarization and cross-recording search on hardware you control, and if you accept tracking an alpha release line. Do not adopt it if you need a hosted service with no server to maintain, or if you cannot supply an ASR backend, since the app is a front end to one. Before committing, run the Docker image, point it at a WhisperX endpoint, and confirm that diarization, voice profiles and custom vocabulary behave as the documentation describes on your own audio.

Frequently asked questions

What is Speakr and who is it for?

Speakr is a personal, self-hosted web application for transcribing audio recordings and turning them into organized, searchable notes. The README describes it as built for privacy-conscious groups and individuals who want recordings to stay on their own infrastructure.

Which transcription engines can Speakr use?

The README lists self-hosted WhisperX, OpenAI, Mistral / Voxtral, AssemblyAI, OpenASR, Alibaba FunASR, or a custom ASR webservice, and says the right connector is auto-detected from your configuration. WhisperX is the recommended backend because it enables speaker diarization and voice profiles.

How is Speakr installed?

The README links to the learnedmachine/speakr image on Docker Hub, and the repository includes a multi-stage Dockerfile that installs the Python requirements and downloads static FFmpeg binaries. There is also documentation under the project's GitHub Pages site.

Does Speakr work without WhisperX?

It runs against other backends, but the README ties several features to WhisperX specifically: voice profiles require it, and custom vocabulary and hotwords are most effective with it. Diarization is available through WhisperX or OpenAI's diarizing models.

Official sources

  1. Issues
  2. License: AGPL-3.0
  3. murtaza-nasir/speakr on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/murtaza-nasir-speakr.svg)](https://hysenlabs.com/projects/murtaza-nasir-speakr)