Open-source project
bugbakery/audapolis avatar
bugbakery/audapolis

audapolis: a wordprocessor for spoken-word media, with no cloud and no named transcription engine

an editor for spoken-word audio with automatic transcription

1,895 stars63 forksTypeScriptAGPL-3.0

At a glance

What is it?
audapolis is a free desktop editor for spoken-word media that gives you a wordprocessor-like editing experience and transcribes your audio automatically, with a stated commitment to keeping everything on your own machine. The material is thin in an instructive way: it names four funding bodies and three desktop platforms, and it never once says which speech-to-text engine does the transcription or whether that engine runs offline.
Who is it for?
audapolis is worth a try if you cut interviews, podcasts or audiobooks by sentence rather than by waveform, because a transcript-driven editor is a genuinely different tool from a timeline editor and it is the right shape for material where what matters is which words survive. It is the wrong choice for anything sync-dependent, for music, or for cutting on a breath rather than on a sentence, and no amount of transcription quality fixes that.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository last received commits 99 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A wordprocessor-like experience, which means the transcript is the editing surface

The product idea fits in one line of the readme and everything else follows from it. The editor is described as giving a wordprocessor-like experience for media editing. Read that together with the next claim, that it can automatically transcribe your audio to text, and the design becomes clear: the transcript is not a companion feature or a subtitle track, it is the interface you edit through. You read the words, you select and delete and reorder them the way you would in a document, and the audio follows. That is a different model from a timeline, where the unit of work is a clip you position by ear and by eye, and the difference is not cosmetic. In a timeline, finding the sentence you want to remove means scrubbing and listening. In a transcript, it means reading. For the material the readme lists, which is radio shows, podcasts, audiobooks and interview clips, the second is dramatically faster, and that is the workflow the project says it exists to make faster and more accessible. The readme also states the tool handles video, audio and mixed editing, so the same cut can span a camera shot and a voice track, which matters for an interview clip where you want the picture and the answer to be the same edit. Two things follow from the model that are worth holding onto. First, the quality of the transcript is the quality of the editor, because every cut you make is a cut in the text. Second, the unit of editing is a sentence or a clause, so anything that has no words in it is outside what this tool is good at, and the readme's own list of uses is the honest boundary.

No cloud, and a transcription engine the readme never names

The privacy claim is one of the bullet points and it is absolute: the data stays in your hands, no cloud whatsoever. For a tool that takes your interview recordings and your podcast masters, that is the right default and it is the reason to consider this project at all. But notice what the claim covers and what it does not. It is a claim about storage, about where your media lives, and it says nothing about where the speech recognition happens. And the readme never says. There is no model named, no engine named, no statement about whether recognition runs on your machine, and no mention of a network connection anywhere in the document. The topic tags include speech to text, so the feature is real, and the absence of an engine name is a documentation gap rather than a sign that there isn't one. This is the first question to answer before you adopt the tool, and it is answerable only by running it with the network off, or by reading the source, because nothing in the readme settles it. The repository gives one indirect clue. The top-level file list contains a Python linter configuration in a project whose primary language is TypeScript, which means there is Python code in the tree that the readme never mentions, and the most likely place for a speech recognition engine to live is exactly there. That is an inference from a filename, not a finding, and the difference matters: if the engine is a Python process wrapping a downloadable model, the first run may need a download, which is a different proposition from a tool that never touches a network. If it bundles everything, the installer is large. Either way you should know which before you promise anyone their recordings never leave the machine.

Download a binary for three platforms, and that is the entire install story

The installation section of this readme is four lines long and its substance is a link to the latest release, where binaries exist for Windows, Linux and macOS. There is no package manager entry, no container image, no command, no configuration file to edit and no list of prerequisites. The readme also asks for bugs and usability problems to be reported, which tells you the project expects people to install the build rather than the source. For a home user or a journalist that is exactly right and there is nothing to criticise. For an organisation, it produces three questions the readme does not answer. Where does the binary come from and how would you verify it, because a desktop application distributed only as a compiled artefact is something your security process will want a provenance story for, and the readme here offers none. Can you build it yourself, because if the answer is no then you are dependent on the project's release cadence rather than on your own, and the release record in the next section suggests that cadence is slow. And how do you deploy updates, since a media editor on a handful of machines is a manual operation unless you have a way to push a new binary. The top-level file list gives a partial answer to the second question: there is a workflow directory, which is where a build and a multi-platform release pipeline would live, and there is no build manifest at the top level, so the manifests are inside the application and server directories or they are generated. Neither answer is complete, and both are the kind of thing you find out by opening the repository rather than by reading the readme.

Funded for six months by three German bodies, then kept alive for four years

The acknowledgements section is one line and it is the most informative line in the readme. The project was funded from September 2021 until February 2022 by the German federal ministry for education and research, the Prototype Fund, and Open Knowledge Foundation Germany. Read what that combination means. The first is a national ministry, so this is not a hobby project with a nice idea; it was funded as research or as a prototype in the German public-interest sense. The second is a fund whose whole purpose is funding prototypes of socially useful software, usually with a civic or educational end. The third is the German chapter of a foundation for open knowledge, which is the kind of funder that attaches conditions about openness and reuse. So the project's funding profile tells you its mission was never commercial, and the readme says so in as many words, describing the project as free with the data kept locally and no cloud at all. What the funding dates also tell you is when the money stopped. The window closed in February 2022. The last push to the repository was on 2026-06-24. That is four and a half years of development with no funder named, which means either someone is doing this in their own time, or there is funding that is not acknowledged here, and either way the sustainability question is the one to have an answer to before you depend on the project. The honest framing for an evaluator is that this is a publicly funded prototype that reached a usable state and has been maintained since by whoever chose to, which is a better position than an abandoned experiment and a weaker one than a maintained product.

Two years between the last two releases, and a survey asking what to build next

The release record tells the same story from the other side. The tagged releases are a pre-release at the zero point two point two stage from April 2023, a zero point three point zero from July 2023, and a zero point three point one from November 2025. So there is a convention of shipping pre-releases, there was a gap of about two and a half years between the minor and the patch release, and the version line is still below one. The pre-release naming is worth a small note of its own, because it tells you the project is comfortable shipping something labelled as not finished, which is honest and is also a signal about stability. The survey reinforces the picture. The readme asks users to fill in a short form about their needs and expectations, and says the point is to build software that is actually useful and to know what you need. A project asking that question is a project that does not yet have a strong sense of its user base, and that is a fact you should weigh rather than a criticism of the maintainers. Put the three signals together, the funding that ended in 2022, the release gap, and the survey, and the picture is of a project with a clear and valuable idea, an unusual amount of documentation about its own status, and no institutional backing. The practical consequence is that you should decide whether to use the current build or to wait, and the honest answer for a spoken-word workflow is to try the build, because the idea is good enough that the tool's usefulness does not depend on it being finished. If it fits your material, the cost of a later migration to something else is your transcript, which you already have as text.

AGPL-3.0 on a local application, and contribution links that point elsewhere

The licence is the GNU Affero General Public License, and for this particular kind of software the practical reach of that licence is narrower than the reputation suggests. A desktop editor that runs on your own machine and talks to nobody is not offering a network service, and the network clause of that licence is about modified versions made available to users over a network. So if you install the released build and use it internally, the obligations are the ordinary attribution ones. The interesting case is if you run a modified instance as a service for other people, which is a scenario a civic project might well end up in, and that is where the copyleft reaches. A publicly funded project choosing that licence is making a deliberate statement that modifications should come back, and the choice is coherent with everything else about the project. Two small documentation inconsistencies are worth noting because they are the kind of thing that tells you how closely the readme is maintained. The contributing guide and the code of conduct are both linked, and both links point at a different repository, an umbrella organisation one directory up, rather than at this project. Both files are also present in this repository's own top-level list, so the files are here and the links are not pointing at them. And the top-level list contains no build manifest of any kind, no package file, no requirements file and no container definition, which is unusual for a project that publishes binaries for three platforms and tells you how to report bugs.

The alternative: a transcription tool plus a general editor, and which clips it suits

The honest comparison is not with another wordprocessor, because there is not one worth naming. It is with doing the same job in two stages with dedicated tools. Transcribe the audio properly in a tool built for transcription, correct the text there where the correction interface is designed for the job, export it, and then cut the media in a general editor using that text as a reference. That route gives you two things audapolis does not try to give you. The first is transcript quality, because a dedicated transcription tool is likely to have a better correction interface, speaker labelling and a model choice you can make. The second is editorial control, because a general video and audio editor gives you frame accurate cutting, effects, levels, crossfades and the ability to cut on something other than a word. In exchange you give up the single most valuable property of this project, which is that the cut and the transcript are the same operation, and you do the work twice. So the decision is a property of your material rather than of the tools, and the readme's own list of uses is the best guide to it. Radio, podcasts, audiobooks and interview clips are all sentence shaped. What you are removing is a sentence, a clause or a whole answer, and the boundaries fall in places where a reader would also put a paragraph break. Against that, consider anything where the edit is not linguistic. Cutting a clip to a music bed, trimming silence to a beat, aligning a clip to a visual cue, cutting mid-word because the speaker stumbles, or assembling a montage where the cut is rhythmic rather than semantic. None of those are things a transcript can express, and in a tool whose unit of editing is a word they are not awkward, they are unavailable. Use audapolis where the edit is the edit. Use a timeline where the edit is the picture.

Editorial conclusion

audapolis is worth a try if you cut interviews, podcasts or audiobooks by sentence rather than by waveform, because a transcript-driven editor is a genuinely different tool from a timeline editor and it is the right shape for material where what matters is which words survive. It is the wrong choice for anything sync-dependent, for music, or for cutting on a breath rather than on a sentence, and no amount of transcription quality fixes that. Two things have to be established before you commit to it. Which transcription engine it uses and whether that engine works with the network disconnected, because the no cloud claim is about where your media lives and says nothing about where the recognition happens, and the readme does not address it. And whether you can build it, because the install path is a downloaded binary for three desktop platforms with no package, no container and no build instructions in the readme, so if you need a recent change or an internal build you are reading the source and the workflows yourself. On maintenance, the honest read is a small publicly funded project that ran out of funding in early 2022 and has been kept alive since, with a two year gap between its last two releases.

Frequently asked questions

What does audapolis actually do differently from a normal editor?

It gives a wordprocessor-like editing experience, and it transcribes the audio automatically. Together those mean the transcript is the editing surface: you select and delete text the way you would in a document and the audio is edited with it. The readme lists radio shows, podcasts, audiobooks and interview clips as the intended material.

Which speech-to-text engine does it use, and does it work offline?

The readme does not say. It states that no cloud is involved and that the data stays on your machine, which is a claim about storage, and it names no recognition engine and says nothing about network access during transcription. The repository contains a Python linter configuration in a project whose primary language is TypeScript, which hints at a Python component, but that is an inference from a filename.

How do I install it?

By downloading the newest build for Windows, Linux or macOS from the releases page. There is no package manager entry, no container image and no build command in the readme. The repository does contain a workflow directory and no build manifest at the top level, so building it yourself means reading the repository.

Is the project still being worked on?

It was last pushed on 2026-06-24 and the newest release is v0.3.1 from 2025-11-26, which came about two and a half years after the previous one at v0.3.0 in July 2023. The stated funding ran from September 2021 to February 2022, so the maintenance since then has no funder named in the readme.

What does the AGPL-3.0 licence mean for a local desktop app?

For a local installation the practical obligations are attribution. The licence's network clause reaches a modified version that you make available to other users over a network, so it matters if you run a modified instance as a service for other people rather than if you run the released build on your own machine.

Official sources

  1. bugbakery/audapolis on GitHub
  2. Issues
  3. License: AGPL-3.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/bugbakery-audapolis.svg)](https://hysenlabs.com/projects/bugbakery-audapolis)