Model or dataset
p0n1/epub_to_audiobook avatar
p0n1/epub_to_audiobook

epub_to_audiobook writes one MP3 per chapter, and its two newest tags are 46 seconds apart

EPUB to audiobook converter, optimized for Audiobookshelf, WebUI included

2,072 stars224 forksPythonMIT

At a glance

What is it?
p0n1/epub_to_audiobook is a Python converter that turns an EPUB into per-chapter MP3 files shaped for Audiobookshelf, with five text-to-speech backends and a Gradio interface. Its in-file changelog holds one line, its homepage points at a generated wiki rather than a site, and the last push on main is dated 2026-03-24.
Who is it for?
This is worth using if you own the EPUB and want it on your own audiobook server, because the output shape is aimed at one specific player and the WebUI exposes the same options as the command line rather than hiding a subset. Three things to know before you start.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Two tags published 46 seconds apart, then nothing for seven weeks

The three most recent releases tell an odd little story. v0.8.5 is dated 2025-08-29. Then v0.8.6 is dated 2026-02-03 at 05:41:02 and v0.8.7 is dated the same day at 05:41:48. Those two are 46 seconds apart.

A patch version that follows another patch version 46 seconds later is not two rounds of work. It is one round of work plus a correction, most likely to packaging metadata or a tag that was created and then recreated, which is the usual reason for a release appearing and being superseded within a minute. Nothing in the repository says which it was.

After v0.8.7 the releases stop. The last push on the default branch is dated 2026-03-24, seven weeks later, with no tag cut after it. So there are three states to be aware of when you pin: the last release is v0.8.7 from February, the branch head is from late March, and whatever sits between them is not described anywhere except in the commit history.

The metadata licence is MIT and the primary language is Python. Nothing about the project is archived.

The in-file changelog has one line and stops before the last three releases

There is a section headed Recent Updates, and it contains exactly one entry: 2025-05-23, added a web interface to the project. That is the whole changelog.

It is worth putting beside the release dates. v0.8.5 came in August 2025, and v0.8.6 and v0.8.7 came in February 2026, all after the one line the file records. So the section a reader is most likely to consult for recent changes stopped months before the last release and has not been extended since.

It is also the only place the WebUI's arrival is dated, which makes the single line load-bearing in a way it should not be. Nothing in the file says which of the thirteen pinned dependencies arrived when, or when the Python 3.14 compatibility note was added, or when Piper and Kokoro became options.

The practical advice is short. Use the platform's release list rather than the readme's update list, and treat the readme as prose documentation rather than as a version history.

The disclosure sentence is in the file, inside a comment

Between the opening description and the Recent Updates section there is a line that says the project was developed with the help of ChatGPT. The whole line is wrapped in HTML comment markers, so it renders as nothing. It is present in the file and invisible on the page.

That is worth stating plainly rather than reading into, because it is a governance choice rather than an accident of prose. A reader of the rendered page sees no mention of it; a reader of the source sees it immediately and unambiguously, which is a different and weaker guarantee than stating it in the open. Neither is wrong on its own. What it does mean is that anyone auditing how the code was produced has to know to look at the markup.

A second small thing belongs in the same section. The homepage recorded for the project is not a site of its own but a page on a generated documentation service, and the title line carries badges for a chat server and that same documentation service. So the project's public face is a community server plus an auto-generated wiki, and there is no site here to speak of.

Chapter titles come from an HTML title tag or the first few words

The output shape is one MP3 per EPUB chapter, with the chapter title written into the file's metadata so that Audiobookshelf shows chapter names when you import the files. That is the entire design target, and it shapes every other decision.

Extracting those titles is the part the readme is most candid about. The method is to parse the EPUB and look for a title tag in each chapter's HTML content. If no title tag is present, a fallback title is generated from the first few words of the chapter text. The readme calls this simple but effective for most files, and then says directly that it may not work for EPUBs with complex or unusual formatting.

The pinned dependency list maps onto that pipeline neatly. A library reads and writes the EPUB, an HTML parser finds the title tag, a metadata library writes the title into each MP3, an audio library plus ffmpeg cut the audio, and a sentence splitter handles chunking. Thirteen entries, all pinned to exact versions, and each one is doing a job you can point at rather than a job you have to infer.

Getting them is three commands and nothing more:

bash
git clone https://github.com/p0n1/epub_to_audiobook.git
cd epub_to_audiobook
bash
python3 -m venv venv
source venv/bin/activate
bash
pip install -r requirements.txt

After which the credentials for whichever backend you picked are exported into the environment:

bash
export MS_TTS_KEY=<your_subscription_key> # for Azure
export MS_TTS_REGION=<your_region> # for Azure
export OPENAI_API_KEY=<your_openai_api_key> # for OpenAI

Five backends, and the options change shape with each one

The engine list is not consistent across the readme, and the differences are informative. The introduction names two engines, Azure with Edge as an alternative, and OpenAI. The web interface offers four tabs. The audio samples demonstrate five.

The fifth is Kokoro, and it is wired in a way worth understanding rather than guessing at: it is used through a local OpenAI-compatible endpoint, which is why the samples point at a hosted demo and why the interface exposes it in the OpenAI tab. That also explains an odd requirement, which is that you must set a dummy OpenAI key such as a placeholder value even when you are not calling OpenAI, because the code checks for the variable rather than for a valid key. The compose file is the exception.

The per-engine option lists differ in a way that shows how little they share. Azure offers language, voice, format and a break duration. Edge offers language, voice, rate, volume and pitch. OpenAI offers model, voice, speed and format. Piper offers local or Docker deployment with voice options, and it is the only one with a separate install requirement: the Piper executable and its models, outside the Python dependency list.

A base URL variable documented only under the web interface

Two environment variable lists appear in the readme, and they are not the same length. The command-line installation section sets a speech service key, a region, and an OpenAI key. The web interface section lists those three and adds a fourth for a custom OpenAI-compatible endpoint, and then says the interface respects the same variables as the command-line tool.

So the one variable that makes a local endpoint possible is only documented in the section that claims to be a subset. If you are running the converter from a terminal against a local model server and you are looking at the command-line section, the setting you need is not there. That is a documentation gap rather than a missing feature, since the interface plainly supports it.

The deployment options are similarly spread across files. There are three compose files in the repository: a general example, one specific to Kokoro, and one for the web interface. The readme refers to a compose file for the dummy-key exemption without the visible text reaching the section that would show it, so the exemption is real but its instructions are past the point where this text ends.

The container splits two directories so a bind mount wins

The Dockerfile does something clever and explains why in the comment. It builds in a source directory rather than the application directory, because users are expected to mount their own copy over the application directory with a bind mount, and if the image put its own files there the mounted ones would be shadowed. So the build happens in one directory and the runtime working directory is the other, which means whatever the user mounts takes precedence.

The same file carries two more admissions worth reading. Output buffering is disabled with an environment variable, and the comment says the reason is to ensure print statements appear in the logs, which tells you the application's logging is print statements rather than a logging library. And the only system package installed is ffmpeg, which the audio library needs, with the package lists cleaned afterwards.

One consequence of the bind-mount design is not addressed anywhere: the image creates no unprivileged user, so a container started with a mounted project directory is running that mounted code as root. The web interface compounds it, since you can bind it to any host and port you like and stop it with a keyboard interrupt, and no authentication option is described anywhere for it.

Editorial conclusion

This is worth using if you own the EPUB and want it on your own audiobook server, because the output shape is aimed at one specific player and the WebUI exposes the same options as the command line rather than hiding a subset. Three things to know before you start. Chapter titles come from an HTML title tag with a fallback to the first few words, so a book with unusual formatting will produce odd filenames, and that limitation is stated in the readme rather than discovered later. Every one of the thirteen dependencies is pinned to an exact version, which is reproducible and also means you inherit those pins rather than current ones. And check where the project stands before you plan around it: the newest tag is v0.8.7 from 2026-02-03 and the last push on main is dated 2026-03-24, so the branch has moved since the last release without a new one.

Frequently asked questions

What does epub_to_audiobook produce?

One MP3 file per EPUB chapter, with the chapter title written into the file's metadata, shaped so that the files import into Audiobookshelf with chapter names attached.

How does epub_to_audiobook decide the chapter titles?

It parses each chapter's HTML and looks for a title tag. If there is none, it falls back to the first few words of the chapter text. This works for most EPUB files but can fail on ones with complex or unusual formatting, as the readme notes.

Which text-to-speech engines can epub_to_audiobook use?

The web interface offers Azure, OpenAI, Edge and Piper, and the audio samples also cover Kokoro, which is driven through a local OpenAI-compatible endpoint. Azure and OpenAI need credentials, Edge needs no key, and Piper needs its own executable and models alongside the Python packages.

Do I need an OpenAI API key to use Kokoro with epub_to_audiobook?

Not a real one. Kokoro runs through a local OpenAI-compatible endpoint, but the code still checks for the key variable, so a dummy value has to be put in the environment, for example a placeholder string, unless you are using the docker compose file.

How do I run the web interface of epub_to_audiobook?

python3 main_ui.py starts the Gradio interface at http://127.0.0.1:7860, and python3 main_ui.py --host 127.0.0.1 --port 8080 changes the host and port. It is stopped with Ctrl+C, and no authentication option is described for it.

Why does installing epub_to_audiobook fail on Python 3.14?

Because earlier installs pinned gradio at 5.33.1, which could force a pydantic-core source build and fail with a PyO3 compatibility error. The requirements file now pins gradio at 5.50.0, and Python 3.14 needs that updated dependency set. Every entry in the file is pinned to an exact version.

Official sources

  1. License: MIT
  2. p0n1/epub_to_audiobook on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/p0n1-epub-to-audiobook.svg)](https://hysenlabs.com/projects/p0n1-epub-to-audiobook)