# abogen: turning EPUBs and PDFs into audiobooks with synced captions

> abogen is a Python tool that reads EPUB, PDF, text, markdown and subtitle files and writes audio plus matching subtitles through Kokoro-82M. It installs from PyPI or with uv, needs espeak-ng on every platform, and runs as a desktop app or a Flask web UI on port 8808.

**denizsafak/abogen** — Generate audiobooks from EPUBs, PDFs and text with synchronized captions.

- Repository: https://github.com/denizsafak/abogen
- Website: https://pypi.org/project/abogen/
- Stars: 6,076 · Forks: 472
- Language: Python
- License: MIT
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/denizsafak-abogen

## The gap abogen fills between an ebook file and a listenable track

Most text-to-speech pipelines assume you already have clean plain text. Books do not arrive that way. An EPUB is a zip of XHTML with navigation metadata, a PDF is a layout format with no reading order guarantee, and a subtitle file is timed fragments. abogen's job is to absorb those formats, decide what the reading order is, and emit audio plus a subtitle file that lines up with it. The pyproject.toml lists ebooklib and beautifulsoup4 for EPUB parsing, PyMuPDF for PDF, and Markdown for markdown input, which tells you the project treats format handling as its own layer rather than a preprocessing chore you are expected to solve first.

The audience is narrower than the README's marketing line suggests. The dependency list includes PyQt6, pygame, Flask and gpustat. That is a desktop application with a bundled web server and GPU monitoring, not a library you import into a service. If you want to call a function and get bytes back, abogen is not shaped for that; if you want to drop a book in, pick a voice, and get a folder of chapter files, it is.

## How abogen turns a book into chapters, audio and captions

The pipeline is visible in the dependency set rather than in a published architecture document. Parsing happens first: ebooklib and beautifulsoup4 for EPUB, PyMuPDF for PDF, the stdlib plus Markdown for text and markdown. The demo folder contains demo.ass alongside demo.wav and demo.txt, which shows the output side is a subtitle format, not just raw audio.

Synthesis runs through Kokoro-82M, pulled in as kokoro>=0.9.4, with misaki[zh] for Chinese text handling and supertonic>=0.1.0 as a second engine in the dependency list. num2words and spacy sit between parsing and synthesis; spacy is a natural language pipeline and num2words expands numerals, so the text is normalized before it reaches the voice model. That normalization step is also where the optional LLM integration attaches. The .env.example exposes ABOGEN_LLM_BASE_URL with the comment that the server root is supplied and /v1 is appended automatically, and the default value points at http://localhost:11434, the Ollama port.

Chapter handling is a first-class concept, not an afterthought. The demo directory carries a chapter_marker.png asset, and the pyproject keywords list chapter-management. Audio is assembled per chapter, which is why the Docker volume layout separates /data/outputs from /data/cache: rendered files land in outputs, scratch audio conversion files land in cache.

The web UI is a Flask application. docker-compose.yaml states the Web UI provides a browser-based interface for audiobook generation and that it is reachable at http://localhost:8808, with ABOGEN_HOST set to 0.0.0.0 and ABOGEN_PORT set to 8808 inside the container. The desktop path is PyQt6, and pygame is present for playback.

## Installing abogen and generating a first audiobook

Every platform needs espeak-ng installed separately first. On Windows the README points at the espeak-ng latest release and says to download and run the .msi file. On macOS it is brew install espeak-ng, and on Debian or Ubuntu it is sudo apt install espeak-ng.

The recommended install path is uv. The README gives a separate command per accelerator, and these are not interchangeable:

```bash
# NVIDIA GPUs (CUDA 12.8)
uv tool install --python 3.12 abogen[cuda] --extra-index-url https://download.pytorch.org/whl/cu128 --index-strategy unsafe-best-match

# AMD GPUs or no GPU
uv tool install --python 3.12 abogen
```

The extra-index-url and index-strategy flags matter. Without unsafe-best-match, uv's resolver can refuse the mixed index setup. On macOS the README adds a --with clause pinning numpy<2 and pulling Kokoro from git, because the released Kokoro package lacks MPS support.

For a container instead of a local install, docker compose up --build builds the Flask web UI and serves it on port 8808. The compose file mounts four host paths: settings at /config, outputs at /data/outputs, cache at /data/cache, and uploads under /data/uploads. Copy .env.example to .env first if you want to change where those land.

```bash
cp .env.example .env
docker compose up --build
```

The .env.example file documents ABOGEN_UID and ABOGEN_GID defaulting to 1000:1000, with the note to find your own values via id -u and id -g. Getting this wrong is the most common way to end up with a container that cannot write into the mounted output directory. Once the UI is up at http://localhost:8808, the workflow is to upload a file, choose a voice, and start the queue; a queue.png asset in the demo folder suggests jobs are batched rather than run one at a time.

## Where abogen breaks down or is simply the wrong tool

The Python version range is the first hard boundary. pyproject.toml declares requires-python = ">=3.10, <3.13". Python 3.13 is excluded on every platform, even though the macOS uv line in the README asks for --python 3.13. That contradiction is worth resolving before you build anything, because a resolver failure on macOS is the likely outcome.

AMD GPU acceleration only exists on Linux. The README states plainly that ROCm is not available on Windows, so an AMD card on Windows falls back to CPU. The pip path for AMD on Linux is also more invasive than the uv path: it instructs you to pip3 uninstall torch and then install a nightly ROCm build, which means abogen is no longer running against the torch version its own metadata resolved.

The pip path for NVIDIA carries a pinned downgrade. The README says PyTorch 2.8.0 must be used until a linked PyTorch issue is fixed, with torchvision 0.23.0+cu128 and torchaudio 2.8.0. If you already have a newer torch in the same environment, that pin will fight you.

The LLM normalization feature depends on an external server. The default ABOGEN_LLM_BASE_URL is http://localhost:11434, and the comment explains that /v1 is appended. If nothing is listening there, the feature has no backend; the README does not describe a fallback for that case.

Finally, the README does not document rollback, and it does not state how partial output is handled if a job is interrupted mid-book. For a single short file that is irrelevant. For a 400-page PDF it is the difference between resuming and starting over.

## abogen against Calibre and Piper

Calibre is the obvious comparison for the input side. It converts between ebook formats and has a mature, long-tested EPUB and PDF handling layer, but it does not synthesize speech. If your problem is a malformed EPUB that will not parse, Calibre is the tool that fixes it, and abogen is downstream of that fix. The two compose: repair in Calibre, narrate in abogen.

Piper is the closer comparison on the output side. It is a standalone neural TTS binary with ONNX voice models, and it is designed to be called from a script with text on stdin and a WAV on stdout. The difference in approach is architectural. Piper gives you a synthesis primitive and nothing else; you own the EPUB parsing, chapter splitting, text normalization and subtitle timing. abogen gives you the whole chain, but in exchange it wants a PyQt6 desktop app or a Flask server, not a subprocess call. If you are building an automated pipeline inside a larger service, Piper's shape is easier to embed. If you want a finished audiobook folder without writing the glue, abogen is the shorter path.

The voice model choice follows from that split. abogen is built around Kokoro-82M, and the dependency list also names supertonic>=0.1.0 as a second engine. Piper's model set is separate and larger in language coverage. Neither is a drop-in for the other.

## Licence, maintenance and what an upgrade costs you

abogen is MIT licensed, stated in both the README badge and pyproject.toml as license = "MIT". That is permissive: you can use the output commercially, and there is no copyleft obligation on the audio you generate. The licence of the voice model is a separate question, and the README does not address it. Kokoro-82M is hosted on Hugging Face, and its own model card governs use of the weights, not abogen's MIT grant. If you plan to publish the audio, check that separately.

The project is not archived, and the last push was on 2026-09-07. The most recent tagged release is v1.3.1 from 2026-02-06, preceded by v1.2.5 on 2025-12-10 and v1.2.4 on 2025-11-28. The gap between the latest tag and the latest push is roughly seven months, which suggests development continues on main without a matching release. That matters for upgrade planning: pinning to a released version means you are not getting whatever landed since February.

Upgrade cost is dominated by dependencies, not by abogen's own code. The torch pin on the NVIDIA pip path, the numpy<2 constraint on macOS, and the nightly ROCm build on AMD Linux all mean an abogen upgrade can force a torch reinstall. Budget for a fresh virtual environment rather than an in-place upgrade. The Docker path is the cheapest to keep current, since docker compose up --build rebuilds the image with the TORCH_INDEX_URL build argument, which defaults to the cu126 wheel index.

## Conclusion

abogen fits anyone who already has a local Kokoro or PyTorch setup and wants chapter-aware audiobooks with subtitle files, especially on Linux with an NVIDIA or AMD GPU. It is the wrong pick if you want a hosted service, a Windows AMD GPU path, or Python 3.13 on Linux. Before committing, confirm your Python version is inside the >=3.10, <3.13 range, that espeak-ng is on PATH, and whether you need the [cuda], [cuda126], [cuda130] or [rocm] extra, since the README treats each as a separate install line rather than a fallback.

## FAQ

### How to use abogen?

Install espeak-ng first, then install abogen with uv or pip for your platform. Either launch the desktop app or run docker compose up --build and open http://localhost:8808, then upload an EPUB, PDF, text, markdown or subtitle file, choose a voice, and start the job.

### Which Python versions does abogen support?

pyproject.toml declares requires-python = ">=3.10, <3.13", so Python 3.13 is excluded. The macOS uv instructions in the README ask for --python 3.13, which conflicts with that constraint.

### Does abogen support AMD GPUs?

On Linux, yes, through the [rocm] extra with a ROCm 6.4 nightly wheel index. On Windows the README states ROCm is not available, so AMD acceleration is not supported there and the install falls back to CPU.

### What does abogen need installed besides itself?

espeak-ng is required on every platform: the .msi installer on Windows, brew install espeak-ng on macOS, and the distro package on Linux. The optional LLM text normalization feature also expects a server at ABOGEN_LLM_BASE_URL, which defaults to http://localhost:11434.

## Sources

- [denizsafak/abogen on GitHub](https://github.com/denizsafak/abogen)
- [License: MIT](https://github.com/denizsafak/abogen/blob/main/LICENSE)
- [Project website](https://pypi.org/project/abogen/)
- [README](https://github.com/denizsafak/abogen/blob/main/README.md)
- [Releases](https://github.com/denizsafak/abogen/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/denizsafak-abogen
