ebook2audiobook: A Self-Hosted E-Book to Audiobook Converter with Voice Cloning
Generate audiobooks from e-books, voice cloning & 1158+ languages!
At a glance
- What is it?
- ebook2audiobook turns EPUB, PDF and other e-book files into chaptered audiobooks using XTTSv2 and other TTS engines, with optional voice cloning and a Gradio interface. It is a local, Docker-friendly tool for people who want narration without a cloud service.
- Who is it for?
- Adopt ebook2audiobook if you already own non-DRM e-books, you have a machine with at least 2 GB RAM and 1 GB VRAM, and you accept that CPU-only synthesis is slow. Skip it if your library is DRM-protected, if you need a hosted service with an SLA, or if you expect the project to handle licensing questions for you.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What ebook2audiobook actually does with your e-book file
The project takes a non-DRM e-book file and produces an audio file with chapters and metadata. That is the whole promise, and the README is direct about the legal boundary: it is "intended for use with non-DRM, legally acquired eBooks only." The supported input list is unusually wide for this category: .epub, .mobi, .azw3, .fb2, .lrf, .pdf, .txt, .rtf, .doc, .docx, .html, .odt, and image formats such as .tiff, .png and .jpg, plus .zip archives. A separate TextArea in the interface converts a short block of pasted text straight to audio, which is useful for testing a voice before committing to a full book.
The audience is narrow but real. It is for readers who already own e-books and want a personal listening copy, for people who want a specific cloned voice reading a text, and for anyone who would rather run synthesis on their own hardware than upload a manuscript to a hosted service. It is not a library management tool, and it does not acquire or decrypt books for you.
The TTS engine stack and how a book becomes audio
The pipeline is split into three visible parts. First, extraction: the e-book is parsed and its text pulled out, with OCR scanning available for pages that are images rather than text. Second, text preparation: the requirements file lists a large set of language-specific libraries (jieba, pypinyin, pycantonese, fugashi, pythainlp, indic-nlp-library, stanza, num2words2) which handle tokenisation, romanisation and number expansion before synthesis. Third, synthesis: one of the supported engines renders the audio, and the result is written out with chapters and metadata. The supported engine list is XTTSv2, Bark, Fairseq, VITS, Tacotron2, Tortoise, GlowTTS and YourTTS, and the README notes that custom models can be supplied for XTTSv2, VITS, FAIRSEQ and PIPER.
The design choice worth flagging is that XTTSv2 is the default and the heaviest of the group. The README states plainly that "Modern TTS engines are very slow on CPU, so use lower quality TTS like YourTTS, Tacotron2 etc." That is an honest admission that the quality and speed settings trade against each other, and it means the engine name is the single most consequential option in a conversion.
Installing ebook2audiobook and running a first conversion
The README points at a Quick Start section and at the releases page for downloads, and the repository root contains per-platform launchers: ebook2audiobook.sh, ebook2audiobook.command, ebook2audiobook.cmd, e2a.sh and e2a.cmd. Python support is declared in pyproject.toml as requires-python = ">3.9,<3.13", so a 3.13 interpreter will not satisfy the constraint.
For a containerised run, the repository ships docker-compose.yml and the README documents accelerator profiles. The compose file itself carries the usage lines in its header comment:
docker compose --profile cpu up
DEVICE_TAG=cu130 docker compose --profile cuda up
DEVICE_TAG=rocm docker compose --profile rocm upThe compose file is explicit that it is a local-build path only. Its header states that the image "is never pulled from a registry" and that pull_policy is set to never, so an `up` uses a locally built image or fails. It also warns that the build context is the whole repository because the Dockerfile runs COPY . /app, which means a lone downloaded Dockerfile will not work. The compose file asks for Compose V2, the `docker compose` plugin, not the legacy `docker-compose` v1, because only V2 honours deploy.resources.devices outside Swarm and the pull_policy key.
Once the app is running, the local entry point is the Gradio web interface, launched through the platform launcher scripts. In that interface you pick an e-book file, choose an engine and a voice, and start the job. The README also documents a headless path and a custom-model upload path for XTTSv2, with a help command documented in the README's help output section.
Where ebook2audiobook breaks down
The first limitation is stated by the project itself: DRM-protected books are out of scope. If your library comes from a store that wraps files in DRM, this tool is the wrong one and no configuration will change that.
The second is hardware. The README gives minimums of 2 GB RAM and 1 GB VRAM, with 8 GB RAM and 4 GB VRAM recommended, and it repeats that modern TTS engines are very slow on CPU. A long novel synthesised on CPU with XTTSv2 is a realistic candidate for a very long wait, and the README's own advice is to drop to a lower-quality engine in that case. If your machine has no usable GPU and you want near-real-time output, this is not the tool for the job.
The third is that the Docker path is not a one-line pull. Because the compose file refuses to pull from a registry and builds from the full checkout, you need the entire repository, a working build toolchain, and the matching DEVICE_TAG for your accelerator. Anyone expecting `docker run` against a prebuilt image will be surprised. The compose header also cautions against `pull_policy: build` because COPY . /app invalidates the cache for the expensive build layer on any source change, which means rebuilds after an edit are not cheap.
How it compares with a plain Piper or Coqui TTS setup
The honest alternative is to skip the wrapper and drive a TTS engine directly. Piper and Coqui TTS both appear in this project's dependency list (piper-tts and coqui-tts[languages]), so the underlying synthesis is not unique to ebook2audiobook. If you have a clean plain-text manuscript and one voice in mind, a short script against coqui-tts or piper-tts gives you the same audio with far fewer moving parts and no Gradio server.
What you lose by going direct is the part ebook2audiobook actually contributes: format parsing across EPUB, MOBI, AZW3, FB2, PDF and the rest; OCR for image-based pages; chapter and metadata handling in the output; the multi-language text preparation libraries; and a UI that ties engine, voice and output format together. The trade is real in both directions. Direct engine use is simpler to automate and easier to debug; ebook2audiobook is the better fit when the input is a real e-book file rather than a text file you already cleaned up.
Maintenance, versions and the licence question
The repository is not archived, and the last push was on 2026-09-17, so the codebase is being touched recently. Releases are frequent and date-versioned: v26.9.7 on 2026-09-09, v26.8.20 on 2026-08-22, and v26.7.27 on 2026-07-31. That cadence has a cost for anyone pinning a deployment, because the README includes a section on reverting to older versions, which implies upgrades are not always frictionless. Budget for reading release notes between versions rather than upgrading blind.
On licensing there is a discrepancy you should resolve yourself rather than assume. The repository LICENSE is Apache-2.0, but the pyproject.toml classifiers declare "License :: OSI Approved :: MIT License". Those are different licences with different patent and notice requirements. This is not legal advice, and the ambiguity is upstream, but if you plan to redistribute the tool or embed it in a product, confirm the intended licence with the maintainers before relying on either declaration. Note also that the bundled ext/py/demucs dependency and the TTS model weights carry their own terms, which may differ from the application code.
Editorial conclusion
Adopt ebook2audiobook if you already own non-DRM e-books, you have a machine with at least 2 GB RAM and 1 GB VRAM, and you accept that CPU-only synthesis is slow. Skip it if your library is DRM-protected, if you need a hosted service with an SLA, or if you expect the project to handle licensing questions for you. Before committing, verify three things: that your GPU stack matches one of the Docker Compose profiles (cpu, cuda, rocm, xpu, jetson), that your e-book format appears in the supported list, and that the Apache-2.0 LICENSE file in the repository root matches what your distribution requires. Note also that the pyproject.toml classifier says MIT while the repository LICENSE is Apache-2.0, so confirm the terms with the upstream project before shipping it inside a product.
Frequently asked questions
How can I convert an ebook to an audiobook for free with ebook2audiobook?
Run it locally from a checkout of the repository, either through the platform launcher scripts such as ebook2audiobook.sh or via the Docker Compose profiles documented in docker-compose.yml. There is no paid tier described in the README, and the project also links to a Hugging Face Space and a free Google Colab notebook for remote runs. The README restricts use to non-DRM, legally acquired e-books.
Can ebook2audiobook convert a PDF to an audiobook for free?
PDF is in the supported input list, and the README also lists OCR scanning for files whose pages are images, which is what makes scanned PDFs usable. The repository requirements include pymupdf and pytesseract for that path. Output can be written as mp3, m4b, flac, wav and other formats.
Is ebook2audiobook free to use?
The project is published under Apache-2.0 according to the repository LICENSE, though pyproject.toml classifiers declare MIT, so the two declarations disagree. The README describes no paid tier and links to a Ko-fi page for supporting the developers. Treat the licence discrepancy as something to confirm upstream before redistribution.
Does ebook2audiobook support voice cloning?
Yes. The README lists optional voice cloning using your own voice file as a feature, and there is a Cloned Voices section in the table of contents. The default engine, XTTSv2, is the one the README describes for custom model uploads.
What hardware does ebook2audiobook need?
The README states minimums of 2 GB RAM and 1 GB VRAM, with 8 GB RAM and 4 GB VRAM recommended. It supports CPU, CUDA, ROCm, XPU, JETSON and Apple MPS, and warns that modern TTS engines are very slow on CPU, suggesting lower-quality engines like YourTTS or Tacotron2 for CPU-only machines.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/drewthomasson-ebook2audiobook)