Alexandria Audiobook Generator: A Local Qwen3-TTS Pipeline for Multi-Voice Books
AI-powered multi-voice audiobook generator — LLM script annotation, voice cloning, voice design, LoRA training, per-line style control, and export to MP3, chaptered M4B, or Audacity multi-track. Built on Qwen3-TTS.
At a glance
- What is it?
- Alexandria turns a book into a cast-voiced audiobook through LLM script annotation, local Qwen3-TTS synthesis, and a browser editor. It is a self-hosted tool for people with a CUDA GPU and a willingness to babysit a pipeline.
- Who is it for?
- Adopt Alexandria if you have an NVIDIA GPU with 8 GB or more VRAM, an OpenAI-compatible LLM endpoint you control, and a book you are willing to annotate line by line. Skip it if you are on AMD Windows, Apple Silicon, or Intel macOS, where the README states the app runs in CPU mode only, or if you need a hosted service with no local setup.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 58 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What Alexandria Solves That a Plain TTS Script Does Not
Feeding a novel to a text-to-speech engine produces one narrator voice reading dialogue tags out loud. The work Alexandria automates is the step before synthesis: deciding who speaks each line and how it should be delivered. The README describes an LLM pass that parses text into JSON with speakers, dialogue, and TTS instruct directions, then an optional second pass that strips attribution tags from dialogue, splits misattributed narration and dialogue, merges over-split narrator entries, and validates instruct fields. That second pass is the interesting part. Anyone who has annotated a book by hand knows the failure mode: the model writes "'Get out,' she said" as a line of dialogue and the narrator reads the tag in character.
The audience is narrow and specific. You need a book in a supported language (English, Chinese, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish, or auto-detect), a GPU, and patience. The README's own note to new users says the project has seen a surge of attention and that the maintainer may not respond to every issue promptly, pointing readers to the Wiki first. That is an honest framing of what this is: a working pipeline maintained by a small team, not a product with a support desk.
The Pipeline: LLM Annotation, Chunking, Voice Assignment, Synthesis
The data flow has four visible stages. First, an LLM connected over any OpenAI-compatible API (LM Studio, Ollama, OpenAI, or another) reads the book and emits structured JSON. Second, the pipeline chunks that JSON: consecutive lines by the same speaker are grouped up to 500 characters so the delivery flows naturally, and each chunk carries the character roster plus the last three script entries so names and style stay consistent across chunk boundaries. Third, voices get attached. Persona Generation has the LLM write a voice description per character, generates reference audio through VoiceDesign, and assigns clone voices automatically. Speaker aliases let you collapse variants, so "YOUNG ELENA" can point at the same configuration as "ELENA". Fourth, Qwen3-TTS synthesizes locally.
The built-in engine loads models directly, and the README states no external TTS server is required, though an external Qwen3-TTS Gradio server mode exists. Batch processing generates many chunks at once, and an optional torch.compile path is described as codec compilation. The README claims 3-6x real-time throughput for batch processing and 3-4x faster batch decoding with codec compilation; those are the project's own numbers, not measurements I can confirm. The editor sits on top of all of it: five core steps (Setup, Script, Voices, Editor, Result) plus Designer, Dataset, and Training tools, with selective regeneration so you re-render one chunk instead of the whole book.
Installing Alexandria with Pinokio or Docker on Port 4200
The README lists two installation paths. Option A is Pinokio, which the project marks as recommended and which the repository supports through pinokio.js, install.js, start.js, and update.js. Option B is Docker. The docker-compose.yml in the repository builds the image, publishes port 4200, sets ALEXANDRIA_CONFIG_PATH, and reserves one NVIDIA GPU.
services:
alexandria:
build: .
ports:
- "4200:4200"
environment:
- ALEXANDRIA_CONFIG_PATH=/alexandria/config/config.json
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]The compose file also mounts volumes for uploads, designed voices, clone voices, LoRA models and datasets, scripts, config, and output, plus a named hf_cache volume. That hf_cache mount matters: the Dockerfile installs qwen-tts==0.1.1 on top of a pytorch/pytorch:2.8.0-cuda12.8-cudnn9-runtime base, and the README says model weights download automatically on first use at roughly 3.5 GB per variant. Without the cache volume you re-download on every container rebuild.
The Dockerfile ends with a CMD that runs the application entry point, and sets ALEXANDRIA_HOST to 0.0.0.0 inside the container while exposing 4200.
ENV ALEXANDRIA_HOST=0.0.0.0
EXPOSE 4200
CMD ["python", "app/app.py"]Once the build finishes and the first model download completes, the editor should be reachable at http://localhost:4200. Expect that first start to be slow; the README puts total disk at around 20 GB, split between an 8 GB venv with PyTorch and roughly 7 GB of weights.
Where Alexandria Falls Down: Platform Gaps and Manual Review
The GPU compatibility table is the first thing to read before installing. NVIDIA on Windows and Linux is full support with driver 550+ and CUDA 12.8. AMD on Linux is full support with ROCm 6.3+. AMD on Windows is CPU only, and the README is explicit that GPU acceleration is not supported there and that Linux is the path to AMD acceleration. Apple Silicon and Intel on macOS are CPU only, with MPS acceleration stated as not currently supported and the app described as functional but slow. On a novel-length book, "functional but slow" is the whole story.
The second limitation is the LLM dependency. Alexandria does not ship a model for annotation; it expects an OpenAI-compatible endpoint. If you run LM Studio or Ollama locally, you are now running two model families on one GPU, and the README's 8 GB VRAM minimum assumes the TTS models get their share. Each TTS model uses about 3.4 GB, and the remaining VRAM determines batch size, so a large local LLM competing for memory directly reduces how many chunks you can synthesize at once.
The third is that annotation is probabilistic. The optional review pass fixes common errors, which is an admission that the first pass produces them. On a book with unusual dialogue formatting, dense attributions, or many minor characters, expect to spend real time in the Chunk Editor correcting speaker assignments. Selective regeneration makes that cheaper, not free. If your source text is a technical manual or anything without a speaking cast, the persona generation and multi-voice machinery are solving a problem you do not have.
Alexandria Against a Cloud TTS Service
The obvious alternative is a hosted TTS API. The difference is architectural, not just financial. A cloud service gives you a voice per request and no annotation layer; you write the code that splits a book into speaker turns yourself, and you accept that the text leaves your machine. Alexandria inverts both: the annotation is the product, and synthesis runs on your hardware through Qwen3-TTS. The README's emphasis that no external TTS server is required, and that weights are downloaded locally, is the project's central claim about where your manuscript goes.
The trade is control for operational burden. A hosted API scales without a GPU and gives consistent latency; Alexandria asks for 8 GB of VRAM minimum, 16 GB recommended, 16 GB of RAM, and about 20 GB of disk. It also asks you to own the LLM side. The upside is that voice cloning from a 5-15 second sample, voice design from a text description, and LoRA fine-tuning on your own dataset all happen locally, and the repository ships builtin_lora/ with pre-trained adapters the README describes as ready to assign to characters. That combination, a trainable local voice identity plus an annotation pipeline, is not something a per-request API offers.
Licence, Maintenance, and the Cost of Upgrading
The repository is MIT licensed, which is permissive and places few obligations on how you use or redistribute the code. That licence covers Alexandria's own source. It does not automatically cover the Qwen3-TTS model weights, the built-in LoRA adapters in builtin_lora/, or any voice you clone from a reference sample. Cloning a real person's voice raises consent and publicity questions that no software licence settles, and the README does not address them. If you plan to publish what you generate, check the terms attached to the model weights and the LoRA assets separately from the MIT grant. This is not legal advice.
On maintenance: the repository is not archived, and the last push was on 2026-08-02. The most recent release listed is v1.5 from 2026-03-03, which the release notes title "M4B Export, Voice Cloning Upload & Stability Fixes", following v1.4 ("Dataset Builder, Built-in Presets & Codebase Polish") and v1.3 ("LoRA Voice Training & Voice Designer Integration"). The gap between the last release and the last push suggests ongoing commits between tagged versions, though the release notes do not say what those commits contain. Upgrade cost is dominated by the model cache: the Dockerfile pins qwen-tts==0.1.1 and the base image pins PyTorch 2.8.0 with CUDA 12.8, so moving to a newer Qwen3-TTS build means rebuilding against a possibly different CUDA base, and a driver below 550 on NVIDIA would need updating first. The README does not document a rollback path if a new model variant changes output characteristics.
Editorial conclusion
Adopt Alexandria if you have an NVIDIA GPU with 8 GB or more VRAM, an OpenAI-compatible LLM endpoint you control, and a book you are willing to annotate line by line. Skip it if you are on AMD Windows, Apple Silicon, or Intel macOS, where the README states the app runs in CPU mode only, or if you need a hosted service with no local setup. Before committing, verify three things: that your driver is 550+ on NVIDIA or ROCm 6.3+ on AMD Linux, that the ~3.5 GB per-model weight download fits your disk alongside the roughly 20 GB install, and that the M4B chapter markers land where your player expects them, since the README documents per-chunk and auto-detected markers but not how auto-detection resolves a chapter boundary.
Frequently asked questions
How can you tell if an audiobook is AI generated?
The README does not describe detection methods. It documents how Alexandria generates audiobooks: an LLM annotates the script and a built-in Qwen3-TTS engine synthesizes the audio locally, with voice cloning, voice design, and LoRA training available.
Does Alexandria use AI?
Yes. The README describes two AI stages: an LLM that annotates the book into JSON with speakers, dialogue, and TTS instruct directions, and a built-in Qwen3-TTS engine that synthesizes the audio locally. Voice cloning, voice design, and LoRA training are also part of the pipeline.
Which AI audiobook narrator is the best?
The README does not rank narrators or compare quality across tools. It describes the voices Alexandria offers: 9 pre-trained voices with instruct-based emotion and tone control, voice cloning from a 5-15 second reference sample, voice design from a text description, and LoRA adapters shipped in builtin_lora/.
What is the most popular audiobook of all time?
The README does not address audiobook popularity or sales figures; it documents Alexandria's own pipeline for generating audiobooks from a book or novel.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/finrandojin-alexandria-audiobook)
Community notes