Pandrator: a local, browser-based workspace for turning EPUBs and PDFs into audiobooks
Turn PDFs and EPUBs into audiobooks; subtitles or videos into dubbed videos (including translation), and more. For free. Pandrator uses local models, including voice-cloning (instant, RVC-enhanced, XTTS fine-tuning) and LLM processing. It aspires to be a user-friendly app with a GUI, an installer and all-in-one packages.
At a glance
- What is it?
- Pandrator bundles document cleanup, voice cloning and speech generation into one browser interface, with the option to let an existing MCP-capable coding agent do the language work. It is free, MIT-licensed and still pre-1.0.
- Who is it for?
- Adopt Pandrator if you want a local, MIT-licensed pipeline that keeps document cleanup, subtitle correction and speech generation in one reviewable interface, and if you are comfortable with a 0.9.x project whose Windows launcher is unsigned. Do not adopt it if you need a stable CLI contract or a signed installer today.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Pandrator is for, and who it is actually aimed at
Pandrator is a browser-based production workspace for audio derived from text and video. The README frames the scope as three families of work: narrated audio and M4B audiobooks from books and documents, reviewed or translated subtitles from audio and video, and dubbed video with synchronized speech. A fourth, smaller path called Quick Transcribe turns a clip or a microphone recording into TXT, SRT or JSON without creating a session.
The audience is narrower than "anyone who wants an audiobook". The project assumes you are willing to run speech and transcription models on your own machine, or connect a cloud provider, and that you want to inspect intermediate results rather than accept a single button-press output. The README repeatedly uses the word "reviewed": reviewed subtitles, reviewed boundaries, a reviewable editor. That is the design centre. If you want a one-shot converter, the interface will feel like it has too many stops.
The stated aspiration is a user-friendly app with a GUI, an installer and all-in-one packages. The current release ships a Manager application for Windows and Linux, which is a step toward that, but the underlying Python package still exposes a CLI entry point (pandrator = "pandrator.web.cli:main") and the repository keeps a pixi.lock alongside requirements.txt, which tells you development is not yet settled on one environment story.
How the pipeline is put together: documents in, validated artifacts out
The mechanism visible in the repository is a Flask application (flask>=3.1,<4) served by waitress, backed by SQLAlchemy and Alembic for state and migrations. That matters because it means Pandrator is not a script that streams text to a model. It is a server that records work, which is why the README can promise that "Pandrator tracks batches, validates submissions, and keeps completed work so an interrupted agent can resume."
Document handling is a chain rather than a single conversion. Import accepts TXT, PDF, EPUB, DOCX, MOBI or pasted text, and the dependencies show what each step costs: PyMuPDF for PDF extraction, EbookLib for EPUB, BeautifulSoup4 for markup, and optional PaddleOCR plus onnxruntime for scanned pages. MOBI is the outlier: the README states that MOBI conversion needs optional Calibre, so that format depends on software outside the package.
Text preparation then runs chapter detection, cleanup and optional AI assistance before any speech is generated. Sentence boundaries come from sentence-splitter and wtpsplit-lite, with num2words and Unidecode handling numbers and transliteration. Speech generation happens in segments, which is what allows the compare-takes and regenerate-selected-passages behaviour the README describes. Output can be WAV, MP3, Opus, FLAC or M4B, with chapters, metadata and cover art where the format supports them; mutagen is the dependency behind that tagging.
Subtitles follow a parallel path. Whisper-family engines produce timestamps with optional speaker diarization, SRT is the working format, and WebVTT, ASS and SSA uploads are recognized but converted. Cues can be corrected, translated, split, merged and retimed. Translation can come from configured language providers (deepl and litellm are both dependencies) or from the passive MCP route described below.
The passive MCP workflow, and why it changes the cost model
The most distinctive design decision in the README is what it calls the in-harness passive MCP workflow. If you already use Codex, Claude Code, OpenCode or another MCP-capable host, that host's model does Pandrator's language work. You do not configure a separate LLM provider or API key inside Pandrator for this route. The README is explicit about the division: Pandrator prepares and tracks the work, the model in your existing conversation processes it and submits results through MCP, and "passive" describes Pandrator's role because it does not make the LLM calls itself.
The flow is a loop. Pandrator prepares a batch, your agent corrects, translates or optimizes it, Pandrator validates and saves, and either more batches follow or the run completes into review, export or speech generation. The README names four uses: cleaning PDF and EPUB text before audiobook preparation, correcting and translating subtitles while keeping cue identities and shared terminology across batches, optimizing text for speech while retaining the source wording as a separate artifact, and coordinating an end-to-end workflow from import to verified deliverables.
The practical consequence is that the language quality is whatever your agent's model produces, and the README states the cost plainly: your host's normal usage limits and costs apply, and if its model runs in the cloud, the text it processes is sent there even when Pandrator runs locally. That undercuts a simple local-only privacy story. Speech recognition and speech generation still use the engines you select in Pandrator, so the audio itself does not leave your machine through this route, but the text does.
One operational caveat the README gives directly: the agent must continue the tool loop, because starting a passive run alone does not process it. A run that looks stalled may simply be an agent that stopped calling tools.
Installing Pandrator and making a first audiobook
The README points to a Manager application rather than a pip install for desktop users. Download the Windows .exe or the Linux .AppImage from the releases page, matching the versions it lists: Pandrator 0.9.2 with Manager 0.9.22, for Windows 10/11 and Linux desktop on x86-64. On Linux, the README says to make the AppImage executable first.
chmod +x PandratorManager-0.9.22-x86_64.AppImage
./PandratorManager-0.9.22-x86_64.AppImageAfter launching, the Manager asks you to choose a parent folder. It creates a Pandrator workspace inside that folder and opens the setup interface in your browser. The README states that Docker, WSL and a separate Python installation are not required for this route.
Inside the setup interface you install Pandrator and then at least one engine. The README suggests Kokoro for ready-made narration voices or CrispASR for transcription, with voice cloning and other engines added later. Local models download separately on first setup, and the README notes that speed and memory requirements vary by engine and hardware, with the Manager showing available compute options.
The first real task is deliberately small. Choose Generate an audiobook and paste a short passage, or open Quick Transcribe to upload a clip or record your microphone, then listen, review and save. If you prefer the Python package, pyproject.toml requires Python >=3.11,<3.13 and exposes a console script, so the entry point is:
pandratorThat is the extent of what the repository documents for a package-level start; the desktop Manager path is the one the README walks through.
Where Pandrator gets in the way
The Windows launcher is unsigned, and the README says so: Windows may show an unknown publisher warning. The installation guide explains the warning and the release checksums, which is the right place for it, but it means every Windows user meets a security prompt on first run. For managed corporate machines with application control policies, that alone can disqualify the tool.
Engine setup is the second friction point. Local models download separately on first setup, and the project does not promise a fixed disk or memory budget because requirements vary by engine and hardware. There is no single documented minimum specification in the README, so capacity planning is guesswork until you have installed the engine you intend to use.
Format support has holes that are easy to miss. MOBI needs Calibre installed separately. OCR is an optional extra (paddleocr and onnxruntime live under the ocr extra), so a scanned PDF without that extra will not yield usable text. Text normalization through nemo-text-processing is marked sys_platform != 'win32', so that particular dependency is not installed on Windows at all.
Finally, the version number. Pandrator is at 0.9.2, and three releases landed within two days in September 2026. The last push to the repository was on 2026-09-10. That is a project moving quickly, which is good for features and bad for anyone who needs a frozen interface. The README does not document rollback between versions, and there is no compatibility statement for the on-disk workspace the Manager creates.
Pandrator compared with Audiblez and plain TTS scripts
The obvious alternative people search for alongside Pandrator is Audiblez, another free tool in the ebook-to-audiobook space. The difference in approach is scope and state. A focused converter takes an EPUB, runs it through a TTS engine and hands back an audio file; the README's description of Pandrator is a workspace that imports six input types, detects chapters, cleans extraction artifacts, generates in segments so you can compare takes and regenerate passages, and exports to five audio formats with chapters and cover art. That is more surface area, and more places for a run to stop and wait for you.
Against a hand-rolled script over a TTS library, the difference is the review loop and the persistence layer. Pandrator keeps a database (SQLAlchemy, Alembic) and tracks batches, which is what makes the passive MCP workflow resumable after an interruption. A script gives you reproducibility and no UI; Pandrator gives you a browser editor and a state store. If your source material is clean and your output is one MP3, the script is less machinery.
The subtitle and dubbing side has fewer direct free equivalents, because it combines transcription, cue-level editing, translation and synchronized speech in the same workspace. That combination, rather than any single capability, is the reason to pick Pandrator over a chain of separate tools.
Licence, maintenance and the cost of keeping up
Pandrator is MIT-licensed, and pyproject.toml declares license = "MIT" with license-files = ["LICENSE"]. MIT is permissive, so the practical implication is that you can use, modify and redistribute it, including in commercial settings, provided you keep the licence notice. That is a general statement about the licence text, not legal advice, and the dependencies carry their own terms: litellm, PyMuPDF, PaddleOCR and the speech engines each ship under their own licences, and model weights often have terms separate from the code that loads them. Check those before shipping anything commercially.
The maintenance picture from the repository alone: not archived, last push on 2026-09-10, releases 0.9.0 through 0.9.2 within three days of each other. That is a fast cadence, and upgrade cost follows from it. The Manager and the Pandrator package are versioned separately (Manager 0.9.22 against Pandrator 0.9.2 in the download labels), so an upgrade is not a single artifact swap. requirements.txt is generated from pyproject.toml by scripts/generate_requirements.py and carries a warning not to edit it directly, which tells you dependency bumps are expected to be routine and that pinning happens upstream. numpy is pinned exactly at 1.26.4 and wtpsplit-lite at 0.2.0, so those two will conflict with anything else in your environment that wants a different version.
Editorial conclusion
Adopt Pandrator if you want a local, MIT-licensed pipeline that keeps document cleanup, subtitle correction and speech generation in one reviewable interface, and if you are comfortable with a 0.9.x project whose Windows launcher is unsigned. Do not adopt it if you need a stable CLI contract or a signed installer today. Before committing, verify that your chosen engine fits your hardware, check the release checksums the installation guide mentions, and read docs/getting-started/installation.md for the unknown-publisher warning.
Frequently asked questions
Does Pandrator need Docker, WSL or a separate Python installation?
No, for the desktop route. The README states that Docker, WSL and a separate Python installation are not required when you use Pandrator Manager, which creates a Pandrator workspace inside a parent folder you choose and opens the setup interface in your browser.
Which engines should I install first in Pandrator?
The README suggests Kokoro for ready-made narration voices or CrispASR for transcription, with voice cloning and other engines added later. Local models download separately on first setup, and speed and memory requirements vary by engine and hardware.
Can Pandrator translate subtitles without me configuring an API key?
Yes, if you use an MCP-capable host such as Codex, Claude Code or OpenCode. The README says the host's model does the language work through the passive MCP workflow, so you do not configure a separate LLM provider or API key in Pandrator for that route, though your host's usage limits and costs apply.
What audio formats can Pandrator export?
The README lists WAV, MP3, Opus, FLAC and M4B, with audiobook chapters, metadata and cover art where the format supports them.
Why does Windows warn about an unknown publisher when I run Pandrator?
The README states that the Windows launcher is currently unsigned, so Windows may show an unknown publisher warning. The installation guide explains this warning and the release checksums.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lukaszliniewicz-pandrator)