Model or dataset
estebanstifli/LocalText2Voice avatar
estebanstifli/LocalText2Voice

Narration first, video second, and a Storyboard the author still calls beta

A complete local production workflow for clean narration, structured learning content, and podcast-ready audio

370 stars29 forksPythonMIT

At a glance

What is it?
LocalText2Voice is a free desktop app for producing long-form audio and video on your own machine, with Piper, Kokoro, Chatterbox, Qwen3 TTS and OmniVoice as local engines and cloud TTS as an option. The audio half is a correction workflow rather than a one-shot render, and the video half is explicitly beta, which is the right way to read both halves.
Who is it for?
Adopt LocalText2Voice if you produce audiobooks, lessons or narrated video on Windows and want the narration step to happen on your own machine with correctable segments rather than one long render. Do not adopt it expecting a finished video pipeline, because Video Storyboard is beta and the project asks for human review of continuity, timing and generated media, and macOS is not supported at all.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Windows gets the installer, Linux gets a source checkout

The distribution story is uneven, and knowing which side you are on saves an afternoon. Windows 10 or 11 at 64-bit is the recommended and fully packaged path, with a LocalText2Voice-Setup.exe on the releases page.

Linux at 64-bit runs from source. It needs Python 3.10 or newer with virtual environment support, and FFmpeg available in PATH. The documented Linux workflow has been tested on Arch and CachyOS with KDE Plasma and Wayland, which is a specific combination rather than a general claim about the distribution.

macOS is not currently tested or officially supported. That is stated in one line and is the single fact to check before anything else, because the rest of the requirements assume a supported platform.

The repository shows what a packaged build involves: build_windows.bat for the Windows build, run_dev.bat and run_dev.sh for development, an installer/ directory, and PyInstaller in the requirements. A bundled ffmpeg/ directory is also present, which is how FFmpeg can be available without the user installing it themselves.

Three hardware profiles, described as guidance rather than limits

The requirements table is framed carefully, as practical recommendations rather than strict compatibility limits, which is the honest way to describe a tool whose needs depend on which engine you pick.

The Basic profile is a 64-bit four-core CPU with 8 GB of RAM and 5 GB free, no dedicated GPU, aimed at editing projects, cloud APIs and lightweight local generation. Recommended is a modern six-core CPU, 16 GB of RAM, an SSD with 20 to 30 GB free and an NVIDIA CUDA GPU with 8 GB of VRAM, for regular use of local neural TTS at a better speed. High performance is an eight-core or better CPU, 32 GB of RAM, at least 50 GB free on an SSD and an NVIDIA RTX GPU with 12 GB or more of VRAM, for larger models, long-form production and optional speech verification.

The NVIDIA GPU is optional. Most local engines can fall back to CPU execution, but generation may be considerably slower, and acceleration currently focuses on NVIDIA CUDA, so other GPU families should not be assumed to speed up every engine.

Disk is the number people get wrong. It grows with each optional engine, because models and isolated dependencies, including PyTorch runtimes, are downloaded on demand, and generated projects and exported audio need space on top of those figures.

Local engines and cloud APIs are two separate decisions

You can run the whole narration stage offline, and the engines for that are named: Piper, Kokoro, Chatterbox, Qwen3 TTS and OmniVoice, plus an optional F5-TTS Russian engine that is explicitly non-commercial. Local engines keep your data on your own computer.

The alternative is a set of cloud APIs, also named: OpenAI TTS, ElevenLabs, Google Gemini TTS and Azure Speech, plus direct video providers for the picture stage. The project describes these as optional and clearly configured, which is the phrasing that matters: nothing is sent anywhere unless you point the app at a provider.

The dependency layout shows how the engines are isolated. There is a requirements-kokoro-engine.txt alongside the main requirements.txt, which is the pattern of per-engine environments rather than one shared environment where five frameworks fight over the same PyTorch install.

The main requirements also hint at what the app does around the audio: soundcard for capture, mutagen for tagging, python-docx for Word scripts, num2words for spelling numbers out, QtAwesome for icons, and litellm for talking to several providers through one interface.

Audio is a correction workflow, not a one-pass render

The audio pipeline is built around the fact that a three-hour audiobook will have a mistake in it. The documented order is: import or write a script, generate narration, review and correct segments, build an audio mix, and only then move to pictures.

Segments are the unit. Version 2.1.2 stores every generated segment in SQLite while batching project snapshots and pausing updates, which is what makes exporting a large audiobook cheaper than the version before it. That is a change to how the export reads and writes the project, not to the audio itself, and it is the kind of fix that only shows up on real book-length projects.

Two smaller signals point the same way. Every project stays editable, and human review is described as part of the workflow before final export rather than an option after it. The markup manual at docs/LTV_MARKUP.md is the contract for how a script is written for this, including the tags that drive segmentation and pronunciation.

For comparison, a text-to-speech API on its own gives you one long audio file and no way to fix paragraph twelve without regenerating everything.

Video Storyboard is beta, and the README says so in three lines

The video stage arrived in 2.0.1 as Video Storyboard and was expanded in 2.1.2. It turns narration into an editable visual timeline, analyses scenes and reusable entities, generates or imports images, creates video clips, and renders an MP4 with the narration audio underneath.

Version 2.1.0 added richer character, location, era and object context, reference-aware prompts, batch video generation, recovery for long-running jobs, and direct video API adapters alongside ComfyUI and Runpod. Version 2.1.1 added source-aware analysis safeguards, manual timed-prompt import, resumable work and clearer controls for long projects.

The limitation is stated rather than buried: this remains a beta, and AI continuity, timing and generated media still need human review. Continuity is the interesting one, because a visual timeline assembled by a model will drift on a face or a costume across scenes, and the reusable entity analysis is the mechanism trying to prevent it rather than a guarantee.

There are 50 built-in image styles, each with a preview, and you can write your own style prompt instead. That is a lot of surface for a feature the project is asking you to check by hand.

MCP exposes the storyboard as 76 tools with a closed UI

The 2.1.2 release puts the whole Video Storyboard workflow behind a local MCP server: 57 new storyboard tools, 76 in total, covering timed narration and scene editing through reference images, clips and the final MP4. The desktop UI can stay closed while it works.

That is a meaningful shift for a tool whose interface is a desktop app. The requirements pin mcp[cli] at 1.28 or newer, and the repository root holds mcp_stdio_bridge.py alongside engine_host.py, so the MCP server is a real component rather than an export format.

The same release works on the failure cases, which is where a long render earns its keep: persistent results let you review generated candidates and resume visual jobs, there is cancellation and retry, and there are safeguards against concurrent project edits. That last one matters as soon as two agents or two windows touch the same project.

Dependencies tell the rest of the story. fastapi and uvicorn are in the main requirements, so the MCP surface is served locally over HTTP, and boto3 is there with a comment about private reference uploads for Runpod public endpoints.

The markup manual is the real interface

There is no command line to learn here, so the documentation is the interface. docs/LTV_MARKUP.md is the markup manual, and it is what tells a script where a segment starts, how a number should be spoken and where a correction belongs.

Around it the repository is organised like a shipped desktop product rather than a script collection: locales/ for translations of the interface, voices/ for voice assets, music/ for background beds, assets/ and capturas/ for images and screenshots, config.example.json as a starting configuration, and an output/ directory for results. tests/ and pytest.ini are present, along with CONTRIBUTING.md and CHANGELOG.md.

One file deserves attention before you redistribute anything: THIRD_PARTY_NOTICES.md, next to a licenses/ directory. An app that bundles FFmpeg, offers six local engines of different provenance and calls several cloud providers is exactly the case where the third-party terms matter more than the project's own MIT licence.

The documentation also branches by platform and language: docs/LINUX.md for the source route, docs/VIDEO_STORYBOARD.md for the beta stage with an Spanish version alongside it.

Version 2.1.2, three releases in a week, and a truncated changelog

Cadence is quick: v2.1.0 on 2026-09-21, v2.1.1 on 2026-09-26 and v2.1.2 on 2026-09-28, with the last push on 2026-09-28. Three releases in eight days during a beta feature is a fast-moving project, so the version in your install is worth knowing before you compare output with a tutorial.

The licence is MIT, which covers this project's own code. The engines are another matter: the F5-TTS Russian engine is described as non-commercial, models are downloaded on demand rather than vendored, and THIRD_PARTY_NOTICES.md plus licenses/ exist to carry those terms. If you are producing paid work, read that file rather than assuming the app's licence covers the voices.

The home page for the project lives at andromedanova.com with a page for this tool at andromedanova.com/localtext2voice.html, and there is an official YouTube channel with the four long walkthroughs linked from the README, including a from-text-to-documentary workflow and an 18-minute narrated video built from biblical source text.

One caveat: the copy of the README available here stops partway through the 2.1.2 list, at the item about seeing the actual generation and export phase and its progress, so read the rest of that section in the repository.

Editorial conclusion

Adopt LocalText2Voice if you produce audiobooks, lessons or narrated video on Windows and want the narration step to happen on your own machine with correctable segments rather than one long render. Do not adopt it expecting a finished video pipeline, because Video Storyboard is beta and the project asks for human review of continuity, timing and generated media, and macOS is not supported at all. Verify first which engine you want, since each one downloads its own models and an isolated set of dependencies onto your disk.

Frequently asked questions

Does LocalText2Voice need a GPU?

No, an NVIDIA GPU is optional and most local engines can fall back to CPU execution, though generation may be considerably slower. Acceleration currently focuses on NVIDIA CUDA, so other GPU families should not be assumed to speed up every engine.

Which platforms does LocalText2Voice support?

Windows 10 or 11 at 64-bit through the packaged installer, and Linux at 64-bit from source with Python 3.10 or newer, virtual environment support and FFmpeg in PATH, tested on Arch and CachyOS with KDE Plasma and Wayland. macOS is not currently tested or officially supported.

Which TTS engines run without a network connection?

Piper, Kokoro, Chatterbox, Qwen3 TTS and OmniVoice, plus an optional non-commercial F5-TTS Russian engine, all run locally and keep data on your computer. Cloud options such as OpenAI TTS, ElevenLabs, Google Gemini TTS and Azure Speech are supported as optional, clearly configured alternatives.

Is the LocalText2Voice Video Storyboard ready for production?

No, it is labelled beta, having arrived in 2.0.1 and been expanded in 2.1.2. It turns narration into an editable visual timeline, generates or imports images, builds clips and renders an MP4, but AI continuity, timing and the generated media still need human review.

How does LocalText2Voice handle a very long audiobook?

Every generated segment is saved in SQLite, and version 2.1.2 batches project snapshots and pauses updates to export large audiobooks more efficiently. The workflow is built for correction: you review and correct segments before building the audio mix, and every project stays editable.

What can the LocalText2Voice MCP server do?

Version 2.1.2 exposes the Video Storyboard workflow through a local MCP server with 57 new storyboard tools, 76 in total, from timed narration and scene editing through reference images, clips and the final MP4, so the desktop interface can stay closed. Jobs keep persistent results with cancellation, retries and safeguards against concurrent project edits.

Official sources

  1. estebanstifli/LocalText2Voice on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/estebanstifli-localtext2voice.svg)](https://hysenlabs.com/projects/estebanstifli-localtext2voice)