# Manim voiceover lists four voice services and ships ten extras

> A plugin that puts narration inside Python animation code, records it from your microphone during rendering, or generates it, and then triggers animations on individual spoken words. The most interesting dependency is the one that transcribes your own voice so that word timing works for recordings too.

**ManimCommunity/manim-voiceover** — Manim plugin for all things voiceover

- Repository: https://github.com/ManimCommunity/manim-voiceover
- Website: https://voiceover.manim.community/en/stable
- Stars: 316 · Forks: 77
- Language: Python
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/manimcommunity-manim-voiceover

## The readme names four voice services and the package ships ten extras

The readme has a short list headed with the words currently supported, and it contains four entries: a cloud speech service from Google, a cloud speech service from Microsoft, a Google translation speech library, and a local offline speech engine.

The packaging metadata has ten optional dependency groups. Four of them correspond to those services. The others are for recording, for translation, for transcription, and three that the readme's list does not mention at all.

Two of the omissions are the ones people look for. There is an extra for a commercial text-to-speech service that also has its own example file in the repository, and there is an extra for a large language model provider's speech service, also with an example file. Both appear in the search terms people use to find this project, and both are absent from the sentence that claims to list what is supported.

So the readme's list is a subset rather than the whole surface. That matters if you install by reading it, since the two services people search for are documented in the package metadata and demonstrated in the example gallery but not advertised in the readme.

The recommended service is named too, and it is the Google one rather than either of the omitted commercial services. That is a defensible default and also a hint that the maintainers consider the free and cheap options sufficient.

## Your own recording gets word timings because it gets transcribed

The fourth feature in the readme is the one that makes this plugin more than a narration track, and it is stated as a single clause with an explanation attached.

Per-word timing means you can trigger an animation at a specific word in the voiceover rather than at a timestamp you guessed. The readme adds that this works for the recordings too, not just the generated voices, and attributes it to a speech recognition model.

The attribution is the whole mechanism. If a generated voice comes from a text-to-speech service, the service can usually hand back timing information. A recording of your own voice cannot, so the only way to know where each word landed is to run recognition over what you just recorded.

That is why there is a transcription extra, and why it is heavier than the name suggests: a speech recognition model plus a second library built on top of it for stable word timestamps. Both are optional, because you only need them if you record rather than generate.

So the feature that differentiates this plugin from a plain text-to-speech call has a cost, and the cost is that your recorded narration is round-tripped through a recogniser before your animations line up. If your recording has any accent, background noise, or unusual proper nouns, that is where the timings will drift.

## The core animation library is the one dependency with no version bound

The runtime dependency list has eight entries and seven of them are bounded with an upper limit. One is not.

That one is the animation library the plugin extends. It is declared with no specification at all, which means a fresh resolve takes whatever the latest release is, and a future major version of the animation library will arrive in your environment without anything in this package objecting.

For a plugin that hooks into an animation framework's rendering and mobject API, that is the dependency where a version break would hurt most, and it is the one left open. Everything else in the list is careful: an audio toolkit, a dotenv loader, an audio metadata library, an audio manipulation library, a data validation library, and a filename slugifier, all with compatible release ranges.

The eighth entry is the odd one. The package manager itself is a runtime dependency, pinned to a minimum release. A library that needs to install something at import time or at runtime can express that by depending on the installer, and at least one of the extras here does need a compiled audio or speech package that a plain wheel resolve will not always satisfy.

The practical advice follows from the list. Pin the animation library yourself in your own project, because this package will not, and expect a resolver invocation somewhere in the workflow if you use the extras that build native code.

## An audio binary arrives as a Python dependency and the readme says nothing

One entry in the runtime dependency list is the name of a command-line audio program, bounded to a compatible range.

There is a Python package by that name, and installing it does not give you the audio binary. The binary is a separate system program that has to be present on the machine, and on a minimal container it usually is not.

This plugin processes audio: it synthesises or records a voice track, splices it, writes metadata, and hands the result to the renderer. Doing that well means shelling out to a real audio tool rather than pretending pure Python can do it.

So the expected first failure on a fresh machine is an audio-related error from inside a render, after the install succeeded and after the plugin was configured, with a message about a missing program rather than about a missing Python package.

The readme does not mention it. The installation section is a single link to the documentation site, and the requirements do not call out a system prerequisite the way they do for the extras that need a microphone or a speech service.

Two other entries point at the same class of problem in milder form. The audio metadata library and the audio manipulation library are pure Python and will install, but the recording extra needs a PortAudio binding and a global hotkey library, and PortAudio is itself a system library on most platforms. Installing that extra is therefore a two-stage operation and the packaging cannot do the first stage for you.

## Alpha status on a zero four release after a long quiet period

The release history has three entries and the gaps between them are the story.

The oldest is a post-release suffix on a zero three version from the spring of 2024. The next is a zero three version seven later. The newest is the zero four release, and it is recent enough that the branch was pushed within days of it.

So the project went from a zero three line to a zero four line with roughly twenty months in between, and the manifest still declares an alpha development status.

That combination is worth reading carefully. A minor version increment in this ecosystem usually signals a feature release, and the optional dependency set now includes two services and a transcription path that the readme's own list does not mention, so zero four did add surface. Declaring alpha at the same time is the maintainers telling you the interface may still move.

For a plugin, alpha is a harsher label than it is for an application, because a plugin is consumed as a library by code you wrote. If the voice service interface or the recording callback signature changes, your scenes break, and you are the one holding the upgrade decision.

The version number in the manifest matches the newest release, so at least the packaging metadata is not drifting here.

## One language classifier on a plugin whose selling point is translation

The classifier list is thorough about platforms and almost silent about language.

There are nine topic classifiers, covering scientific and engineering work, video, graphics, audio capture, audio synthesis, speech, and two that place the package in both artificial intelligence and visualisation. There are three Python version classifiers. There is one natural language classifier, and it declares English.

That is the classifier doing its job for a package whose documentation is in English, and it is also the least informative line in the metadata for this particular project. The plugin's translation feature uses a machine translation service to render the voiceover in other languages, and the example gallery includes a file whose name indicates a right-to-left script.

So the package can produce narration in languages it does not claim to be documented in, and a package index filtered by language will show it as English only.

The version classifiers have the opposite problem: three Python versions listed, and no upper bound in the requirement, so the metadata implies a range it does not actually state.

None of this affects the software. All of it affects whether you find it, and whether you find the right version of it, from an index rather than from a recommendation.

## Eleven examples, a second docs tree, and a path ownership file

The example gallery is the best documentation this project has, and it is worth reading in a particular order.

There are eleven entries. One demonstrates approximating a mathematical constant with a voiceover, which is the canonical use for this plugin. One demonstrates bookmarks, which is the per-word timing feature in its most direct form. One demonstrates recording from your own microphone. One demonstrates translation, and it is a directory rather than a file, which suggests it carries more than one language or more than one configuration.

Then there is one example per service: one for each of the four listed services and one each for the two the readme omits. That is the discovery path for those two.

The most useful filename in the gallery is the quadratic formula example in Arabic. It is the one artefact in the repository that demonstrates what happens to layout, font selection, and timing when the narration is not in a left-to-right Latin script, and it exists without being called out anywhere in the readme.

The repository also carries two documentation directories, one of them named as unverified, plus a file that configures a type checker, a file configuring some release or workflow tool with an unfamiliar name, a file assigning ownership by path, and an agent instruction file.

Ownership by path means contributions are reviewed by whoever owns that area, which is the mechanism behind the maintenance status of an alpha package with a single sponsoring organisation behind it.

## Conclusion

Manim voiceover is worth reaching for if you animate mathematical or technical material and want the narration to live in the same file as the scene graph, because the alternative is exporting a silent video and hand-timing an audio track in a separate editor. The per-word triggering is the part that changes how you write a scene. Two things to plan for before you install. SoX appears as an ordinary Python dependency while the actual audio binary is a system program the readme never mentions, so that is the first thing to check when nothing renders. And the core animation library is the one dependency with no version constraint at all, so pin it yourself rather than inheriting whatever a fresh resolve picks.

## FAQ

### What is manim-voiceover?

It is a plugin for the Manim animation library that adds narration to videos written in Python, so you do not need a separate video editor. You can generate voices from cloud or local speech services, or record your own through a command line interface during rendering.

### Which text-to-speech services does manim-voiceover support?

The readme lists four: Google speech, Microsoft Azure speech, a Google translation speech library, and a local offline engine. The package metadata ships ten optional dependency groups in total, including extras for a commercial speech service and for a large model provider's speech service, both of which have example files but are not in the readme's list.

### How does manim-voiceover time animations to speech?

Per word rather than per timestamp, so an animation can fire at a specific word in the voiceover. For generated voices the timing comes from the service. For your own recording it works because the audio is transcribed, which is why there is a separate optional extra holding a speech recognition model and a timestamp library.

### Does manim-voiceover need an audio program installed?

Yes, and the readme does not say so. An audio toolkit appears in the runtime dependency list as a Python package, but the program itself is a separate system binary that has to be present. The recording extra has the same shape of requirement, since it needs a PortAudio binding and a global hotkey library.

### What are manim-voiceover's dependencies like?

Seven of the eight runtime dependencies carry an upper bound, including the audio, metadata, validation, and filename libraries. The animation library the plugin extends has no version specification at all, and the package manager itself is listed as a runtime dependency with a minimum release.

## Sources

- [License: MIT](https://github.com/ManimCommunity/manim-voiceover/blob/main/LICENSE)
- [ManimCommunity/manim-voiceover on GitHub](https://github.com/ManimCommunity/manim-voiceover)
- [Project website](https://voiceover.manim.community/en/stable)
- [README](https://github.com/ManimCommunity/manim-voiceover/blob/main/README.md)
- [Releases](https://github.com/ManimCommunity/manim-voiceover/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/manimcommunity-manim-voiceover
