Rhubarb Lip Sync: turning voice recordings into 2D mouth animation from the command line
Rhubarb Lip Sync is a command-line tool that automatically creates 2D mouth animation from voice recordings. You can use it for characters in computer games, in animated cartoons, or in any other project that requires animating mouths based on existing recordings.
At a glance
- What is it?
- Rhubarb Lip Sync analyzes an audio file, recognizes the speech, and writes out mouth-shape data for six to nine standard 2D mouth positions. It is a CLI first, with integrations for After Effects, Moho, OpenToonz, Spine and Vegas Pro.
- Who is it for?
- Adopt Rhubarb Lip Sync if your characters use the six-to-nine Hanna-Barbera-style mouth shapes and you want mouth data generated from an existing recording rather than keyed by hand; the CLI plus a TSV, XML or JSON output drops into most 2D pipelines. Do not adopt it if you need real-time lip sync from a live microphone, if your character has a 3D rig driven by visemes or phoneme curves, or if you cannot ship a compiled binary alongside your project.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 107 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Rhubarb Lip Sync actually produces
The project solves a narrow, repetitive problem in 2D animation: given a voice recording, decide which of a small set of mouth drawings should be on screen at each moment. The README describes the pipeline as analyzing audio files, recognizing what is being said, and then automatically generating lip sync information. The output is not a rendered animation. It is data that names a mouth shape over time, which your animation tool then applies to a character that already has those drawings.
The set of drawings is fixed and comes from traditional television animation. Six basic shapes (A through F) were invented at Hanna-Barbera for shows such as Scooby-Doo and The Flintstones, and the README calls them a de-facto standard for 2D animation. Three more (G, H and X) are optional: G for upper teeth on lower lip, H for a visible tongue on long L sounds, X for a relaxed idle mouth during pauses. That last one matters more than it looks. If your character is standing in a scene not talking, X is the drawing you want, and the README notes it is almost identical to A but with less pressure between the lips.
Who it is for: small animation teams, game developers and solo creators who already have a mouth chart drawn and want the timing filled in. It is not a character animation system, and it does not draw anything.
How the audio becomes mouth-shape data
The tool takes an audio file as its main argument and writes an output file whose format you choose. The README's simplest invocation is `rhubarb -o output.txt my-recording.wav`. Everything between reading the recording and writing the file is speech recognition plus a mapping step: recognized sounds are assigned to the mouth shapes whose descriptions in the README match them. The shape table is explicit about this. B is for most consonants and the EE vowel, C for EH and AE, D for AA as in father, E for AO and ER, F for UW, OW and W. G covers F and V, H covers long L, X covers silence.
That mapping is why the output is a sequence of shape letters rather than a waveform or a phoneme string. The tool has already made the artistic decision about which drawing to show. Your job is to make those drawings and to make the transitions between them look right. The README gives two specific notes on transitions: C is used as an in-between when animating from A or B to D, and E is used as an in-between from C or D to F. It warns that the mouth should not be wider open for E than for C. Those are constraints on your artwork, not on the software.
Output formats are TSV, XML and JSON through the CLI. For Moho and OpenToonz the tool can write .dat switch data files instead, with `--datFrameRate` controlling the frame rate and `--datUsePrestonBlair` controlling the shape names. That last flag is the clearest sign of who the tool is aimed at: it assumes the reader knows what a Preston Blair mouth chart is.
Installing Rhubarb Lip Sync and running a first recording
There is no package manager step in the README. The instructions say to download the latest release for your operating system from the GitHub releases page and unpack the file anywhere on your computer. Windows, macOS and Linux builds are listed. Once unpacked, the `rhubarb` executable is called directly from the command line.
The minimal run passes an audio file and an output path. The README's own example writes a text file next to nothing else:
rhubarb -o output.txt my-recording.wavAfter that command, `output.txt` holds the generated lip sync data. If you want a format your engine can parse directly rather than a plain text file, the CLI supports TSV, XML and JSON output, and the README points to the options list for the flags that select them.
For a Moho or OpenToonz workflow, the output is a .dat switch data file rather than a text file. The README names two options for that path: `--datFrameRate` to set the frame rate, and `--datUsePrestonBlair` to switch the shape naming.
If you would rather not touch a terminal, the repository ships graphical and plugin front ends under `extras/`. There is an Adobe After Effects integration in `extras/AdobeAfterEffects`, a Spine tool in `extras/EsotericSoftwareSpine`, and two Vegas Pro plugin scripts in `extras/MagixVegas`. Each has its own README.adoc in the download. The Spine integration is described as a graphical tool that imports a Spine project, performs the lip sync, and re-imports the result into Spine.
Where the approach breaks down
The mouth shape set is the biggest constraint, and it is a design choice rather than a bug. Six to nine discrete drawings means the animation cannot show a mouth partway between two shapes unless you draw the in-between yourself, which is exactly why the README tells you to make A-C-D and B-C-D transitions smooth. A production that wants continuous mouth deformation, or that uses a phoneme set outside the Preston Blair tradition, will spend its time fighting the output rather than using it.
The second constraint is the input. The README frames the tool around existing recordings: you bring a voice file, it produces data. Nothing in the README describes live microphone input, streaming audio, or real-time operation, so a game that needs the mouth to move as a player speaks into a mic is outside what this tool is documented to do. The workflow is offline and file-based.
The third is the delivery model. The README's install path is a downloaded archive per operating system, and the repository is a C++ project built with CMake. There is no documented package manager install. If your build pipeline expects to pull a dependency from npm, pip or a distro repository, you will be adding a binary artifact and a version pin to your own release process instead.
Finally, the project does not draw. If you have no mouth chart, the tool's output names shapes you do not have. That is a real prerequisite, not a formality.
Rhubarb Lip Sync compared with an engine's built-in viseme track
The closest alternative for many teams is not another standalone tool but the animation engine's own audio-driven viseme system, which maps audio to a viseme set defined by that engine and applies it to a 3D facial rig. The difference is in what is being named. Rhubarb Lip Sync names 2D drawings from a hand-drawn chart; an engine viseme system names blend shapes on a mesh. If your character is a 3D model with a jaw rig, Rhubarb's output is the wrong vocabulary, and the search interest in Rhubarb Lip Sync combined with 3D or Blender reflects that mismatch.
Against a manual workflow, the difference is time and consistency. Keying mouth shapes by hand gives an animator full control over emphasis and timing, which matters for stylized delivery. Rhubarb Lip Sync trades that control for speed and repeatability: re-record a line, re-run the command, get new data. For a game with hundreds of short voice lines, that trade is usually worth it. For a handful of hero lines in a short film, an animator may prefer to key them.
The integrations are the other axis. Because the repository ships adapters for After Effects, Moho, OpenToonz, Spine and Vegas Pro, plus a community integration with Visionaire Studio, the practical comparison is often between Rhubarb Lip Sync and whatever lip sync feature already exists inside the tool you are animating in. If that tool has no such feature, Rhubarb fills a gap rather than replacing something.
Maintenance, licence and upgrade cost
The repository is not archived, and the last push was on 2026-06-16. That is recent enough that the project is not abandoned, but the release history is uneven: v1.14.0 shipped on 2025-04-03, v1.13.0 on 2022-06-14, and v1.12.0 on 2022-03-12. Three years separated the two most recent releases, so pinning to v1.14.0 and treating upgrades as an event rather than a stream is the realistic posture.
Upgrade cost depends on which surface you use. The CLI's documented options are stable and few, and the output formats are plain text, so a CLI integration is cheap to move between versions. The `extras/` integrations are more exposed: each is a separate adapter with its own README.adoc, and the After Effects, Spine and Vegas Pro adapters depend on host applications that version independently of Rhubarb. A host application update is as likely to break the adapter as a Rhubarb update is.
The licence is the item to check before anything else. The repository metadata reports NOASSERTION rather than a recognized SPDX identifier, and there is a LICENSE.md at the top level. Read that file directly and have whoever handles licensing at your organisation confirm what it permits for your distribution model. This article cannot tell you what the licence allows, and the metadata alone is not enough to assume a permissive one.
Editorial conclusion
Adopt Rhubarb Lip Sync if your characters use the six-to-nine Hanna-Barbera-style mouth shapes and you want mouth data generated from an existing recording rather than keyed by hand; the CLI plus a TSV, XML or JSON output drops into most 2D pipelines. Do not adopt it if you need real-time lip sync from a live microphone, if your character has a 3D rig driven by visemes or phoneme curves, or if you cannot ship a compiled binary alongside your project. Before committing, verify three things: that the audio you plan to feed it matches what the tool expects, that your target integration (After Effects, Moho, OpenToonz, Spine or Vegas Pro) is one of the ones the repository ships an adapter for, and that your character's mouth drawings match the shape semantics the README describes, since the tool names shapes rather than drawing them.
Frequently asked questions
How do I install Rhubarb Lip Sync?
Download the latest release for your operating system from the GitHub releases page and unpack the file anywhere on your computer. The README lists Windows, macOS and Linux builds, and the tool is then called as `rhubarb` from the command line. No package manager install is documented.
How do I use Rhubarb Lip Sync?
Pass it an audio file and an output path, for example `rhubarb -o output.txt my-recording.wav`. The tool analyzes the recording, recognizes the speech, and writes lip sync data naming mouth shapes over time. Additional command-line options select other output formats and control settings such as the .dat frame rate.
Is Rhubarb Lip Sync free?
The repository metadata reports the licence as NOASSERTION and there is a LICENSE.md at the top level, so the terms are not identifiable from the metadata alone. Read LICENSE.md in the repository for the actual terms.
Is there a lip sync add-on for Blender built on Rhubarb Lip Sync?
The README lists integrations with Adobe After Effects, Moho, OpenToonz, Spine by Esoteric Software, and Vegas Pro by Magix, plus an external Visionaire Studio link. Blender is not among them. Rhubarb Lip Sync targets 2D mouth shapes, so a 3D Blender rig would need its own viseme mapping.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/danielswolf-rhubarb-lip-sync)