Open-source project
receptron/mulmocast-cli avatar
receptron/mulmocast-cli

MulmoCast CLI: turning MulmoScript JSON into video, podcast and slideshow output

AI-powered podcast & video generator.

475 stars82 forksTypeScriptLicense varies

At a glance

What is it?
MulmoCast is a TypeScript CLI that compiles a JSON script called MulmoScript into video, podcast, slideshow, PDF and manga formats using external AI providers. It is aimed at people who would rather write structured JSON than edit a timeline, and it needs an OpenAI key plus ffmpeg before it does anything.
Who is it for?
Adopt MulmoCast if your source material is already text and you want the same script rendered as video, podcast, PDF or manga without rebuilding it per format. Do not adopt it if you need a visual timeline editor or if you cannot put provider API keys on the machine.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem MulmoCast solves is format, not content

Most presentation tools assume a human is arranging slides. MulmoCast assumes a language model wrote the content and a program will render it. The unit of work is a JSON or YAML document called MulmoScript, and the README describes it as an intermediate language that "functions like a screenplay or web markup". That framing matters: the script is not the output, it is the input to a renderer that can emit several outputs.

The README's own diagram shows the loop. A creator talks to an LLM, the LLM produces MulmoScript, and MulmoCast turns that script into video, podcast, slideshow, PDF, manga or swipe anime. The audience is therefore narrow and specific. It is someone who already has text (a research report, a pitch, a children's story) and wants it in more than one medium without redrawing anything. If you enjoy placing text boxes by hand, this is the wrong tool for you.

MulmoScript is the contract between the model and the renderer

A MulmoScript document carries a version block and an array of beats. The README's Hello World is the whole specification you need to start:

json
{
  "$mulmocast": {
    "version": "1.1"
  },
  "beats": [
    { "text": "Hello World" }
  ]
}

Everything else hangs off that structure. Speakers, images and layout are declared in the same document, and provider settings such as imageParams and speechParams sit alongside them. That is the design decision worth noticing: model selection is data, not a command-line flag. A script can name provider google with model gemini-2.5-flash-image, or provider openai with model gpt-image-1.5, and the renderer routes accordingly.

The cost of that decision is that scripts are not portable across providers by accident. Azure OpenAI is supported, but the README states plainly that Azure deployment names must match model names exactly, giving the example of a deployment named gpt-image-1.5 for a model named gpt-image-1.5. A script moved between accounts can fail on naming alone, before any generation happens.

Installing MulmoCast and producing a first podcast

Installation is two steps: the npm package and ffmpeg. The README gives the global install and the Homebrew line for macOS, pointing other platforms at the ffmpeg download page.

bash
npm install -g mulmocast
bash
brew install ffmpeg

Before anything runs, the CLI needs credentials. Create a .env file in your project directory. OPENAI_API_KEY is the only required entry; GEMINI_API_KEY, ANTHROPIC_API_TOKEN, REPLICATE_API_TOKEN and ELEVENLABS_API_KEY are optional and unlock other providers.

bash
OPENAI_API_KEY=your_openai_api_key

The package installs three binaries: mulmo, mulmocast and mulmo-mcp. The README's Docker example uses mulmo with the tool scripting subcommand, an input flag, a template name, an output directory and a story flag:

bash
docker run -e OPENAI_API_KEY=<your_openai_api_key> -it mulmo-cli mulmo tool scripting -i -t children_book -o ./ -s story

If you prefer Docker, the repository's Dockerfile is built on node:24, installs ffmpeg via apt, installs mulmocast globally and optionally adds the Google Cloud CLI for Google's image models. Build it with docker build -t mulmo-cli . as the README shows. What you should see after a successful run is generated media in the output directory you passed with -o, plus, if you set MULMOCAST_DUMP_USAGE, a JSON usage dump on stdout or at the path you gave.

Usage tracking is the feature most CLI tools forget

Generating video and audio through paid APIs means the cost of a script is not obvious until the bill arrives. MulmoCast addresses this with two environment variables. Setting MULMOCAST_DUMP_USAGE=1 prints a JSON usage dump to stdout after each CLI action; pointing it at a path instead, as in MULMOCAST_DUMP_USAGE=/tmp/usage.json, writes the same dump to a file. The README says the dump reports token, per-second and per-char usage per provider:model.

That granularity is the useful part. A script that mixes a cheap text model with an expensive image model produces a breakdown you can attribute, rather than a single number. The README does not describe a budget cap or a hard stop when usage crosses a threshold, so this is measurement, not cost control. You still have to read the dump and decide.

Where MulmoCast gets in your way

The dependency chain is the first real constraint. ffmpeg must be installed separately, and on anything other than macOS with Homebrew the README sends you to the ffmpeg download page rather than giving a command. Container users get it for free through the Dockerfile; everyone else assembles it.

The second constraint is credentials. There is no local-only path described in the README. Every generation route it documents runs through a hosted provider: OpenAI, Google Gemini, Anthropic Claude, Replicate or ElevenLabs. If your organisation forbids sending source material to third-party APIs, MulmoCast as documented cannot be used, regardless of how good the script format is. The related searches around local video generation with Ollama point at a need this project does not claim to meet.

The third is the beta caveat at the top of the README. The Quick Start section tells readers who want to try the beta version to follow the release notes in docs/beta1_en.md and docs/beta1_ja.md, which means the documented entry point and the current release line are not the same thing. Anyone following the README literally should read those notes first.

Finally, the licence is not stated anywhere in the repository listing or package.json, which contains no license field and no LICENSE file at the top level. Treat the terms as unknown until you check the repository directly.

How it compares with Podcastfy and GraphAI

Podcastfy is the closest comparison in the related searches, and the difference is scope. Podcastfy is a podcast generator. MulmoCast's README positions the same script as a source for video, podcast, slideshow, PDF, manga and swipe anime, which means the script format has to carry enough structure to satisfy all of them. That breadth is the argument for choosing MulmoCast, and also the reason its script format is more elaborate than a two-speaker dialogue file.

GraphAI appears in the same search results and sits at a different layer. MulmoCast is a CLI with an MCP server binary, mulmo-mcp, in its bin map, so it can be driven by an agent. GraphAI is a workflow engine. Choosing between them is choosing whether you want a finished renderer or a graph to build your own. MulmoCast also exposes a programmatic surface: package.json exports a browser entry point and a tools/complete_script module alongside the Node entry, so it is usable as a library, not only as a command.

Maintenance, upgrades and what the release history shows

The last push to the repository was on 2026-08-23, and release 2.12.0 carries the same date. Release 2.11.0 and @mulmocast/[email protected] landed earlier the same day. The repository is not archived. That is a same-day cluster of package and types releases, which suggests the types package is versioned in step with the CLI rather than independently.

Upgrade cost is mostly the MulmoScript version field. The Hello World pins "$mulmocast": { "version": "1.1" }, and the presence of CHANGELOG-0.x.md, CHANGELOG-1.x.md and CHANGELOG.md in the repository root shows the format has already moved through major revisions. A script written against 0.x is not automatically a 1.x script. Before upgrading, read the changelog that matches your current version rather than the newest one.

The second upgrade cost is provider drift. Because model names live in the script and, for Azure, must match deployment names exactly, a provider renaming or retiring a model can break a script that was working. Keeping provider identifiers in one place per project limits the blast radius.

Editorial conclusion

Adopt MulmoCast if your source material is already text and you want the same script rendered as video, podcast, PDF or manga without rebuilding it per format. Do not adopt it if you need a visual timeline editor or if you cannot put provider API keys on the machine. Before committing, verify three things: that ffmpeg is on PATH, that OPENAI_API_KEY is set in a .env file in the working directory, and that the model names in your MulmoScript match the provider deployments you actually have.

Frequently asked questions

How do I install MulmoCast CLI?

Install the npm package globally with npm install -g mulmocast, then install ffmpeg separately. The README gives brew install ffmpeg for macOS with Homebrew and points other platforms at the ffmpeg download page. A Dockerfile in the repository installs both.

Which API keys does MulmoCast need?

OPENAI_API_KEY is the only required key, placed in a .env file in your project directory. GEMINI_API_KEY, ANTHROPIC_API_TOKEN, REPLICATE_API_TOKEN and ELEVENLABS_API_KEY are optional and enable Google image and TTS, the htmlPrompt feature, movie models and ElevenLabs TTS respectively.

What is MulmoScript and what does a minimal script look like?

MulmoScript is the JSON or YAML format MulmoCast renders, described in the README as an intermediate language that works like a screenplay or web markup. The minimal example declares a $mulmocast version of 1.1 and a beats array containing a single object with a text field.

Can I use MulmoCast without sending data to a cloud provider?

The README documents only hosted providers: OpenAI, Google Gemini, Anthropic Claude, Replicate and ElevenLabs. It describes no local model path, so there is no documented way to run generation entirely offline.

Official sources

  1. Official README
  2. Project repository
  3. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/receptron-mulmocast-cli.svg)](https://hysenlabs.com/projects/receptron-mulmocast-cli)