Model or dataset
alibaba/lumenx avatar
alibaba/lumenx

LumenX is a local-first pipeline wrapping a dozen hosted video models

LumenX Studio,AI 短漫剧一站式生产平台。它能够将小说文本转化为动态视频,打通了从剧本分析、角色定制、分镜绘制到视频合成的完整创作链路。

1,297 stars303 forksTypeScriptMIT

At a glance

What is it?
The interesting engineering in this project is not the video generation, it is the model catalogue: a YAML directory compiled to JSON, an adapter layer per provider, and a capability matrix so the interface can say which models support text-to-image, image-to-video or reference-to-video. Everything the tool actually generates comes from someone else's API and the readme is clear about which key buys what.
Who is it for?
Adopt LumenX if you want to assemble a short comic-style video from a script and want the orchestration local so your source text and generated assets stay on your machine, because the pipeline stages and the export step are the parts that are genuinely yours. Do not adopt it expecting it to generate video by itself, since every model in the catalogue is a hosted service you pay for, and one API key is the mandatory configuration.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 50 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Two products in one repository, and the split is deliberate

The readme describes two modules and the table says the split is about context. The first is a pipeline-first production tool for short comic-style drama, running script to storyboard to assets to video to compositing to export, and it assumes you have a script. The second is a standalone generation workbench that needs no script context and works immediately, offering six generation modes: image generation, text-to-video, image-to-video, reference-to-video and video editing. The distinction matters because the second is a thin interface over the first's adapters, and that is where the interesting code is. The first is a pipeline with stages, which is orchestration. The second is a catalogue browser, which is a capability matrix and a form. If you are evaluating this project, the question is which of the two you need, because the pipeline without a script in hand is a lot of machinery, and the workbench without the pipeline is a nice way to try a dozen video models side by side. The repository layout confirms the split at the source level, with a backend directory for each, a models directory holding the adapters, an audio directory, a model catalogue configuration directory and an output directory. Two applications, one model layer, one catalogue.

The catalogue is the actual architecture

The model catalogue is described as a YAML directory compiled to JSON, with a documentation directory dedicated to it, including an onboarding guide for adding a new model and a design document for the catalogue architecture with a dated filename. That is more infrastructure than a hobby project usually builds, and it is the right infrastructure here, because the readme's model table is really a capability matrix. Each row names a provider, a model and the operations it supports, and the operations are a fixed vocabulary: text-to-image, image-to-image, image-to-video, reference-to-video, text-to-video and video-to-video. A model that supports only the last two will render controls it cannot honour, and a model with a four-times-higher resolution for images than for video will need per-model parameters, which the readme lists as a feature in its own right. So the catalogue carries three things per model: which operations are available, what parameters that model takes, and how to reach it. The adapters directory names four of them after providers or routers, which is a small number for a dozen models, and suggests the adapters are grouped by access path rather than by model. A team adding their own model is pointed at a written guide, which means the extension point is documented rather than reverse-engineered from the code.

One key, four upgrades, and the provider matrix underneath

The configuration section is the most honest part of the readme and deserves to be read as a cost model. The architecture is described as local-first with a minimum configuration of one API key, and there is a table of five configuration levels. The base level needs a single key and covers several image and video models plus the speech models, with two of the video models marked as proxied. The next level adds a second provider through either a login command or a key in a particular format, and brings two more models including one with high-resolution image output. Then two levels add direct access to two video providers with their own credential pairs, and the last adds object storage credentials for cloud media mirroring and signed URLs. So the readme is telling you that a single key gives you a usable subset and that direct access to the best video models means separate accounts with separate billing. That is worth internalising before installing, because the difference between proxied and direct is not a technical detail, it is a rate limit and a price. Two configuration locations are documented: a file in the project root for development, and a settings page that saves to a per-user configuration path. The second is what makes the local-first claim useful, since the tool can run without a project checkout.

The pipeline stages, and the compositing step that is not a model call

The studio capability list is short and ordered like a production schedule. Script analysis, where a language model extracts characters, scenes and props and produces a structured storyboard script. Art direction, where you set a visual style and the whole film is held consistent. Asset generation across multiple models, covering three-view character sheets, scene establishing images and prop reference images. Video generation from the storyboard, with two modes named for image-to-video and reference-to-video, plus a batch feature described as drawing many candidates. Voice synthesis with two named speech model families for multi-voice dialogue. Then compositing, described as a timeline editor plus a command-line tool stitching the pieces into a finished file. That last stage is the one that is not a model call and is therefore the part most likely to surprise. Reading the project layout tells you the video work is done by an external tool installed on the system, which means this project orchestrates media it does not produce and hands it to a separate encoder. The consequence is a hard dependency on a system-level binary and on the two matching, and the requirements file shows the project is not shy about heavy Python dependencies for media handling. The pipeline is a reasonable decomposition of the problem, and the stages are the ones a studio would recognise.

Two runtimes, two dev servers, and a desktop wrapper

The quick start tells you the stack before it tells you the commands: a Python runtime at a stated floor, a Node runtime at a stated floor, and a video tool installed system-wide. The one-line start is:

bash
npm run dev

The single development command starts the backend, the frontend and a browser launcher together, with a pre-step that runs a setup script, and the compose file shows the same split as two services where the backend gets its environment from a file and the frontend is served by a minimal image. The separate-path alternative installs Python requirements and runs a shell script for the backend, then installs and runs the frontend in its own directory. So a development run is two servers on two ports, and the readme documents the addresses for both applications and for the generated API documentation. The desktop story is separate and larger. The repository carries a Tauri configuration directory, a build template for packaging, a build script for macOS, one for a sidecar binary, a PowerShell script for Windows, an icon file in the macOS bundle format, a development container directory, two Dockerfiles, and a directory of packaging hooks. That is a project intending to ship a desktop application, not only a web one. The compose file also mounts a local output directory into the backend, which is the practical meaning of local-first: your generated media lands on your disk, not in someone else's bucket, unless you configure object storage.

Editorial conclusion

Adopt LumenX if you want to assemble a short comic-style video from a script and want the orchestration local so your source text and generated assets stay on your machine, because the pipeline stages and the export step are the parts that are genuinely yours. Do not adopt it expecting it to generate video by itself, since every model in the catalogue is a hosted service you pay for, and one API key is the mandatory configuration. Four things to verify. Which models you will actually use, because the catalogue mixes direct provider access with two proxied through one key and two that need their own credentials, and the capability matrix differs per model. What the local-first claim means in practice, since it describes where configuration and output live rather than where inference happens. Whether a desktop build is what you want, because the repository carries a Tauri shell, a build template for packaging, and build scripts for two desktop platforms in addition to the browser path. And whether the audio dependency surprises you, since the requirements include a source separation library and a sound file library for what the readme describes as voice synthesis. The licence is MIT, there are no published releases, and the last push was on 2026-08-11.

Frequently asked questions

What does LumenX do and what are its two modules?

It turns creative text into a publishable animated video. LumenX Studio is the pipeline-first production path, running script analysis, storyboarding, asset generation, video generation, compositing and export. LumenX Playground is a standalone generation workbench that needs no script context and offers six modes including image generation, text-to-video, image-to-video, reference-to-video and video editing.

Which AI models does LumenX support?

The readme lists video and image models from one cloud platform including a video model, an image model and two speech model families, plus two video providers available both through that platform and directly, one other provider reached through a router, and a language model used for script analysis and prompt polishing. Each row states which of the six operations the model supports.

What configuration does LumenX need to run?

One API key is mandatory, which gives access to a subset of the image and video models plus speech. The readme's table lists four further levels: adding a second provider via a login command or a key, adding direct access to each of two video providers with their own credential pairs, and adding object storage credentials for cloud media mirroring and signed URLs. Configuration can live in a project file or be saved from a settings page to a per-user configuration path.

How do I start LumenX locally?

Clone the repository, copy the example environment file and fill in the required API key, then run the single development command, which starts the backend, the frontend and a browser launcher together. You can also start the two separately by installing the Python requirements and running the backend script, then installing and running the frontend in its own directory.

Does LumenX have a desktop application?

The repository carries a Tauri configuration directory with development and build scripts, a build template for packaging, a macOS build script and a Windows build script, an icon file, packaging hooks and two Dockerfiles, alongside a compose file for the two-service web path. The frontend is described as a Next.js application.

What licence is LumenX released under?

MIT. The repository publishes no GitHub releases, and the last push to the main branch was on 2026-08-11. The project requires Python 3.11 or later, Node 18 or later, and a system-level video processing tool.

Official sources

  1. alibaba/lumenx on GitHub
  2. Issues
  3. License: MIT
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/alibaba-lumenx.svg)](https://hysenlabs.com/projects/alibaba-lumenx)