Video Material GEN Workstation: a management shell for short video production, written by someone who tells you it is mostly a management tool
A short-video generation workstation combining content planning, automatic AI copywriting, batch TTS voiceover, AI image material synthesis, ASR-based subtitle extraction, and free-form AI creation, making it easy to manage each video project.
At a glance
- What is it?
- A Chinese-language project that batches together copywriting, TTS dubbing, image generation and subtitle extraction per video episode. The author credits borrowed code in two places and labels the Docker path as buggy.
- Who is it for?
- Treat this as a working folder with a web front end, not as a pipeline product. Its real value is showing one author's arrangement of the pieces: a project card per episode, structured scene scripts, batch TTS with emotion prompts, prompt reuse for image generation, and subtitle extraction, all sharing one directory on disk.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What the project is, in the author's own framing
Video Material GEN Workstation is a workstation for producing short videos. The README's opening line lists what it combines: content planning, AI-assisted copywriting, batch TTS dubbing, AI image material composition, ASR extraction of speech subtitle scripts, and free-form AI creation, all in one place for managing per-episode video projects.
The author is direct about what this is and is not. A later section states that the project is mainly oriented toward management, put plainly that the features will not be as practical as you imagine, and that defining video content and making something popular still requires thinking. That candor is rare and it changes how you should evaluate the repository. You are not getting an automated content factory. You are getting a consistent workspace where the assets for one episode live together.
The quick overview at the top of the README adds three specifics. Projects can be generated in bulk from templates, with scripts, AI image materials, subtitles and audio produced together. Text-to-speech uses Gemini plus TTS synthesis, so the same setup can rewrite a script or output emotional-sounding voice directly. Images, subtitles and audio are managed on separate tracks, so you can replace any one of them in the interface and preview the result.
That last point is the real design idea. Keeping the three media types independent means a bad generated image does not cost you the whole episode, you swap the one asset and re-preview.
Deploying from source: one dependency and one config file
The source deployment path is the one the author documents most fully. The steps are to copy the example config, fill in your keys, install dependencies, and start the server.
cp env.example.yaml env.yamlYou then fill in a Gemini key, a base URL, a model name, a TTS key and the prompt templates, all of which live in `env.yaml`. Without these the interfaces cannot be called, as the README states. One optional setting matters more than the rest: `Default-Project-Root`, which is where generated scripts, audio and image files are written. Set it to a path you control if you care about where the output lands.
npm install
npm startOr double-click `start.bat` on Windows. Either way the server comes up on port 8765 at `http://localhost:8765`.
What makes this easy to pick up is the dependency count. The package manifest declares exactly one runtime dependency, `js-yaml` at 4.1.0, with `server.js` as the entry point, an engine requirement of Node 12 or newer, and a version of 1.0.0. There is no framework, no build step and no client-side bundler in the manifest. The tree confirms it: `server.js`, `index.html`, a `js/` directory, a `css/` directory, an `asr/` directory, `env.example.yaml`, `start.bat` and a small `img/` folder of screenshots. A web server that reads YAML and serves static files, which is the entire deployment story.
That also means the Node version requirement in the manifest, 12 or newer, is far below what the Docker image actually uses, which is `node:20-alpine`.
The Docker path, and the README's own warning about it
The Docker section exists, and its heading says in plain terms that there are currently bugs. The author put the warning in the section title rather than in a footnote, which tells you how much to trust it.
The documented flow is to copy the config, set `Default-Project-Root` to `/data/projects` inside the container so it maps to the local `./data` directory for persistence, then start with:
docker compose up -d --buildThe compose file does the mapping work for you. It binds `./env.yaml` to `/app/env.yaml` read-only so the container cannot rewrite your credentials, binds `./data` to `/data` for persistence, publishes port 8765, sets `NODE_ENV=production` and restarts the service unless stopped.
services:
video-workstation:
build: .
container_name: video-workstation
ports:
- "8765:8765"
environment:
- NODE_ENV=production
volumes:
- ./env.yaml:/app/env.yaml:ro
- ./data:/data
restart: unless-stoppedThe image build is a four-step Alpine sequence: Node 20 Alpine as the base, `/app` as the working directory, `npm ci --omit=dev` after copying just the package files so the dependency layer caches across rebuilds, then the rest of the source copied in, port 8765 exposed and `npm start` as the command.
Two operational notes from the README deserve attention. If the Node image fails to pull, the author recommends pulling `node:20-alpine` first and then re-running the compose command. And the container has no desktop, so buttons like open project directory or open TTS folder will not launch a file manager; the interface returns the path and you navigate there yourself on the host.
Two features that belong to other people
This is the section that should shape your expectations, because the author documents it clearly instead of burying it.
Subtitle extraction depends on the author's separate project, an n8n HTTP tools repository linked directly from the feature list. The README explains that the previous n8n workflow file could not be found, so that part should be treated as unavailable, and that the copywriting feature needs n8n to operate. In other words, two of the eight listed features depend on an external toolchain that is not bundled here.
The ASR-based automatic extraction of subtitle files is described as implemented by reverse-engineering an interface, and the README states that this part of the code was open-sourced by another author. So the subtitle extraction feature is credited code, not original work by this repository's author.
Both disclosures are to the author's credit. They are also the reason to be careful about maintenance. If the n8n side of subtitle generation is not runnable from this repository, then the subtitle feature is a documented integration point rather than a working feature, and you should verify it works before designing around it.
For image generation, the README names the model as NanoBanana, reached through a locally deployed reverse proxy for the AIStudio interface. The author's assessment is that it is stable enough to generate images that then feed into Sora. Note what that implies: the image pipeline expects a self-hosted proxy you supply, not a hosted endpoint the project configures for you.
The eight features, and where the value actually sits
Walking the README's feature list gives a fair picture of what you get, in the author's order.
Project overview presents projects as cards you manage in bulk, showing the output directory, creation time and a delete action. Copywriting generation displays structured scene scripts, lets you copy a single line or the whole script, and links checkbox selection on the left to the prompt on the right. Subtitle extraction is the n8n-dependent feature. TTS synthesis supports both single and batch modes, taking synthesis text plus an emotion prompt. Image generation centralizes character descriptions and scene descriptions as prompts that can be checked and copied in bulk to an image task.
Character art and background generation takes a prompt, a reference image upload, an aspect ratio setting and a history record, so assets can be reused. The ASR subtitle feature sits below the TTS interface in the UI and opens a separate subtitle tool. Prompt library and free creation round it out, with a place to save frequently used prompts for one-click copying plus a free creation panel.
The author then says the remaining features were not written out one by one and invites you to deploy it and see. Taken together the list describes an asset assembly surface rather than an automation engine: the value is that prompts, reference images, aspect ratios, generated audio and generated images are all attached to a per-episode record instead of living in scattered folders.
If your current process already keeps episode assets organized this way, the gain is modest. If it does not, the organization is probably the feature you would notice first.
Licensing, language labeling and what to check first
Two metadata details do not match the repository contents, and both are worth resolving before you build on this.
The package manifest declares `license` as MIT, which is a clear and permissive statement. The license field in the repository metadata is empty, and there is no LICENSE file in the tree listing. When a manifest says MIT and there is no license text, the manifest claim is still the only license information available, and it is worth confirming with the author if you plan to redistribute.
The repository's primary language is recorded as Python, but the application runs on Node. `server.js` is the entry point, the Dockerfile is Node 20 Alpine, the only dependency is js-yaml, and the build runs `npm install` and `npm start`. Python appears only in the `asr/` directory, which is consistent with the README's statement that the automatic speech recognition code came from another author and was likely contributed in Python.
So when you read the language label, expect a small Node web server with one borrowed Python component, not a Python application. That also tells you the security surface: js-yaml parses your config, server.js serves it, and the asr component runs separately.
There is also a disclaimer at the end of the README stating the project is provided for reference and study, and that the author accepts no responsibility for problems arising from its use. Combined with the buggy-Docker note and the missing n8n workflow, the sensible first step is the source install on a spare directory, with `Default-Project-Root` pointed at a throwaway path so you can see the output structure before committing to it.
Editorial conclusion
Treat this as a working folder with a web front end, not as a pipeline product. Its real value is showing one author's arrangement of the pieces: a project card per episode, structured scene scripts, batch TTS with emotion prompts, prompt reuse for image generation, and subtitle extraction, all sharing one directory on disk. Two things to know before you spend an afternoon on it. The Docker path is labeled as having bugs in the README heading itself, so start from the source install unless you need containers. And two features are credited elsewhere: subtitle generation depends on the author's separate n8n tooling project, and the ASR extraction code was written by another author. The licensing is also uneven, with the package manifest declaring MIT while the repository license field is empty. The last push was on 2026-06-02.
Frequently asked questions
How do I run this locally?
Use the source path. Copy `env.example.yaml` to `env.yaml`, fill in the Gemini key, base URL, model, TTS key and prompts, set the optional `Default-Project-Root`, then run `npm install` followed by `npm start`. The server comes up on port 8765. The Docker path exists but the README's own heading for it notes there are currently bugs.
Does the subtitle extraction feature work out of the box?
Two separate features are involved and neither is self-contained. Subtitle extraction is described as depending on the author's separate n8n HTTP tools project, and the README notes the previous n8n workflow file could not be found so that part should be ignored. The automatic speech recognition feature that produces subtitle files is credited to code open-sourced by another author.
Which image model does it use?
The README names NanoBanana, reached through a locally deployed reverse proxy for the AIStudio interface that you supply yourself. The author's note is that it is stable enough to generate images that then feed into Sora.
What license does the project use?
The package manifest declares MIT. The license field in the repository metadata is empty and no LICENSE file is present in the tree, so the manifest declaration is the only license statement available. Confirm with the author before redistributing.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/norsico-video-materials-autogen-workstation)