Pixelle-Video's dependency list has two exact pins, and one of them is there because of a bug
🚀 AI 全自动çŸè§†é¢‘引擎 | AI Fully Automated Short Video Engine
At a glance
- What is it?
- Pixelle-Video turns a topic into a finished short video: it writes the script, generates an illustration or video clip per line, synthesises narration, adds music and composes the result. Windows users get a bundled build; everyone else installs uv and ffmpeg and runs a Streamlit app. Behind the automatic promise is a dependency list where almost everything is a floor and exactly two things are pinned.
- Who is it for?
- Pixelle-Video fits someone who publishes short explainer-style videos and wants the mechanical work done for them, and who is comfortable handing an API key to several different providers at once. Do not adopt it expecting a single model to do the work, because the design is explicitly swappable at every stage and you will be configuring more than one service.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 109 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
One topic in, a finished video out
The pitch is a single sentence: enter one topic and the system does the rest.
The five things it claims to handle are writing the video script, generating AI images or video, synthesising the voice narration, adding background music, and composing the video with one click. The README describes the result as zero barrier and zero editing experience, making video creation a one-sentence job.
Each of those five is a separate capability with its own provider, and the configuration screen reflects that. You are asked for a language model, for a ComfyUI address or a RunningHub key if you want workflow-based generation, and for direct image and video model providers including DashScope, OpenAI, Seedream, Seedance and Kling.
Speech is likewise a choice rather than a fixed voice: Edge-TTS and Index-TTS are both named, and one of the update entries mentions adding multi-language TTS voices. So the automatic path is automatic in orchestration, not in the number of external services it depends on.
The pipeline is four stages, and each one is swappable
The video generation flow is given as a four-step chain: script generation, image planning, frame-by-frame processing, and video composition.
The design is described as modular, and each stage can use a different AI model, audio engine or visual style. That modularity is the feature list's last entry, described as atomic capabilities that can be combined flexibly, and it names two ways of doing the same job: ComfyUI or RunningHub workflows, or direct API models.
That distinction matters more than it first appears. A workflow means you are running or renting a graph that someone else assembled, with its own models to manage. A direct API model means one HTTP call to a provider you already have a key for. The project supports both and lets you replace the image, video, TTS or vision capability independently.
Beyond the core path there are three extension modules named in the update log: digital human spokesperson, image-to-video, and motion transfer, the last of which takes a reference video and an image. Custom materials were added so you can upload your own footage and have the script written from it.
Two exact pins in a list of mostly floors
The dependency list is worth reading rather than skimming, because two entries behave differently from the rest.
Everything is a lower bound except two. Almost every requirement uses a greater-than-or-equal form: pydantic, httpx, streamlit, openai, fastapi, playwright, dashscope, numpy and the rest. Exactly two are pinned with an equality: edge-tts at 7.2.7 and moviepy at 1.0.3.
The update log explains the first one. An entry dated 2025-12-10 records that the sidebar got a built-in FAQ and that the edge-tts version was locked to fix an instability in the TTS service. So that pin is a scar, not a preference: a newer edge-tts changed something the pipeline depended on.
The second pin is older technology pinned harder. moviepy 1.0.3 is the previous major line of a library whose later versions changed its API, and holding it exactly is the cheapest way to keep a composition step working.
Pillow is the third entry that needs reading: it is capped as well as floored, at 10 or newer and below 12, so it is bounded from both sides.
Windows users get a bundle, everyone else installs uv and ffmpeg
There are two install paths and one of them exists specifically to avoid the other.
For Windows, the README offers a one-click bundle and says plainly that you do not need to install Python, uv or ffmpeg. You download the bundle from the releases page, extract it, double-click start.bat to launch the web interface, and the browser opens at http://localhost:8501. Then you configure the language model API and the image generation service in the system configuration panel. The bundle is stated to contain every dependency, so the only setup left is API keys.
For macOS and Linux, or anyone who wants to customise, you install the package manager and the video tool first. On macOS:
brew install ffmpegOn Ubuntu or Debian:
sudo apt update
sudo apt install ffmpegThen clone and launch, and uv handles the dependencies:
git clone https://github.com/AIDC-AI/Pixelle-Video.git
cd Pixelle-Videouv run streamlit run web/app.pyThe README asks you to verify both tools with uv --version and ffmpeg -version before going further.
The compose file has an init service that undoes Docker's own behaviour
The container setup contains a service that exists to work around a specific Docker quirk, and the comment above it says so.
The init service runs Alpine and mounts the project directory, then its command checks whether config.yaml is a directory and removes it if so, and creates config.yaml by copying config.example.yaml when the file does not exist but the example does. The stated reason: mounting a non-existent file creates a directory.
That is a real failure mode. The api service bind-mounts ./config.yaml into the container, and if the host file is missing Docker will happily create a directory at that path, after which every write to the config fails in a confusing way. The init service runs once and exits, with restart set to no.
The api service then declares that it depends on init with the condition that the dependency completed successfully, which is the mechanism that makes the ordering real rather than best-effort. It runs uvicorn's application module directly from the virtual environment on port 8000, and the config file is mounted read-write so the web interface can save settings back to it.
One build argument rewrites apt, pip and uv
The Dockerfile is short and has a single switch that changes where everything is downloaded from.
A build argument named USE_CN_MIRROR defaults to false. When it is true, the build rewrites the Debian package sources to point at mirrors.aliyun.com, and the comment explains that Debian 12 writes them in the DEB822 format under etc/apt/sources.list.d, which is why a plain sed against the older sources list would not be enough.
The same switch changes how the package manager itself is installed. With the mirror, uv is installed through pip from the Tsinghua index; without it, it comes from the official installer script fetched over curl. Either way the build ends by printing the uv version, so a failure to install it stops the image build.
The system packages are three: curl, described as being for health checks and downloads, ffmpeg for video and audio processing, and fonts-noto-cjk for CJK character support. That last one is a reminder that the default templates render text into the video itself.
The compose file exposes the switch the same way, and a comment at the top shows both invocations.
The tags say 0.1.15, the manifest says 0.2.0
Two version numbers, and they do not agree.
The three most recent releases are v0.1.12 on 2026-01-14, v0.1.14 on 2026-01-26 and v0.1.15 on 2026-01-27. Every one of their titles ends with the same phrase, referring to the Windows one-click bundle, so the release notes are about the packaged build rather than the library.
The project manifest says version 0.2.0. So the published tags stopped in mid-January at 0.1.15 while the manifest has moved on a minor version.
The last push was on 2026-06-14, close to five months after the newest tag, and the update log's most recent entry is dated 2026-06-01, which adds direct API configuration for image and video providers in the web interface. That work is in the branch and not in any tag.
There is also a naming detail: the repository carries the organisation name in its links, but the clone instructions and the release page point at a different one. Check which organisation you are actually in before you clone.
Editorial conclusion
Pixelle-Video fits someone who publishes short explainer-style videos and wants the mechanical work done for them, and who is comfortable handing an API key to several different providers at once. Do not adopt it expecting a single model to do the work, because the design is explicitly swappable at every stage and you will be configuring more than one service. Before you start, decide which install path you are on: the Windows bundle carries its own dependencies, while a source install puts uv, ffmpeg and a set of provider keys in your hands. And check the version numbers first, because the newest tag and the manifest version do not match.
Frequently asked questions
What is Pixelle-Video?
It is a Python project described as a fully automated short video engine. You enter one topic and it writes the script, generates images or video, synthesises the narration, adds background music and composes the result. The pipeline runs script generation, image planning, frame-by-frame processing and video composition, and the README says no video editing experience is required.
How do I run Pixelle-Video on Windows?
Download the one-click bundle from the releases page and extract it, then double-click start.bat. The bundle contains every dependency, so Python, uv and ffmpeg are not needed. The browser opens at http://localhost:8501, where you configure the language model API and image generation service before generating.
How do I run Pixelle-Video on macOS or Linux?
Install uv and ffmpeg first, then clone the repository and run uv run streamlit run web/app.py, which installs the dependencies and opens http://localhost:8501. The README asks you to verify the tools with uv --version and ffmpeg -version, and gives brew install ffmpeg for macOS and apt for Ubuntu or Debian.
Which AI services does Pixelle-Video call?
The README names DashScope, OpenAI, Seedream, Seedance and Kling for direct image and video model APIs, ComfyUI and RunningHub for workflow-based generation, Edge-TTS and Index-TTS for speech, and GPT, Qwen, DeepSeek and Ollama among the language models. All of it is configured in the web interface's system configuration panel rather than in a file you edit.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ath-maas-pixelle-video)