AI Media Assistant: a local browser tool for turning scripts into narrated short videos
AI-Powered Automate Self-Media Video Generation Tools
At a glance
- What is it?
- AI Media Assistant is an MIT local short-video generator for Chinese creators. It handles script editing, per-line images, TTS, BGM and export in the browser, binds to 127.0.0.1 by default, and keeps project data on your machine.
- Who is it for?
- Use AI Media Assistant if you produce narrated, script-driven short videos and want an automated local pipeline, script to per-line images to TTS to export, that keeps your projects and outputs on your own machine and binds to 127.0.0.1 by default.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 94 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A short-video pipeline that runs on your own machine
AI Media Assistant, formerly ai_caption_video, is a local short-video generation tool aimed at Chinese content creators. The README describes a browser-based editor that handles script editing, subtitle templates, per-line images, voiceover, BGM and video export, with all project data and generated results saved locally by default.
The user is a solo creator producing narrated short videos, the kind of script-plus-images-plus-voiceover content common on platforms like Bilibili, who wants a local tool rather than a cloud service. The README makes a point of the privacy posture: the service listens only on `127.0.0.1` by default, so it is meant for personal use on your own computer.
That local-first design is the distinguishing choice. Rather than uploading scripts and media to a cloud renderer, AI Media Assistant keeps the whole pipeline, editing, TTS, rendering and the SQLite project database, on the user's machine. It is MIT licensed, with a React and TypeScript frontend, a FastAPI backend and a separate render worker.
How the pipeline is structured
The architecture separates editing, background work and media logic cleanly, which the README lays out directory by directory. The frontend under `apps/web/` is React, TypeScript and Vite; the backend under `apps/api/` is FastAPI with SQLAlchemy and Alembic; a `workers/render_worker/` runs TTS and video rendering as background tasks so they do not block API requests; and `packages/media_core/` holds the subtitle layout, templates, TTS bridging and audio-video composition core.
That background-worker split matters for a media tool, because TTS and video rendering are slow, and running them in a worker rather than inline keeps the editor responsive and lets tasks be cancelled, retried and previewed. The README lists exactly those task controls.
The content features are specific: multi-line script management where each line is an independent segment with a UUID tracking its image, color and audio; subtitle templates including large centered text, scrolling queues and emotion templates; automatic per-line image matching from a local library that avoids reusing an image within a project; keyword highlighting in subtitles; TTS supporting OmniVoice, Qwen3-TTS, reference audio and preset voices with speed control; and BGM from a built-in library or uploads. A unified backend preview means the browser preview and the final export share the same subtitle layout and cropping logic, so what you see matches what you get.
Installing and running locally
The README provides platform setup scripts run once, then a start script. On Windows, first-run setup and subsequent launch are:
powershell -ExecutionPolicy Bypass -File scripts\setup_windows.ps1
powershell -ExecutionPolicy Bypass -File scripts\start_windows.ps1On macOS the equivalent scripts are made executable and run:
chmod +x scripts/setup_macos.sh scripts/start_macos.sh
./scripts/setup_macos.sh
./scripts/start_macos.shEither way you then open the app in a browser:
http://127.0.0.1:8123For development the README uses a single command after the setup script:
npm run devwhich starts the web app on `http://127.0.0.1:5173`, the API docs on `http://127.0.0.1:8123/docs`, a health check, and the worker consuming TTS and render tasks in the background. The requirements are Windows 10/11 or macOS, Python 3.11 or newer, Node.js 20 or newer, and FFmpeg on the PATH. The README warns macOS users that the system Python may be 3.9, which will not work, so a 3.11+ install is needed.
Models are yours to supply, and platform-bound
The honest limitation is that the AI models are not in the repository or the release. The README states model weights are not included and users must prepare them, configuring the model location in the settings, and recommends placing them under a `models/` folder in the app root. So the TTS quality you get depends on models you download and wire up yourself, and the tool ships the pipeline, not the voices.
The README is also explicit that model runtimes are platform-bound: the Windows and macOS model environments cannot be mixed, because a Windows model package's `python.exe`, DLLs and scripts will not run on macOS, and Mac users need macOS arm64 model packages. That is a real friction point: moving a working setup between operating systems means re-provisioning the model environments, not just copying files.
The tool's scope is another boundary. It is built for narrated, script-driven short videos with per-line images and subtitles, the specific format its templates and image-matching target. It is not a general video editor, and content that does not fit the script-line-plus-image structure is outside its design. The local-only, 127.0.0.1 default also means it is a personal tool, not a multi-user service, by intent.
Against a cloud video generator or a manual editor
The alternatives are a cloud short-video generation service or a manual editor like a timeline-based video tool. A cloud service is turnkey and needs no local setup or model provisioning, but it uploads your scripts and media and charges per use, and you do not control the models. A manual editor gives total creative control but no automation of TTS, per-line image matching or templated subtitles.
AI Media Assistant sits between them: it automates the narrated-short-video pipeline, script to subtitles to voiceover to render, while keeping everything local, with project data and outputs on your machine and the service bound to 127.0.0.1. The cost is that you provision the models yourself and accept the platform-bound model environments. Choose a cloud service for zero setup when you are comfortable uploading content and paying per render. Choose a manual editor when you want full control and no automation. Choose AI Media Assistant when you want an automated, templated narrated-video pipeline that runs entirely on your own machine and keeps your scripts and outputs local.
MIT, local data, and where to start
AI Media Assistant is MIT, so it can be forked, self-hosted and modified with attribution, which fits a local tool creators may want to adapt. The README is clear about what is data versus code: distributable resources like the BGM library, fonts and preset voices live under `storage/resources/`, while runtime artifacts, the SQLite project database, uploads and exported videos, live under `storage/` and are not meant to be committed or shipped.
Upgrade cost is mostly the model environments, since those are downloaded separately, platform-specific and not covered by the app's own updates. The app itself updates through its scripts and, for maintainers, a packaging script that bundles the built frontend and distributable resources while excluding generated videos, uploads and the database.
The concrete first step is to get the environment right before expecting output: on macOS confirm you have Python 3.11 or newer rather than the system 3.9, install FFmpeg on the PATH, run the platform setup script, and open `http://127.0.0.1:8123`. Then provision a TTS model, OmniVoice or Qwen3-TTS, into the `models/` folder and set its location in Settings, since without a model the pipeline can edit and lay out but cannot produce voiceover.
Editorial conclusion
Use AI Media Assistant if you produce narrated, script-driven short videos and want an automated local pipeline, script to per-line images to TTS to export, that keeps your projects and outputs on your own machine and binds to 127.0.0.1 by default. It is the wrong tool if you want a zero-setup cloud service or a general-purpose editor, or if you cannot provision the TTS models yourself, since weights are not included and model environments are platform-bound between Windows and macOS. Before expecting output, install Python 3.11 or newer (not macOS's system 3.9) and FFmpeg, run the platform setup script, open http://127.0.0.1:8123, and place a TTS model in the models/ folder with its location set in Settings.
Frequently asked questions
Does AI Media Assistant run locally or in the cloud?
Locally. The README says all project data and generated results are saved on your machine by default and the service listens only on 127.0.0.1, so it is meant for personal use on your own computer.
Are the AI models included with AI Media Assistant?
No. The README states model weights are not in the repository or the release and must be prepared by the user, placed in a models/ folder and configured in Settings. Windows and macOS model environments cannot be mixed.
What do I need to run AI Media Assistant?
The README lists Windows 10/11 or macOS, Python 3.11 or newer, Node.js 20 or newer, and FFmpeg on the PATH. macOS users must avoid the system Python 3.9. You run a platform setup script, then open http://127.0.0.1:8123.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/alexchan197611-ai-media-assistant)