Open-source project
SamurAIGPT/Text-To-Video-AI avatar
SamurAIGPT/Text-To-Video-AI

Text-To-Video-AI: a Python pipeline that assembles narrated shorts from a topic string

Generate video from text using AI

833 stars321 forksJupyter NotebookMIT

At a glance

What is it?
SamurAIGPT/Text-To-Video-AI is an MIT-licensed Python and Remotion pipeline that turns a topic into a scripted, voiced, captioned video. It is glue, not a diffusion model: the generated footage comes from paid Muapi video models, and Pexels is required for every run.
Who is it for?
Adopt it if you want a readable, hackable pipeline for narrated vertical shorts and you are willing to hold API keys for Pexels, an LLM provider, and Muapi. Do not adopt it if you expect a free or offline generator, or if you need a stable CLI contract.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 37 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Text-To-Video-AI actually generates, and for whom

The project takes a single topic string and produces a finished video file at rendered_video.mp4. The README frames the output for YouTube Shorts, Instagram Reels and TikTok, and the default orientation is portrait at 1080x1920, with landscape at 1920x1080 as the alternative. The intended user is someone who wants a repeatable short-form video assembly line rather than a prompt box.

The word "generated" needs unpacking. The repository does not ship a video diffusion model. The pipeline writes a script with an LLM, synthesizes a voiceover, fetches background clips from Pexels, and can call Muapi models for AI B-roll. The sample videos in the README are described as combining local EdgeTTS voiceover, whisper-timestamped word captions, Suno AI background music, and Veo3 B-roll segments. Suno is not a configuration option in .env.example, so music is part of the author's own workflow rather than a feature you can switch on.

If your goal is a text-to-video generator you run with no accounts and no keys, this is the wrong shape of tool. Pexels is listed as always required.

The stage chain: script, speech, captions, B-roll, render

Configuration is entirely environment-driven. .env.example exposes three swappable providers: LLM_PROVIDER accepts openai, groq or gemini; TTS_PROVIDER accepts edgetts or elevenlabs; STT_PROVIDER accepts whisper or deepgram. Each choice changes which key the run needs, which is why the setup step is copying .env.example to .env rather than editing code.

The data flow runs one direction. An LLM produces the script from your topic. A TTS provider turns that script into audio. An STT provider produces word-level timings that drive captions, and the README credits whisper-timestamped captions for the sample videos. Pexels supplies background footage. Muapi supplies AI-generated clips when MUAPI_VIDEO_MODEL is set. Rendering is handled by a React/Remotion composition engine in the remotion-composer directory, which is why the repository is classified as Jupyter Notebook but contains a Node-side renderer and a Colab notebook, Text_to_Video_example.ipynb.

The Muapi model list is the most detailed part of the README. It spans Google Veo (veo3-fast-text-to-video through veo3.1-lite-text-to-video), xAI Grok, ByteDance Seedance, Wan, LTX, Kling, Vidu, OpenAI Sora, MiniMax and Alibaba, with the default set to veo3-fast-text-to-video. That breadth is a routing layer over other vendors' models, and each model name carries its own latency and price that the README does not list.

Installing Text-To-Video-AI locally and rendering a first short

Prerequisites are Python 3.8+, FFmpeg and ImageMagick. Windows users are pointed at INSTALL_WINDOWS.md, which matters because ImageMagick and FFmpeg are the two dependencies most likely to fail on that platform.

Clone the repository and install the pinned dependency set. The requirements file pins exact versions, including torch==2.3.1 and openai-whisper==20231117, so expect a large download on the first install.

bash
git clone https://github.com/SamurAIGPT/Text-To-Video-AI.git
cd Text-To-Video-AI
pip install -r requirements.txt
cp .env.example .env

Open .env and fill in the keys for the providers you selected. Pexels is required on every run. If you leave LLM_PROVIDER=openai, you need OPENAI_API_KEY; the defaults in .env.example are OPENAI_MODEL=gpt-4o, STT_PROVIDER=whisper and TTS_PROVIDER=edgetts, so no Deepgram or ElevenLabs key is needed for a first pass.

env
LLM_PROVIDER=openai
OPENAI_API_KEY=your_openai_api_key_here
PEXELS_API_KEY=your_pexels_api_key_here
TTS_PROVIDER=edgetts
STT_PROVIDER=whisper
VIDEO_ORIENTATION=portrait

Then run the entry point with a topic. The README gives this exact form, and the output lands in the working directory as rendered_video.mp4.

bash
python app.py "Your topic here"

Expect the first run to be slow: Whisper downloads model weights, EdgeTTS needs network access, and Remotion renders frame by frame. If you would rather not install anything, the README offers a Colab notebook and a hosted API at docs.vadoo.tv as alternatives to local setup.

Where the pipeline breaks, and who should not use it

The dependency on Pexels for every run is the sharpest constraint. A Pexels outage, a revoked key or a rate limit stops the run even when your LLM and TTS providers are healthy, and the README does not document a fallback or a cache. For a pipeline whose selling point is automation, that is a single point of failure sitting outside your control.

The Muapi path is metered. Every model name in the supported list points at a third-party service, and the README gives no pricing, no per-model latency and no failure guidance. There is also no documented rollback: if a render fails partway through, the README does not describe resuming from an intermediate artifact, so a failed run means starting over. The same silence applies to retries and rate-limit handling.

Caption styling is a fixed enum rather than free-form. CAPTION_POSITION accepts center, top, bottom, bottom_center, bottom_left or bottom_right, and font color is limited to white, yellow, cyan, red, green, blue or magenta. If your brand needs a hex color or a custom font file, you are editing the renderer, not the config.

One more thing to check before trusting the interface: the repository contains app.py, but the README's usage section does not document command-line flags, so it is unclear what app.py accepts beyond a topic string. Treat the CLI as unstable and read the source if you plan to wrap it.

How it differs from MoneyPrinterTurbo and short-form editors

The closest comparison is MoneyPrinterTurbo, which targets the same job: topic in, narrated vertical short out, with LLM scripting, TTS and stock footage. The difference is the rendering layer. MoneyPrinterTurbo renders with MoviePy in Python, which keeps the whole stack in one language and makes the composition logic ordinary Python you can read and patch. Text-To-Video-AI splits rendering into a Remotion composition in remotion-composer, which gives you React-based timeline control and reusable components at the cost of a Node toolchain alongside the Python one. The repository still lists moviepy==1.0.3 in requirements.txt, so both worlds are present.

The second difference is the AI-footage layer. Text-To-Video-AI routes through Muapi to Veo, Sora, Kling, Seedance, Wan and others behind one MUAPI_API_KEY and one MUAPI_VIDEO_MODEL switch. That is a real convenience if you want to swap models without rewriting integration code, and a real dependency if Muapi changes its catalog or pricing.

Against a browser editor like CapCut, the trade is inverted. You give up direct manipulation of the timeline and gain a scriptable command that can run unattended. That only pays off if you are producing many videos from a template, not one video you want to fine-tune by hand.

Maintenance, upgrade cost and the MIT licence

The last push to the default branch was on 2026-08-24, so the repository is not dormant, but there are no retrieved releases, which means upgrades arrive as commits rather than versioned tags. The requirements file pins exact versions across roughly fifty packages, including torch, numpy and numba, so a fresh install resolves to a fixed, older set rather than current releases. Upgrading any single package risks conflicting with the pins around it, and the README does not describe a supported upgrade path.

The practical cost is the external surface. A pipeline that calls an LLM provider, a TTS provider, an STT provider, Pexels and Muapi can break without a single commit to this repository, when a provider deprecates a model name or changes an endpoint. The .env.example defaults (gpt-4o, llama3-70b-8192, gemini-2.5-flash) are the names most likely to age.

The licence is MIT, which permits commercial use, modification and redistribution with the copyright notice and permission notice retained. That covers this repository's code only. The Muapi models, Pexels footage, ElevenLabs voices and any music you add carry their own terms, and the README does not address them. Check those separately; this is not legal advice.

Editorial conclusion

Adopt it if you want a readable, hackable pipeline for narrated vertical shorts and you are willing to hold API keys for Pexels, an LLM provider, and Muapi. Do not adopt it if you expect a free or offline generator, or if you need a stable CLI contract. Before committing, run one topic end to end and confirm rendered_video.mp4 exists, then read .env.example to see which providers each stage will bill.

Frequently asked questions

Is Text-To-Video-AI free to use?

The code is MIT-licensed, and the default TTS and STT providers, edgetts and whisper, are marked free in .env.example. Pexels is listed as always required, and Muapi AI video generation needs MUAPI_API_KEY, so a full run is not cost-free.

How do I turn text into video with Text-To-Video-AI for free?

Clone the repository, install requirements.txt, copy .env.example to .env, and set PEXELS_API_KEY plus one LLM key, keeping TTS_PROVIDER=edgetts and STT_PROVIDER=whisper. Then run python app.py "Your topic here" and the output is saved as rendered_video.mp4.

What is Text-To-Video-AI?

It is an MIT-licensed pipeline that generates a script from a topic with an LLM, voices it, times captions with Whisper or Deepgram, pulls background footage from Pexels, and renders the result through a React/Remotion composition engine.

Is there an AI that turns text into video?

Text-To-Video-AI does this by chaining an LLM script, a TTS voiceover, timed captions, Pexels background footage and optional Muapi AI clips, then rendering the result to rendered_video.mp4.

How do I make text to video AI for free?

Keep the free defaults, TTS_PROVIDER=edgetts and STT_PROVIDER=whisper, and supply PEXELS_API_KEY plus one LLM key. Free AI video generation through Muapi still requires MUAPI_API_KEY, so the AI B-roll step is the paid part.

What is the best text to video AI?

The README does not rank providers. It exposes many Muapi model names, such as veo3-fast-text-to-video, kling-v3.0-pro-text-to-video and openai-sora-2-pro-text-to-video, and leaves the choice to MUAPI_VIDEO_MODEL without publishing quality or latency comparisons.

Official sources

  1. Issues
  2. License: MIT
  3. Project website
  4. README
  5. SamurAIGPT/Text-To-Video-AI on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/samuraigpt-text-to-video-ai.svg)](https://hysenlabs.com/projects/samuraigpt-text-to-video-ai)