# html-video: Turn HTML, CSS, and Data into MP4 with Coding Agents

> html-video is a meta-layer that lets a coding agent take a prompt, an article URL, or a GitHub repository link and produce a real MP4 video, rendered locally using headless Chromium and ffmpeg. It ships 21 curated templates, supports 14 coding agents including Claude Code and Cursor, and is licensed under Apache-2.0 with no per-render fees.

**nexu-io/html-video** — Programmatic video for coding agents — HTML to video on your laptop. Turn HTML, CSS & data into real MP4s with pluggable render engines, 21 templates, AI soundtrack. Apache-2.0, no per-render fees. An official project by the Open Design team.

- Repository: https://github.com/nexu-io/html-video
- Website: https://open-design.ai/html-video
- Stars: 4,625 · Forks: 563
- Language: HTML
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/nexu-io-html-video

## What html-video solves and who it targets

Every HTML-to-video rendering engine has its own authoring model. Remotion uses React components. Motion Canvas uses TypeScript generators. Manim is built around mathematical and 3D animation. Learning each model and stitching them into a single workflow costs real engineering time, and most teams pick one engine and accept its constraints.

html-video sits above all of these engines as a meta-layer. The idea is that a developer or a coding agent describes a video in natural language or provides a link; the tool picks the right engine and template, fills in the content, and produces an MP4 on the local machine. The README describes the target users as developers who work with coding agents and want to generate video without managing the rendering engine directly. The 14 supported agents listed in the README include Open Design (Vela), Windsurf CLI, Trae CLI, Claude Code, Cursor Agent, Codex CLI, Gemini CLI, Grok Build, Qwen Code, OpenCode, GitHub Copilot CLI, Aider, Hermes, and the Anthropic Messages API. Agents are auto-detected on the system PATH and can be switched from the studio interface.

## How the pipeline converts a prompt to an MP4

The README describes the pipeline in four steps. First, the studio fetches the source: if the input is a URL or a GitHub repository link, the server-side fetcher pulls it and flattens it to Markdown, including WeChat public account articles. If the input is a plain prompt, this step is skipped.

Second, the coding agent reads the source and the selected template's style, then emits a content-graph describing the storyboard. The content-graph is a multi-frame intermediate representation: nodes represent entities, data, or text, connected by edges that describe sequence, dependency, or contrast. The nodes are sorted topologically into frame order and timing. Third, each node in the content-graph becomes a self-contained animated HTML block for one frame. Finally, the Hyperframes engine records those animated HTML frames using headless Chromium and encodes the result to MP4 using ffmpeg with libx264. This entire loop runs locally with no cloud render step and no per-clip fee.

## Project setup and monorepo structure

The repository is a pnpm workspace monorepo. The package.json at the root specifies the required runtime environment:

```json
{
  "engines": {
    "node": ">=20",
    "pnpm": ">=9"
  }
}
```

Building all packages runs the build script recursively across the workspace:

```bash
pnpm -r build
```

The top-level directories include packages/ for the library packages, templates/ for the 21 video templates, docs/, notes/, and research/. The CLAUDE.md at root provides instructions for the Claude Code agent integration, and ATTRIBUTIONS.md tracks third-party asset credits.

The README references a quick start section and a CLI named html-video. The CLI package is identifiable from the smoke test script in package.json as @html-video/cli. The studio is a local browser interface; both the studio and the CLI are described in the README's feature table as available tools for working with the system.

## The 21 templates and what they cover

The template gallery ships 21 curated, license-clean patterns. The README describes several by name and category. On the data visualization side, frame-data-chart-nyt provides an editorial NYT-style animated line chart with a headline, annotated data points, and a source line. For title cards, frame-glitch-title gives a chromatic-aberration glitch effect with scanlines. On the hero side, frame-liquid-bg-hero is an aurora liquid-gradient hero with a centered headline. The frame-light-leak-cinema template produces a warm film-grain cinematic frame.

VFX templates include vfx-text-cursor, which adds a typewriter effect with a blinking terminal cursor. An outro template, frame-logo-outro, provides a clean animated logo end card. The full gallery of 21 templates covers multi-scene product promos, kinetic type, Swiss-grid and Vignelli data cards, decision-tree explainers, Takram-organic motion, and warm-grain editorial frames. All templates are previewed live in the studio gallery.

## Engine architecture: what works now and what is planned

The README is explicit about which rendering engines are available. Hyperframes is the only fully wired engine at the time of the last push on 2026-06-21. It uses headless Chromium to record animated HTML frame-by-frame and ffmpeg to encode the result to libx264. This engine is the default when a render is requested.

Three other engines appear in the README's comparison table. Remotion (React components) is planned. Motion Canvas and Revideo (TypeScript generators on canvas) are planned. Manim (math and 3D animation) is marked as being researched. The README calls the status column in the engine table "the single source of truth for what's actually runnable today," meaning the adapter interface is designed for these engines but their adapters have not been built. Adding a new engine only requires implementing the `render(input, ctx)` adapter interface, so the architecture supports extension without requiring changes to the template layer.

## Current limitations

The three planned engine adapters (Remotion, Motion Canvas/Revideo, and Manim) are not yet available. Teams that have existing Remotion codebases cannot migrate their templates to html-video yet, and teams that need the mathematical animation style of Manim cannot use the meta-layer to drive it. The Hyperframes engine requires headless Chromium and ffmpeg installed locally, which adds setup steps beyond a simple npm install.

The AI soundtrack feature uses MiniMax for background music and narration mixed into the MP4 at export time. The README describes it as optional, but the README does not document which MiniMax pricing tier or API key setup is required for the feature. The repository has no GitHub releases as of the last push on 2026-06-21, which means version tracking and upgrade paths must be managed directly from the main branch.

## How html-video compares to using Remotion directly

Remotion is a well-known open-source library for creating videos using React. Each video is a React component rendered frame-by-frame. Remotion's source-available model applies different license terms above four developers, and it is a dedicated video authoring tool rather than an agent orchestration layer.

html-video's relationship to Remotion is explicitly framed in the README: Remotion appears in the engine table as a planned adapter. Once that adapter is built, html-video would be able to drive Remotion-authored templates through the same content-graph pipeline that currently drives Hyperframes. For developers who already know Remotion's API and are satisfied with it, using Remotion directly avoids the extra abstraction layer. html-video adds value when the goal is agent-driven video generation that can switch between engines depending on the use case.

## Conclusion

html-video is a practical choice for developers who want to generate video from HTML and data locally, without a cloud rendering service or per-clip fee. The critical limitation to verify first is the engine support table: as of the last push on 2026-06-21, only the Hyperframes engine is fully wired up. Remotion, Motion Canvas, Revideo, and Manim adapters are on the roadmap but their adapters have not been built yet. Teams that already use Remotion for video should stick with it directly; html-video adds value when you want to switch engines or drive the workflow through an agent.

## FAQ

### How do I turn HTML to video with html-video?

The html-video pipeline takes a prompt, an article URL, or a GitHub repository link. The studio fetches the source, a coding agent produces a content-graph and per-frame HTML, and the Hyperframes engine uses headless Chromium to record the frames and ffmpeg to encode the result to MP4. The rendering runs locally with no cloud step.

### What coding agents does html-video support?

The README lists 14 supported agents: Open Design (Vela), Windsurf CLI, Trae CLI, Claude Code, Cursor Agent, Codex CLI, Gemini CLI, Grok Build, Qwen Code, OpenCode, GitHub Copilot CLI, Aider, Hermes, and the Anthropic Messages API. Agents are auto-detected on the system PATH and can be switched from the studio interface.

### Is html-video free to use and does it charge per render?

html-video is licensed under Apache-2.0, which allows commercial use without restriction. The README explicitly states there are no per-render fees and no vendor lock-in. The rendering runs locally using the user's own hardware and the installed Hyperframes engine.

## Sources

- [Issues](https://github.com/nexu-io/html-video/issues)
- [License: Apache-2.0](https://github.com/nexu-io/html-video/blob/main/LICENSE)
- [nexu-io/html-video on GitHub](https://github.com/nexu-io/html-video)
- [Project website](https://open-design.ai/html-video)
- [README](https://github.com/nexu-io/html-video/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nexu-io-html-video
