# anything2explainer: A Claude Code Skill That Produces Motion-Graphics Explainer Videos

> anything2explainer is a Claude Code or Codex skill that takes a topic as input and produces a 1280x720 H.264 MP4 explainer video with TTS voiceover, word-aligned subtitles, and a chapter progress bar. Every frame is rendered from code using Remotion (React and TypeScript), with parallel build agents writing one shot per component.

**Vincentwei1021/anything2explainer** — Topic in, narrated explainer video out. A Claude Code / Codex skill that turns any topic into a black-canvas motion-graphics explainer video with TTS voiceover, subtitles and a chapter progress bar. Chinese or English; every frame drawn in code with Remotion.

- Repository: https://github.com/Vincentwei1021/anything2explainer
- Stars: 2,010 · Forks: 283
- Language: TypeScript
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/vincentwei1021-anything2explainer

## What anything2explainer Does and Who It Is For

Stock footage licensing, generative video models, and screen recording workflows each have their own constraints. anything2explainer avoids all of them by drawing every frame in code. The README states explicitly: "No stock footage, no generative video model, no frames lifted from anyone else's work."

The skill takes a topic as input (for example, "explain vector databases"), or an article or document you want turned into a video, plus a target length and language. It produces a 1280x720 H.264 MP4 with synchronized voiceover, word-boundary-aligned subtitles, chapter cards, a top HUD, and a bottom chapter progress bar. The full paper trail (research document with sources, narration, storyboard, per-shot source code, QC reports, delivery notes) is written to the examples/ directory.

The target users are developers, technical educators, and content creators who want a deterministic, auditable video pipeline and are comfortable running an AI coding agent for an extended session. The skill is not a one-click app: the README describes it as "the whole method an AI coding agent needs to finish the film."

## Installing anything2explainer

Clone the repository and link it as a Claude Code or Codex skill:

```bash
git clone https://github.com/Vincentwei1021/anything2explainer.git
ln -s "$PWD/anything2explainer" ~/.claude/skills/anything2explainer
ln -s "$PWD/anything2explainer" ~/.codex/skills/anything2explainer
```

Then install system and Python dependencies:

```bash
brew install ffmpeg

python3 -m venv ~/.venvs/a2e && source ~/.venvs/a2e/bin/activate
pip install 'edge-tts==7.2.8' numpy pillow scipy
```

For English narration using the kokoro-82m voice:

```bash
pip install kokoro soundfile && brew install espeak-ng
```

The Node package dependencies (Remotion 4.0.507, React 19) are installed by the template's own npm install when the first video is built. The README pins edge-tts at 7.2.8 specifically: the package tracks a Microsoft endpoint and breaks across upgrades, and 7.2.8 is the version tested. scipy is used only by the QC script frame_metrics.py. The shell scripts are zsh and Python 3; the README notes that Linux should work but Windows is untested.

## Output Specification: What the Video Contains

The README specifies the output format precisely:

- Frame rate: 1280x720 at 30 fps, H.264
- Length: 2 to 8 minutes, at the user's choice
- Language: Chinese or English, configurable via the lang setting in src/config.ts
- Look: black canvas with one of two backdrops (star field plus fog gradient, or dot-field wave from video-talkcraft), white line art with purple accents, ultra-bold headline type
- Persistent layers: 44px white-on-black-stroke subtitles, chapter progress bar at the bottom, capsule HUD at the top, optional pipeline rail, and a "built by Anything2Explainer skill" end credit (removable by setting builtBy to empty string)
- Voiceover (Chinese): edge-tts zh-CN-YunxiNeural at approximately 5.5 chars/s
- Voiceover (English): kokoro-82m am_liam locally

Chapter count follows the content and is constrained by length: under 3 minutes uses a single chapter, 3 to 5 minutes uses 3 to 4 chapters of at least 60 seconds each, 5 to 8 minutes uses 4 to 6. The progress bar splits evenly across chapters.

## How the Pipeline Runs: Research, Storyboard, Parallel Build, QC

The README describes the pipeline in four phases. The agent researches the topic with sources and writes a narration. It then generates the voiceover and a frame-accurate timeline, storyboards every shot, and dispatches parallel build agents that write one Remotion component per shot. QC agents review the rendered frames against written criteria before delivery.

The user is consulted at exactly four checkpoints during the process, and the rest runs autonomously. Wall-clock time depends on video length:

- 2 to 3 minutes: approximately 1 hour, about 2 GB disk
- 3 to 5 minutes (the reference tier with 40 to 50 shots): approximately 2 hours, about 2 GB disk
- 5 to 8 minutes: approximately 2 to 3 hours, about 3 GB disk

Paragraphs in the narration mark shot boundaries: sentences within a paragraph are separated by 10 frames, paragraph ends by 30 frames. The finished video runs 5 to 8% longer than the raw speech by design, because each shot holds 1 to 1.5 seconds after its last element lands.

The reference film for the 3 to 5 minute tier is a RAG and Knowledge Bases explainer: 44 shots, 785 words in English, 5 minutes 2 seconds. The Chinese cut of the same content runs 4 minutes 54 seconds and uses 1,490 characters.

## Linux and Raspberry Pi (ARM64) Support

The README documents a verified setup on Raspberry Pi 5 (ARM64, Python 3.13). Three differences from macOS require attention. First, zsh must be installed explicitly (`sudo apt install zsh`). Second, Remotion has no linux-arm64 headless browser, so system Chromium must be installed and pointed to via the REMOTION_BROWSER_EXECUTABLE environment variable read by template/remotion.config.ts. Third, kokoro (the default English TTS) does not install cleanly on ARM or Python 3.13 due to dependency issues with old numpy and spaCy.

Two alternative TTS engines work on Linux ARM without these issues: kokoro_onnx (via onnxruntime, no torch or spaCy required, using .onnx model files from github.com/thewh1teagle/kokoro-onnx), and piper (described as the fastest local option but noted as robotic). Both are passed via the TTS_ENGINE environment variable.

## Limitations of the Approach

The pipeline is slow. A 3 to 5 minute video takes approximately 2 hours of wall clock time with 8 parallel build agents. This is workable for planned production work, but not for fast iteration on content.

The generated visual style is fixed: black canvas, white line art, purple accents. The README describes two backdrop options (star field and dot-field wave). Teams that need brand colors, custom fonts, or a different visual language will need to modify the Remotion template directly. The README does not document a configuration interface for these parameters beyond the bg and lang settings in src/config.ts.

The bring-your-own-TTS path requires a TTS service that outputs audio with word-boundary timing information, because subtitle alignment depends on those timestamps. The README notes that the edge-tts version must be pinned at 7.2.8; newer versions may break word-boundary requests.

## Comparison with Other Explainer Video Approaches

AI-generated video tools (generative video models) take text or image prompts and produce video frames directly. They are faster than anything2explainer but produce frames that cannot be modified as code and whose exact content is not reproducible from the same prompt. anything2explainer's frames are Remotion components: they can be diffed in git, edited precisely, and re-rendered deterministically.

Slide-based presentation tools (PowerPoint with narration, Google Slides with audio) are a common alternative for technical explainers. They produce content faster than anything2explainer but have a fundamentally different look: slides with bullets rather than animated motion graphics. The anything2explainer style is closer to a produced explainer animation than a recorded slide deck.

The license in the repository is marked NOASSERTION in the detected language field, but a LICENSE file is present in the repository root. The repository has no GitHub releases. The last push was on 2026-09-18 and the repository is not archived.

## Conclusion

anything2explainer suits developers and technical communicators who want a repeatable, code-driven pipeline for short explainer videos on technical topics and who are comfortable running Claude Code or Codex for 1 to 3 hours of wall-clock time per video. The output is fully programmatic: every frame is source code, the full paper trail (research, narration, storyboard, QC reports) is saved to files, and the video can be reproduced or modified by editing the Remotion components. The pipeline is macOS-tested and verified on Linux and Raspberry Pi (ARM64), but the README marks Windows as untested. Before starting, verify that ffmpeg is installed and that the chosen TTS engine is reachable, because voiceover generation runs before frame rendering and a failed TTS step stops the pipeline early.

## FAQ

### What TTS engines does anything2explainer support?

The default engines are edge-tts using zh-CN-YunxiNeural for Chinese and kokoro-82m using am_liam for English. On Linux and ARM, kokoro_onnx and piper are documented as working alternatives. A bring-your-own-TTS path is also supported for teams that provide their own finished audio with word-boundary timing data.

### How long does it take to generate an explainer video with anything2explainer?

The README documents approximately 1 hour for 2 to 3 minute videos, 2 hours for the 3 to 5 minute reference tier (with 8 parallel build agents), and 2 to 3 hours for 5 to 8 minute videos. Most of the time is agents building shots in parallel.

### Does anything2explainer require a GPU?

The README does not mention GPU requirements. The English TTS engine (kokoro-82m) runs locally on CPU. The Remotion rendering uses Chromium for frame capture. The README notes a verified run on a Raspberry Pi 5 (ARM64), which has no discrete GPU, suggesting GPU is not required.

## Sources

- [Issues](https://github.com/Vincentwei1021/anything2explainer/issues)
- [README](https://github.com/Vincentwei1021/anything2explainer/blob/main/README.md)
- [Vincentwei1021/anything2explainer on GitHub](https://github.com/Vincentwei1021/anything2explainer)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/vincentwei1021-anything2explainer
